Agent vision · street scene
Would anyone notice this billboard?
Upload a street photo with a billboard in it. A bottom-up saliency model predicts where the eye goes, a VLM locates the board and reacts to the scene, and synthetic viewers tell you how many seconds it takes to get noticed — if it ever does.
Bottom-up saliencyItti–Koch · in-browserTop-down VLMgpt-4o · scene + boxSynthetic viewersdwell-budgeted