Frontier models can now clear mathematical olympiad problems and write production code, but Scale AI and the research group Elorian have put a number on something they still cannot do: look at a photograph and read the room. Their new benchmark, Humanity's Sixth Sense (HSS), released on 7 October, shows humans scoring 93.1 percent on intuitive visual reasoning while the strongest AI model manages 53.6 percent — a 40-point gap on tasks most adults would call easy.
The benchmark consists of 522 open-ended tasks: 288 image-based and 234 video-based, drawn from 17.6 hours of footage with a median clip length of 76 seconds. Each item pairs a visual with a human-written prompt probing something people infer at a glance — the spatial layout of a scene, a likely cause and effect, an unwritten social rule, or an abstract pattern. Tasks span four domains (temporal and causal dynamics, physical and spatial logic, social understanding, and abstract and contextual inference) across 11 subdomains. Model outputs are graded by Claude Opus 5 acting as an automated judge against a rubric where a response must meet every criterion to pass.
The results are lopsided. Twenty human annotators established the 93.1 percent baseline. Across the 25 models tested, the median score was 30.9 percent. The strongest performer, OpenAI's GPT-6-astra running at its maximum reasoning setting, reached 53.6 percent — still roughly 40 points behind the human baseline.
The breakdown is more revealing than the headline. Social understanding — theory of mind, emotional states, who defers to whom — was the single worst domain for 21 of the 25 models, averaging just 24.4 percent accuracy against 34.1 percent across the other three domains combined. Video proved harder than static images, with models dropping an average of 7.3 percentage points on moving footage. Scale also observed that models consume an average of about 4,000 reasoning tokens per task, frequently overthinking without reaching the correct conclusion.
The gap matters more than it might sound. Spatial reasoning, causal inference and social understanding are exactly the skills a warehouse robot needs to avoid bumping into a person, or a customer-facing assistant needs to register that a user is frustrated. Frontier labs spent 2026 posting record scores on math and coding benchmarks — OpenAI's own release notes have GPT-6-astra clearing 98 percent on FrontierMath Tier 4 and posting near-perfect results on ARC-AGI-3 interactive reasoning — while the same model, on Scale's evidence, misreads social situations a six-year-old navigates without instruction.
The dataset is available for research on Hugging Face, and Scale has released the full evaluation harness, including the model registry and grading prompts, so new models can be tested reproducibly. Scale AI is not new to this exercise: it built Humanity's Last Exam with the Center for AI Safety, a benchmark of roughly 2,500 expert-level questions that has climbed from single digits at its early 2025 launch to above 53 percent by mid-2026.
The honest reading of HSS is not that today's models are dumb, but that intelligence is not one thing. A system can out-prove a mathematician and still misjudge why two people in a photo look uncomfortable standing next to each other. For anyone building agents meant to operate in the physical world rather than in a chat window, 93.1 versus 53.6 is the number to watch.
Comments (0)
Log in to join the discussion
Log InNo comments yet