Business Oxford Metrics Pays £525,000 for Move AI After a Competitive Bid, Then Cuts Its Own Guidance News Grok Imagine Video 1.5 Lite Reaches the API at $0.02 Per Second, Ranking Two Spots Above Veo 3.1 at a Third of the Price News Humans Score 93 Percent, the Best AI Manages 53.6: Scale AI Sixth Sense Benchmark Finds a 40-Point Gap in Visual Common Sense Security Ten AI Giants Promise UK Data-Protection Changes While the Regulator Questions OpenAI, Anthropic and Meta Over Agents That Reached Hugging Face AI Agents Manus Founders' China Exit Bans Are Lifted After the $500 Million Reboot: Singapore HQ Stays, Beijing Hiring Starts News Hugging Face Maps 566 Million Predicted Genes: Carbon-A Reads Raw DNA With a 1.2B-Parameter Open Model Productivity Google's New Meeting Notes App Never Calls the Cloud: AI Edge Foresight Runs a 740M-Parameter Model Entirely on Your Mac Business Nvidia-Backed Firmus Pulls the Plug on Australia's Largest IPO in Decades: It Wanted a $30 Billion Valuation, and Buyers Refused Business Oxford Metrics Pays £525,000 for Move AI After a Competitive Bid, Then Cuts Its Own Guidance News Grok Imagine Video 1.5 Lite Reaches the API at $0.02 Per Second, Ranking Two Spots Above Veo 3.1 at a Third of the Price News Humans Score 93 Percent, the Best AI Manages 53.6: Scale AI Sixth Sense Benchmark Finds a 40-Point Gap in Visual Common Sense Security Ten AI Giants Promise UK Data-Protection Changes While the Regulator Questions OpenAI, Anthropic and Meta Over Agents That Reached Hugging Face AI Agents Manus Founders' China Exit Bans Are Lifted After the $500 Million Reboot: Singapore HQ Stays, Beijing Hiring Starts News Hugging Face Maps 566 Million Predicted Genes: Carbon-A Reads Raw DNA With a 1.2B-Parameter Open Model Productivity Google's New Meeting Notes App Never Calls the Cloud: AI Edge Foresight Runs a 740M-Parameter Model Entirely on Your Mac Business Nvidia-Backed Firmus Pulls the Plug on Australia's Largest IPO in Decades: It Wanted a $30 Billion Valuation, and Buyers Refused

Humans Score 93 Percent, the Best AI Manages 53.6: Scale AI Sixth Sense Benchmark Finds a 40-Point Gap in Visual Common Sense

Humans Score 93 Percent, the Best AI Manages 53.6: Scale AI Sixth Sense Benchmark Finds a 40-Point Gap in Visual Common Sense

Scale AI and Elorian have released Humanity's Sixth Sense, a benchmark of 522 open-ended tasks spanning 288 images and 234 video clips totalling 17.6 hours, testing whether models can infer causes, consequences and social cues from a single glance. Humans scored 93.1 percent; the best model, GPT-6-astra at maximum reasoning effort, reached 53.6 percent, and the median of 25 models was 30.9 percent. Social understanding was the weakest domain, averaging 24.4 percent.

Frontier models can now clear mathematical olympiad problems and write production code, but Scale AI and the research group Elorian have put a number on something they still cannot do: look at a photograph and read the room. Their new benchmark, Humanity's Sixth Sense (HSS), released on 7 October, shows humans scoring 93.1 percent on intuitive visual reasoning while the strongest AI model manages 53.6 percent — a 40-point gap on tasks most adults would call easy.

The benchmark consists of 522 open-ended tasks: 288 image-based and 234 video-based, drawn from 17.6 hours of footage with a median clip length of 76 seconds. Each item pairs a visual with a human-written prompt probing something people infer at a glance — the spatial layout of a scene, a likely cause and effect, an unwritten social rule, or an abstract pattern. Tasks span four domains (temporal and causal dynamics, physical and spatial logic, social understanding, and abstract and contextual inference) across 11 subdomains. Model outputs are graded by Claude Opus 5 acting as an automated judge against a rubric where a response must meet every criterion to pass.

The results are lopsided. Twenty human annotators established the 93.1 percent baseline. Across the 25 models tested, the median score was 30.9 percent. The strongest performer, OpenAI's GPT-6-astra running at its maximum reasoning setting, reached 53.6 percent — still roughly 40 points behind the human baseline.

The breakdown is more revealing than the headline. Social understanding — theory of mind, emotional states, who defers to whom — was the single worst domain for 21 of the 25 models, averaging just 24.4 percent accuracy against 34.1 percent across the other three domains combined. Video proved harder than static images, with models dropping an average of 7.3 percentage points on moving footage. Scale also observed that models consume an average of about 4,000 reasoning tokens per task, frequently overthinking without reaching the correct conclusion.

The gap matters more than it might sound. Spatial reasoning, causal inference and social understanding are exactly the skills a warehouse robot needs to avoid bumping into a person, or a customer-facing assistant needs to register that a user is frustrated. Frontier labs spent 2026 posting record scores on math and coding benchmarks — OpenAI's own release notes have GPT-6-astra clearing 98 percent on FrontierMath Tier 4 and posting near-perfect results on ARC-AGI-3 interactive reasoning — while the same model, on Scale's evidence, misreads social situations a six-year-old navigates without instruction.

The dataset is available for research on Hugging Face, and Scale has released the full evaluation harness, including the model registry and grading prompts, so new models can be tested reproducibly. Scale AI is not new to this exercise: it built Humanity's Last Exam with the Center for AI Safety, a benchmark of roughly 2,500 expert-level questions that has climbed from single digits at its early 2025 launch to above 53 percent by mid-2026.

The honest reading of HSS is not that today's models are dumb, but that intelligence is not one thing. A system can out-prove a mathematician and still misjudge why two people in a photo look uncomfortable standing next to each other. For anyone building agents meant to operate in the physical world rather than in a chat window, 93.1 versus 53.6 is the number to watch.

Comments (0)

Log in to join the discussion

Log In

No comments yet