Day 2 of OpenAI's 28-day shipping pledge turned out to be a four-for-one. The pledge, made publicly by Codex chief Tibo Sottiaux, commits the company to shipping one clear improvement for most Codex and Work users every day for 28 days — or, on any day it misses, triggering a full usage reset for everyone. Day 1 raised default output speed for GPT-6 Astra and GPT-6.1 Sol by roughly 50% through pure infrastructure optimization. Day 2 shipped four changes at once, spanning Codex, ChatGPT and the developer API.
The headline change: Auto-review — the "Approve for me" option in Codex's permission menu — is now free for all users signed in with a ChatGPT account, and the review process no longer counts against plan usage. Sottiaux said the review step previously could consume 2% to 10% of a plan's quota. The feature targets the approval-fatigue problem that sets in when an agent runs a long task: under the default sandbox mode, every operation that crosses a permission boundary — an outbound network request, a file modification outside the workspace — used to demand a manual click.
With Auto-review, a second agent does the watching. The primary agent keeps executing; whenever it requests to cross the sandbox boundary, a separate reviewer agent receives a condensed chat history, the relevant tool calls and results, and the exact permission request, then judges the action against the user's original intent. It focuses on intercepting high-risk behavior — data exfiltration, destructive deletion, running untrusted code. If it approves, execution continues; if it rejects, the main agent either finds a safer route or stops and asks the user.
The numbers OpenAI attached are company-internal and unaudited. Earlier data put the interruption frequency of auto-review at roughly 1/200 of manual approval mode. A deployment diagram Sottiaux posted with the announcement shows 10,000 actions, of which 9,280 ran directly inside the sandbox; the remaining 720 went to auto-review, which approved 713 and rejected 7 — four of those seven continued via safer alternatives, and just three ended up in front of a human. OpenAI notes the ratios shift with task, environment and sandbox configuration. Sottiaux recommended Auto-review as a replacement for Full access, while cautioning that the reviewer can misjudge, does not cover everything that happens inside the sandbox, and is not a security guarantee.
The second change simplifies the API's usage ladder from five paid tiers to three — Build, Launch and Grow — with upgrade conditions now at $5, $100 and $500 in cumulative credit purchases. The threshold for the highest tier drops from $1,000 to $500, halving the amount a team must spend to unlock top-tier rate limits. OpenAI says the change affects call capacity only; per-token model prices are untouched.
Third, a Meetings plugin is coming to the ChatGPT desktop app on macOS in beta for Pro and Business users, with Enterprise support to follow — OpenAI's help documentation currently lists only a small subset of Enterprise customers in alpha testing. The plugin records meetings and, combined with what ChatGPT already knows about the user and their ongoing work, saves personalized summaries and next steps into ChatGPT Space. Records can be kept private or shared with a team, and ChatGPT can then update project plans or draft follow-up materials — turning the meeting itself into a context entry point for subsequent agent work rather than a transcript someone has to file.
Fourth, the Decisions API — previewed at DevDay on September 29 — is now in public beta. Powered by GPT-6 Luna, it accepts text and image input and returns structured judgments in three forms: Predicates (the probability a statement is true), Choices (a pick from predefined options, with confidence) and Scores (a mapping to a numeric range). OpenAI positions it for complaint triage, image checks, model and tool selection, and risk screening before agent tool calls, and says it runs up to 10 times faster than calling Luna through the Responses API — a figure from the company's own DevDay demonstration, which measured roughly 150 milliseconds against 1.6 seconds. Engineer Steven Heidel noted the timeline: first prototype on September 22, public launch two weeks later.
The framing writes itself: Decisions API is OpenAI's answer to Jev, the fast, cheap decision model from TypeSafe AI that has been quietly absorbing classification and routing workloads across the industry. TypeSafe CEO Diego Almeida, an ex-OpenAI engineer, joked on X about the start of a "clone war," and argued the move validates building products the "System One" way — fast, intuitive judgment rather than deliberate reasoning. Whatever the moats turn out to be, the pattern is clear: Jev found a real market, and now everyone — OpenAI included — is shipping their own. In the meantime, the 28-day clock keeps running: every day without a shipped improvement is a day everyone's quotas get reset.
Comments (0)
Log in to join the discussion
Log InNo comments yet