OpenAI's Codex chief has put his team on a public scoreboard. Thibault Sottiaux, the core products and platform lead who runs Codex and ChatGPT Work, posted on X late on October 4 US Eastern time: for the next 28 days, "each day we'll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset." Based on the posting date, the window runs at least from October 5 through November 2. The post cleared a million views within hours, and OpenAI support engineer Romain Huet chimed in: "DevDay was just day one. 28 days of shipping ahead!"
The fine print matters more than the headline. The bar is not "ship something" — by the text of the pledge, a niche CLI flag does not count; the improvement must be relevant to the median Codex or Work user. And the reset is the fallback, not a bonus: this is an either/or, not 28 guaranteed giveaways, despite viral summaries that claimed otherwise. Sottiaux also left the mechanics vague — he did not explain how a reset would work, which users or quotas it would cover, or whether it would stack with existing allowances.
The pledge landed one day after a scope-freeze post that reads as an internal admission. Sottiaux said the team is "locking in" and that the only things being worked on are "simplifications, more efficiency for more usage, groundbreaking features or new models." He conceded the point users had been making for weeks: "sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler." That feedback is not abstract — DevDay on September 29 shipped more than twenty updates at once, and community commentary since has zeroed in on the sprawl: too many changes to follow, tiers that got smaller, and a $500 Pro 500 plan that many users find hard to justify.
Why now? Because September and early October were bruising. A Codex and Work outage on September 25–26 was followed by a paid-account limits reset; a quality postmortem around GPT-6 Astra brought another. The GPT-6.1 Sol launch drew a load spike that degraded speeds for days — Sottiaux apologized on October 2 and announced a global reset for all paid ChatGPT accounts, then spent October 3 chasing reports that some Pro 500 subscribers had missed it. User complaints about usage limits have been persistent and specific: in extreme cases a single runaway agent task can burn most of a week's quota, and frequent, unpredictable resets make it hard to plan real work. Notably, the pledge does not cancel the scheduled Pro 200 reduction — the plan's multiplier is still set to drop from 20x to 10x on October 30.
There is also a competitive shadow over the whole exercise. The pledge arrived days after Claude Opus 5.5 began winning developer mindshare — one community tally even placed it 0.84 points ahead of GPT-6 Astra atop the Epoch Capabilities Index, a narrow but quotable margin. The banter turned direct: Lauren Tan of the Grok Bot team needled OpenAI as "The Ship Company," and Sottiaux publicly challenged her to make the same 28-day commitment. Her reply: "only you can win this battle."
Stripped of the drama, the bet is a clever piece of trust accounting. Weeks of resets have taught Codex users to read Sottiaux's posts the way airline passengers read delay notices; converting that reflex into a daily public commitment makes reliability measurable instead of apologetic. It also puts usage limits — the single most complained-about part of Codex — inside the product roadmap rather than in apology threads. What it cannot guarantee is the harder thing: that 28 days of shipping closes a model gap, or that "a clear improvement" survives contact with users' actual workloads. The pledge's own wording sets the test. If a day's update would not matter to most Codex and Work users, they are owed a reset — and for once, the burden of proof sits with the shipper. All commitments described here, it should be noted, are company statements rather than independently verified outcomes.
The bet immediately collided with its own comment section. In a feature-request thread posted the same weekend — "what is the one thing missing from Codex that you most wish we had?" — the replies queued up behind a single answer that does not ship from OpenAI: Claude Opus 5.5. Chinese-language coverage of the thread counted close to 7,000 replies, with the top-voted responses chanting the competitor's model name. The irony writes itself: OpenAI's own developer-relations exercise doubled as an organic poll on what its users actually want, and the winning write-in was Anthropic's flagship.
Comments (0)
Log in to join the discussion
Log InNo comments yet