Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag

NVIDIA Says AI Coding Agents Flunk 81% of BlueField DPU Tasks - Its Fix Is Not a New Model but a Stack of Skill Files

NVIDIA Says AI Coding Agents Flunk 81% of BlueField DPU Tasks - Its Fix Is Not a New Model but a Stack of Skill Files

NVIDIA has published DOCA agent skills - SKILL.md files carrying verified API signatures, pkg-config names and hardware constraints for its BlueField DPUs. In the company's own 65-prompt evaluation, agents without them satisfied just 19% of checklist items and misused APIs on 59 prompts; with them, 100%. The results are vendor-graded, so the honest test is your own tasks.

NVIDIA has published a set of agent skills for DOCA, the software platform for its BlueField DPUs, on its NVIDIA/skills GitHub repository. Each skill is a SKILL.md file scoped to one DOCA component - Flow, GPUNetIO, PCC, RDMA and more - carrying the component's real function signatures, pkg-config module names, build-container constraints and known failure modes. NVIDIA's framing is pointed: these are machine-readable specifications an agent can reason against directly, not documentation summaries.

The problem being addressed is familiar to anyone who has pointed a general-purpose coding agent at niche infrastructure. When asked to set up a DOCA comm channel or configure an RDMA context, an agent is pattern-matching across general training data - not reasoning from verified API contracts or hardware capability manifests. The DOCA library surface is large, fast-moving and hardware-specific in ways that training data does not capture, so agents confidently invent flags, image tags and functions that do not exist.

NVIDIA quantified the gap with 65 real developer prompts - ranging from one-line questions to detailed multi-requirement tasks - each graded against a pass/fail checklist. Without the skills, agents satisfied only 19% of graded checklist items. The failures clustered: API or flag misuse on 59 of 65 prompts, unverified hardware capability on 46, wrong tool routing on 39, skipped smoke tests on 34, guessed versions on 30. With the skills loaded, NVIDIA says agents satisfied 100% of checklist items across all 65 prompts, and won every one of the 63 prompts testing API accuracy. All of these are company-reported figures from NVIDIA's own test.

The side-by-side demo is the most tangible part. Two agents were asked to build a Go-based RDMA application on a BlueField-3 DPU; both succeeded, but the with-skills agent wrote 189 lines of handwritten code versus 695 - 73% less - and ran 20 hardware commands versus 37, a 46% reduction. On the highest-complexity task, a firmware-level parameter change on a production DPU, the with-skills agent produced a full discipline - preflight inventory, an out-of-band path assumption, an explicit maintenance window, a rollback plan, and the note that writes of this class take effect only on a cold power cycle - while the unaided agent met none of those requirements.

The caveats are real and worth stating plainly. NVIDIA ran and graded its own evaluation, and a perfect score on a vendor's own checklist is the least surprising number in the announcement - the checklist rewards process steps, like hardware verification and smoke tests, that the skills themselves prescribe. The 19-to-100 gap says nothing about a team whose tasks differ from NVIDIA's 65 prompts, and an answer can tick every checklist item and still fail to build on the card. There is also a maintenance risk specific to this format: a stale signature in a SKILL.md hands the agent a wrong answer labeled as verified, which is arguably worse than an honest guess.

Still, the pattern here is bigger than one product line: instead of retraining a model or fine-tuning on docs, put the domain knowledge in open files the agent loads, where anyone can correct or extend it without waiting on a vendor release. For teams writing DOCA code with AI agents, the practical takeaway is trivially cheap - load the skills. The interesting question is whether NVIDIA publishes the 65 prompts and checklists so the comparison can be rerun with independent agents and, eventually, judged on code that actually compiles and runs on real hardware.

Comments (0)

Log in to join the discussion

Log In

No comments yet