Microsoft has published a field manual for a problem most SDK teams have only started to notice: their software is now being used primarily by machines that learned it from stale documentation. The AX Practitioner Playbook, released October 9 by the Developer Relations group behind Microsoft's Agent Experience article series, packages the company's method for evaluating what AI coding agents do with a technology, diagnosing the failure, and correcting it at the source — the docs, MCP tools, skills, plugins, CLIs and APIs agents actually read. It is downloadable as a PDF from aka.ms/ax-playbook.
The premise is blunt. Ask a coding agent to build something with your platform and it will often generate code that looks right and might even compile, while quietly picking the wrong SDK version, a deprecated authentication pattern, or a setup no product engineer would recommend. The agent did not hallucinate at random; it did what its training data taught it. Microsoft argues that waiting for models to improve is not a strategy, because a model's knowledge cutoff says little about what it knows about your product — what you can change is the material agents rely on.
A large part of the playbook is about making evaluations honest. The team recounts evaluations that lied convincingly: perfect scores awarded to code that never compiled, and a "used Platform X" check that passed whether or not the agent used Platform X. The playbook's evaluation model defines what a result needs before it can be trusted, from criteria that judge meaning to gates that prove the code runs, and walks through how to write criteria a judge rules on consistently, calibrate them before relying on them, and version them as the product changes. It also flags the seductive trap of letting a model write your evaluation criteria for you.
Diagnosis gets its own framework. A readout tells you what failed; the trajectory tells you why. The playbook trains practitioners to distinguish three cases that require different fixes — an extension that never loaded, one that loaded but was never called, and one that was called but applied incorrectly — and catalogs nine failure patterns that recurred across every technology the team evaluated, each pointing to the surface worth inspecting first.
The method has already produced measurable changes. Evaluations on Cosmos DB, SharePoint Framework and Microsoft 365 Copilot extensions led to dozens of shipped fixes to documentation and agent tooling, including 46 improvements to the Azure Cosmos DB Agent Kit tracked as closed issues on GitHub. A SharePoint Framework project upgrade scenario runs end to end through the playbook, from initial scenario design to the fixes that shipped, and serves as the worked example.
Alongside the PDF, Microsoft released the AX Practitioner skill, which developers can install into their own coding agent to coach them through an evaluation — how to word a criterion, why an agent ignored their extension — and review their scenarios against the playbook's standards. The skill answers only from the playbook: when a question goes beyond its contents, it says so. Across a 330-question test, its answers averaged 95% fidelity to the source document.
The audience is anyone whose technology is consumed by agents: SDK, API, service, CLI and MCP server builders, skill and plugin authors, and the technical writers behind all of it. The method is deliberately tool-agnostic — the playbook lists the capabilities an evaluation system needs rather than mandating Microsoft's own tooling.
What the release marks is the formal arrival of agent experience as an engineering discipline. Just as the last decade turned developer experience — docs quality, quickstarts, error messages — into a competitive surface, the agent era adds a second consumer of every SDK: a model that will follow the median of the public internet rather than your intended happy path. Teams that measure how agents actually use their platforms, and feed the evidence back into docs and tools, will be the ones agents build correctly with. The playbook is vendor-authored, but its core claim is hard to argue with: you already know what correct looks like for your technology. Now there is a method for finding out whether agents do too.
Comments (0)
Log in to join the discussion
Log InNo comments yet