GitHub has extended Copilot beyond the editor and into the whole screen. Computer use is now in public preview for the Copilot CLI and the Copilot app on macOS and Windows, letting the agent operate desktop applications on a developer's behalf: reading accessible app content and visual context, clicking controls, typing and editing text, pressing keys, scrolling and dragging - including across multiple applications in one workflow.
The target is a category of software that agents have largely been unable to touch: legacy and GUI-only applications that offer no API, no command line and no MCP integration, the kind of tools that until now could only be operated by a human at the keyboard. GitHub's suggested uses include summarizing a pile of browser notifications, updating the contents of a presentation deck, and shuttling data between desktop apps; GitHub's Pierce Boggan demonstrated expense workflows, travel booking and end-to-end testing on the feature.
The control model is deliberately conservative. Copilot must be granted access before it takes over any given application, and the approval is per-app; users who mark an application "always allowed" can go back later and review or reset that list. On macOS, enabling the feature walks users through granting two system-level permissions - Accessibility and Screen Recording - and the capability can be switched on either with the /computer on command in the CLI or from Settings in the desktop app. Enterprises get their own gate: organization-managed settings can disable computer use outright.
GitHub's guidance on prompting is itself a signal of where the feature stands. The company recommends specifying the desired outcome, the applications involved and any constraints - in other words, the system performs best when a human has already decomposed the task, and a vague instruction is likely to produce unpredictable clicking rather than reliable automation. That is an honest framing of current agentic computer use: strong within a well-scoped envelope, brittle outside it.
The strategic read is bigger than one preview. GitHub has spent the past year turning Copilot from an autocomplete engine into an agent platform - coding agents, managed runtimes, dynamic workflows - and screen control is the missing primitive that lets it operate in the enterprise environments where most business software still lives in thick clients. It also aligns GitHub with the broader industry trajectory: Meta's Muse, OpenAI's Dots and Grok's Bot all pitch persistent agents that act on real software, and the regulatory temperature around exactly this capability is rising, from California's new No Robo Bosses Act to the FTC's first probe into rogue agent behavior.
For now the preview is a developer tool with guardrails, not an unsupervised operator. The interesting question is how quickly the approval model - human confirms each app, enterprise can pull the switch - becomes the template other agent vendors are measured against, or gets eroded by convenience as the capability matures.
Comments (0)
Log in to join the discussion
Log InNo comments yet