AMD used an online press briefing on Monday to make its case for what it calls the "agentic PC" — a machine built not for chatbots but for autonomous AI agents that load models, call tools and hold long context windows entirely on-device. The timing was pointed: Nvidia and Microsoft host their RTX Spark launch event on October 7, and AMD spent the session two days earlier drawing a line between its approach and the incoming Arm-based competition.
The hardware centerpiece is the Ryzen AI Max PRO 400 series, led by the Ryzen AI Max+ PRO 495. Built on the Strix Halo design, it offers up to 192GB of unified memory — a 1.67x jump from the previous generation's 128GB — of which up to 160GB can be allocated to the GPU through AMD's variable graphics memory control. The chip pairs 16 Zen 5 CPU cores with RDNA 3.5 graphics, a 55 TOPS NPU and 40 GPU compute units. AMD claims it is the first x86 client processor able to run models with more than 300 billion parameters locally at 4-bit quantization without offloading to the cloud — a vendor claim made under its own specified conditions, not an independent benchmark.
Why memory rather than raw compute? Because agents, in AMD's telling, are memory-hungry in a way chatbots are not: a single agent may hold a base model, a retrieval model and a tool-calling model at once while keeping tens of thousands of tokens of context live. Run out of capacity and the system starts swapping models in and out, and the experience collapses. AMD's briefing included local throughput figures for two open models — GLM-5.3-Flash at 20 tokens per second and Qwen3.8-Flash-Next at 42 tokens per second — with the caveat that multi-agent workloads stress prefill and sustained throughput far more than single-model peak speed.
The company framed its lineup in three tiers: consumer and commercial notebooks for everyday AI, Threadripper workstations and Radeon AI PRO hardware for heavy professional compute, and the agentic PC in between for developers and AI innovators. The platform runs Windows and Linux natively — a jab at Nvidia's Arm-based RTX Spark, which relies on Microsoft's Prism emulation layer on Windows — and supports PyTorch, vLLM, llama.cpp, Ollama, ComfyUI and LM Studio out of the box. AMD's marketing has taken to expanding its own acronym as "Agentic Micro Devices since 2025."
The traction numbers, all company-reported, were the session's other headline: AMD says it has shipped tens of millions of AI PCs across more than 250 designs, and that more than 500,000 machines capable of running 100-billion-parameter-class models locally have shipped across over 50 designs. One customer example carried the economics: Emmy-winning VR studio LightSail VR, which produces 16K 90fps stereo 3D, built a "production coordinator" agent on a Ryzen AI Max desktop that manages six to nine projects at once — work it says used to take two people — with zero cloud spend, versus thousands of dollars per month in tokens for the equivalent cloud workload.
AMD is openly positioning local compute as a hedge against agentic AI's token bill rather than a replacement for the cloud, pointing buyers to its Tokenomics Calculator to model hybrid deployments. The bet is that when enterprises see agents consuming tokens by the millions on repeated reasoning and tool calls, a 192GB desk-side machine starts looking like infrastructure rather than a PC. Nvidia will present the opposite pole on Wednesday — maximum accelerated compute with unified memory as the enabler. The interesting question for buyers is not which chip wins, but where enterprises actually draw the line between data that never leaves the building and tokens that are cheaper in someone else's data center.
Comments (0)
Log in to join the discussion
Log InNo comments yet