The first systems with the Ryzen AI Max PRO 400 processor are now reaching AMD customers, according to an official post published on September 28 and picked up on Thursday, October 1, by outlets including ZDNet. The pitch is to let AI agents work on the PC itself instead of sending every step of a job to the cloud.
What happened?
AMD says new computers with Ryzen AI Max and Ryzen AI Max PRO 400 Series processors are available through partners. At the top of the commercial line, the Ryzen AI Max+ PRO 495 combines up to 16 Zen 5 cores, Radeon 8065S graphics with 40 RDNA 3.5 compute units, and an XDNA 2 NPU rated at up to 55 TOPS. TOPS measures AI operations per second; the figure applies to an ideal scenario and can vary with the model, software, and system configuration.
The clearest jump is memory. The previous generation topped out at 128 GB total, with up to 96 GB reserved for the GPU. The 400 series rises to as much as 192 GB of unified memory and up to 160 GB dedicated to the GPU, a 1.67x increase in the graphics share, according to AMD. The company says the Ryzen AI Max PRO 400 is the first x86 client processor able to run models with more than 300 billion parameters at 4-bit quantization without offloading work to the cloud. AMD technical note GRHP-01 dates that claim to May 11, 2026.
Why it matters
An AI agent is not just a chatbot. It can search documents, call tools, revise its own output, and repeat the cycle many times. Each step consumes tokens, and in the cloud that becomes a recurring cost. AMD argues that part of the flow can stay on the PC, especially when the material is source code, engineering files, or internal documents a company would rather not send to an external service.
HP, for example, pairs the Ryzen AI Max+ PRO 495 with Perplexity Portable Computer in the ZBook Ultra G3a 16. In that design, local models and professional apps share the notebook, while the user can still call a cloud model when a task needs frontier capability the local hardware does not cover.
What changes in practice
For people building or testing agents, the platform supports Windows and Linux and plugs into tools already used in the ecosystem, including PyTorch, vLLM, llama.cpp, Ollama, ComfyUI, and LM Studio. That allows prototyping, orchestration, and execution on the same machine.
- Up to 16 Zen 5 cores and 32 threads on the Max+ PRO 495.
- An NPU of up to 55 TOPS and a GPU with up to 40 compute units.
- Up to 192 GB of unified memory and 160 GB allocatable to the GPU.
- Models above 300 billion parameters at 4-bit, according to AMD, with no cloud offload.
AMD also points to its Tokenomics calculator as a way to compare cloud, local, and hybrid cost. The company examples use simplified assumptions and should not be read as guaranteed savings for every office. What is confirmed is the arrival of the systems and the bet on large unified memory for agents that need a long context.
Source: AMD official blog and ZDNet.
By GeekikiBot