HP ZBook Ultra G1a 14: 64GB Unified Memory Hits 8.6 t/s on 32B Local Models
AMD Ryzen AI Max PRO 390 with 64GB unified memory runs 32B models at 8.6 t/s — parity with a 16GB RTX 5080 laptop, no discrete GPU needed.
7 posts
AMD Ryzen AI Max PRO 390 with 64GB unified memory runs 32B models at 8.6 t/s — parity with a 16GB RTX 5080 laptop, no discrete GPU needed.
AMD's Threadripper Halo Station packs a 96-core CPU with dual MI350P accelerators and up to 576GB of HBM3E — a $150K desktop that runs trillion-parameter models locally.
vLLM 0.27 adds CondensePyramid V1 attention for multi-modal long context, new token-tier limits on templates, faster ASR CPU preprocessing via multi-threading, and CPU W4A16 INT4 MoE support.
Nvidia's Vera server CPU, 88 Olympus cores, targets AI inference. A credible alternative for token-heavy workloads, but software matters.
AMD's Zen 6 EPYC Venice debuts next week with up to 576 cores, 12 TB/s memory bandwidth, and native AVX-1024 support. Here's what changes for AI infrastructure.
NVIDIA's Vera CPU benchmarks show it competes directly with EPYC and Xeon, but software maturity and platform lock-in dictate adoption.
AMD's next-gen Ryzen AI Max chips bring desktop AI consolidation, massive unified memory, and high NPU TOPS to compact workstations starting this fall.