Tech news in 3 minutes
Hot Chips 2026: Intel Outlines Architectures for Agentic AI
Nvidia Groq 3 LPX, the interactive AI inference accelerator now in full production, delivers world-record token generation speeds for agentic coding and latency-sensitive workloads, with Nebius becoming the first AI cloud to adopt the platform for its Token Factory inference service. An extension of the Nvidia Vera Rubin platform, Groq 3 LPX dramatically increases token generation rates for Vera Rubin NVL72 systems, enabling ultrafast responsiveness for agentic AI that requires massive token volumes across hundreds of inference steps. In Artificial Analysis benchmarking, Groq 3 LPX achieved a record 3,400 output tokens per second running Gemma 4 31B with a 100,000-token context—the fastest performance ever recorded for that model. This enables agentic tasks like coding to complete in minutes versus hours, providing 4x faster responsiveness than the nearest alternative platform. “Inference is the growth engine of AI,” said Jensen Huang, Nvidia CEO, noting that Vera Rubin extends the vision with workload-optimized AI factory configurations for the agentic AI era. Nebius plans to integrate Groq 3 LPX into its production inference platform, giving developers instant token generation through existing APIs without stack migration. Purpose-built AI inference cloud Groq also plans to be among the platform’s earliest adopters. The Vera Rubin platform, codesigned across seven chips and five racks, includes BlueField-4 DPUs, Vera CPU racks, Spectrum-6 SPX Ethernet, and STX storage to optimize multi-agent systems for highest throughput per watt and lowest-latency inference.