Tech news in 3 minutes
IBC 2026: Tuxera and ATTO Collaborate to Deliver Blistering macOS Network Performance
Lightbits Labs has unveiled Inferra, an intelligent KV cache orchestration engine that dramatically improves GPU utilization for AI inference, making its public debut at the AI Infra Summit on September 15th. Designed to tackle memory bottlenecks from exploding KV caches in long-context LLMs, Inferra enables neoclouds and enterprises to run inference faster, scale bigger, and reduce AI infrastructure costs. The AI inference market is projected to top $117 billion this year, and Lightbits co-founder Avigdor Willenz emphasized that retrofitting legacy training systems for inference doesn't work, as Inferra was engineered from the ground up to solve GPU efficiency. Key breakthroughs include infinite HBM virtualization with 16x session density, serving up to 16x more concurrent inference sessions without hardware upgrades; over 100x latency improvement with 10M-token context windows by pre-fetching attention states from storage; and AI-native security with tenant isolation. Lightbits, inventors of NVMe over TCP, virtualizes GPU memory across memory and storage tiers, transforming the KV cache into a persistent data layer. OVHcloud CPTO Yaniv Fdida noted substantial GPU utilization gains from Inferra’s intelligent KV cache tiering, enabling scalable, cost-effective AI agent and RAG workloads. Solidigm’s Avi Shetty highlighted how network-attached storage with Inferra’s software virtualization breaks the memory wall for agentic AI, while ICC CTO Alexey Stolyar sees Inferra as an innovative solution for inference and token efficiency. Lightbits SVP Ramesh Chettuvetty stated Inferra delivers instant payback and net positive savings from day one, eliminating GPU stalls for larger models and longer conversations at lower cost. AI infrastructure teams, hyperscalers, and GPU cloud providers can visit booth 219 at the AI Infra Summit for live demonstrations.