Tech news in 3 minutes

Apple Mac Studio M5 Ultra and Mac Mini M6 Launch

8 d ago

Oxmiq Labs presented its High-Bandwidth Flash (HBF) specification for AI inference at Hot Chips 2026, positioning the technology as a capacity-tier alternative to HBM rather than a cheaper substitute. The company argues that HBF delivers 8 to 16 times the capacity of HBM at the same cost, targeting workloads where bandwidth demand is low and memory capacity is the bottleneck. The HBF hardware specification spans three grades, with maximum user bandwidth ranging from 0.384 TB/s to 3.072 TB/s and UCIe rates climbing from 8 GT/s to 32 GT/s. Capacity tops out at 512 GiB on a 16-high stack. Oxmiq frames HBF within a memory technology landscape using alpha and beta metrics, where beta tracks cost and alpha tracks bandwidth, emphasizing that HBF is a distinct capacity point, not a cheaper HBM variant. Oxmiq's analysis focuses on serving cost per token, breaking down the economics of hold and feed bandwidth. The company simulated a 72-GPU rack using a decode-centric Kimi-K2 1T model at FP4 with 1M tokens in and 1K out, finding that HBF buys roughly 14x the capacity for about 0.6x the bandwidth at the same rack cost. Software constraints include 64 KB chunk access for maximum bandwidth, approximately 24 hours of power-on data retention at 85 degrees Celsius, and host-managed lifecycle handling. Since HBF is read-optimized and write-constrained, placement becomes a software problem. Oxmiq proposes a vLLM plugin using HBF in place of host CPU pinned memory for KV cache and MoE expert pools, with a GPU configuration reaching 2.2 TB capacity at 17.4 TB/s peak bandwidth. The company concludes HBF wins only where bandwidth demand is low, such as MoE models with small batch sizes and long-context sparse KV scenarios, with attention sparsity models like DeepSeek Sparse Attention as genuine fits.

View original article

Timeline