Tag: llm-inference
All the articles with the tag "llm-inference".
- AI Signals
Kimi K3 Is Here: What the 2.8T Open Model Actually Changes
Moonshot's Kimi K3 combines 2.8T parameters, sparse MoE routing, native vision, and a 1M context window. Agent reliability and serving cost remain open.
- AI Engineering
Ray 2.55 Makes Google TPUs a First-Class Cluster Resource
Ray 2.55 adds official TPU support across its release pipeline. The key engineering change is atomic scheduling of complete TPU slices on GKE.
- AI Engineering
Google Made Qwen 397B Up to 4.7x Faster Without Training a New Model
Google's Ironwood optimization work shows how sharding, fused kernels, and memory-aware serving can change the economics of a frontier-scale open-weight model.