Hy4 preview
Product Hunt · Aug 29, 2026
Qwen3.8-Flash-Next
Simon Willison · Aug 26, 2026
Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency"
The Decoder · Aug 26, 2026
Granite 4.2 LLMs: How They're Built
Hugging Face · Aug 25, 2026
Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated
The Decoder · Aug 25, 2026
Alibaba, DeepSeek push China’s AI model race towards lower costs
AI News · Aug 5, 2026
Inside the Model Factory — Eiso Kant, Poolside AI
Latent Space · Jul 23, 2026
Latent-MoE — Implementation of LatentMoE,Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts (Elango et al., NVIDIA
GitHub · Jul 21, 2026