← Todas las noticias

not much happened today

**Z.ai** released the **GLM-5.3** open-weight model family, optimized for **agentic coding** and **cyber defense**, with impressive specs like **744B total / 40B active parameters**, **1M context window**, and a **239GB 2-bit** variant retaining **81% accuracy**. **Tencent** launched **Hy4-preview**, a top-tier open-source MoE model with **770B total / 49B active parameters** and **1M context**, showing strong benchmark performance and innovative serving design. **Alibaba** introduced **Qwen3.8-Flash**, a cheaper, long-context MoE with **125B total / 6B active parameters** and multimodality, though early user reports noted some stability issues resolved by switching KV cache to **BF16**. On the systems side, **vLLM** published a detailed speculative decoding benchmark across multiple models and hardware, emphasizing no one-size-fits-all solution. Additionally, search systems like **Perplexity Search** are gaining prominence as evaluated subsystems with strong economic and performance metrics. *"There is no universal winner"* in speculative decoding, highlighting the need for workload-specific tuning.
Leer el original en AINews / smol.ai →