← Все новости

not much happened today

**Inference optimization** is increasingly architectural, with **EAGLE 3.1** improving speculative decoding and long-context handling, collaborating with **vLLM** and **TorchSpec**. **Perplexity** open-sourced a rebuilt **Unigram tokenizer** cutting CPU use by **5–6×** and achieving **63 µs at 514 tokens**. **Qwen3.5** hits **580 tokens/s** via joint efforts from **Alibaba**, **LightSeek**, **NVIDIA**, **Mooncake**, and **FlashAttention-4** contributors. Price cuts in APIs from Chinese labs are sustainable due to structural KV-cache and attention improvements, exemplified by **DeepSeek V4-Pro** and **Xiaomi MiMo** reducing caching costs significantly. Agent engineering shifts focus from model quality to model-harness-memory fit, with **LangChain** releasing **Deep Agents v0.6** and tools like **LangSmith Engine** automating evaluation loops. **Trajectory** launched a continual learning platform with **$15M funding** and partners like **Clay** and **Harvey**, supporting large models including a **397B-parameter model** deployed on autoscaled **H100** infrastructure. Open-source memory-centric agents and minimal training harnesses also gained attention.
Читать оригинал на AINews / smol.ai →