← Все новости

not much happened today

**NVIDIA’s Nemotron 3 Super** is a **120B parameter / ~12B active** open model featuring a **hybrid Mamba-Transformer / SSM Latent MoE** architecture and **1M context window**, delivering up to **2.2x faster inference than GPT-OSS-120B** in FP4 with strong throughput gains. It supports agentic workloads and is unusually open with weights, data, and infrastructure details released. The model scored **36 on the AA Intelligence Index**, outperforming GPT-OSS-120B but behind Qwen3.5-122B-A10B. Community and infrastructure support from projects like **vLLM**, **llama.cpp**, **Ollama**, **Together**, **Baseten**, **W&B Inference**, **LangChain**, and **Unsloth GGUFs** was immediate. Key technical innovations include **native multi-token prediction (MTP)** and a significant **KV-cache efficiency** advantage. On the product side, a shift towards **persistent agent runtimes and orchestration layers** is highlighted, with **Andrej Karpathy** advocating for a "bigger IDE" concept where agents replace files as the unit of work, enabling legible, forkable agentic organizations with real-time control. New launches fitting this vision include **Perplexity’s Personal Computer**, an always-on local/cloud hybrid running on Mac mini, and **Computer for Enterprise** orchestrating 20 specialized models and 400+ apps. **Replit Agent 4** offers a collaborative, canvas-like workflow with parallel agents, while **Base44 Superagents** provide integrated solutions for nontechnical users. The engineering focus is increasingly on the orchestration harness rather than just the model.
Читать оригинал на AINews / smol.ai →