← Все новости

not much happened today

**Nemotron-H** model family introduces hybrid Mamba-Transformer models with up to **3x faster inference** and variants including **8B**, **56B**, and a compressed **47B** model. **Nvidia Eagle 2.5** is a frontier VLM for long-context multimodal learning, matching **GPT-4o** and **Qwen2.5-VL-72B** on long-video understanding. **Gemini 2.5 Flash** shows improved dynamic thinking and cost-performance, outperforming previous Gemini versions. **Gemma 3** now supports **torch.compile** for about **60% faster inference** on consumer GPUs. **SRPO** using **Qwen2.5-32B** surpasses DeepSeek-R1-Zero-32B on benchmarks with reinforcement learning only. **Alibaba's Uni3C** unifies 3D-enhanced camera and human motion controls for video generation. **Seedream 3.0** by **ByteDance** is a bilingual image generation model with high-resolution outputs up to **2K**. **Adobe DRAGON** optimizes diffusion generative models with distributional rewards. **Kimina-Prover Preview** is an LLM trained with reinforcement learning from **Qwen2.5-72B**, achieving **80.7% pass@8192** on miniF2F. **BitNet b1.58 2B4T** is a native 1-bit LLM with **2B parameters** trained on **4 trillion tokens**, matching full-precision LLM performance with better efficiency. Antidistillation sampling counters unwanted model distillation by modifying reasoning traces from frontier models.
Читать оригинал на AINews / smol.ai →