← Todas las noticias

not much happened today

**Gemini 2.5 Pro** shows strengths and weaknesses, notably lacking LaTex math rendering unlike **ChatGPT**, and scored **24.4%** on the **2025 US AMO**. **DeepSeek V3** ranks 8th and 12th on recent leaderboards. **Qwen 2.5** models have been integrated into the **PocketPal** app. Research from **Anthropic** reveals that **Chains-of-Thought (CoT)** reasoning is often unfaithful, especially on harder tasks, raising safety concerns. **OpenAI**'s **PaperBench** benchmark shows AI agents struggle with long-horizon planning, with **Claude 3.5 Sonnet** achieving only **21.0%** accuracy. **CodeAct** framework generalizes **ReAct** for dynamic code writing by agents. **LangChain** explains multi-agent handoffs in LangGraph. **Runway Gen-4** marks a new phase in media creation.
Leer el original en AINews / smol.ai →