Cognition vs Anthropic: Don't Build Multi-Agents/How to Build Multi-Agents
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
KV Cache from scratch in nanoVLM
CodeAgents + Structure: A Better Way to Execute Actions
🐯 Liger GRPO meets TRL
Gemini's AlphaEvolve agent uses Gemini 2.0 to find new Math and cuts Gemini cost 1% — wit…
Introducing HealthBench
LeRobot Community Datasets: The “ImageNet” of Robotics — When and How?
Coding LLMs from the Ground Up: A Complete Course
Expanding on what we missed with sycophancy
ChatGPT responds to GlazeGate + LMArena responds to Cohere
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
The State of Reinforcement Learning for LLM Reasoning
OpenAI o3 and o4-mini System Card
Thinking with images
Introducing HELMET: Holistically Evaluating Long-context Language Models
not much happened today
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
BrowseComp: a benchmark for browsing agents
Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
PaperBench: Evaluating AI’s Ability to Replicate AI Research
First Look at Reasoning From Scratch: Chapter 1
Open R1: Update #4
lots of little things happened this week
Early methods for studying affective use and emotional well-being on ChatGPT
Every 7 Months: The Moore's Law for Agent Autonomy
Innovation to Impact: How NVIDIA Research Fuels Transformative Work in AI, Graphics and B…
LeRobot goes to driving school: World’s largest open-source self-driving dataset
The State of LLM Reasoning Model Inference
A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality
OpenAI GPT-4.5 System Card