PaperBench: Evaluating AI’s Ability to Replicate AI Research
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
First Look at Reasoning From Scratch: Chapter 1
Open R1: Update #4
lots of little things happened this week
Early methods for studying affective use and emotional well-being on ChatGPT
Every 7 Months: The Moore's Law for Agent Autonomy
Innovation to Impact: How NVIDIA Research Fuels Transformative Work in AI, Graphics and B…
LeRobot goes to driving school: World’s largest open-source self-driving dataset
The State of LLM Reasoning Model Inference
A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality
OpenAI GPT-4.5 System Card
The Ultra-Scale Playbook: Training LLMs on GPU Clusters
Introducing the SWE-Lancer benchmark
LLaDA: Large Language Diffusion Models
Reasoning Models are Near-Superhuman Coders (OpenAI IOI, Nvidia Kernels)
What Are Foundation Models?
The Open Arabic LLM Leaderboard 2
AI-Designed Proteins Take on Deadly Snake Venom
s1: Simple test-time scaling (and Kyutai Hibiki)
How To Scale Your Model, by DeepMind
DABStep: Data Agent Benchmark for Multi-step Reasoning
π0 and π0-FAST: Vision-Language-Action Models for General Robot Control
What Is Retrieval-Augmented Generation, aka RAG?
Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial
Open-R1: a fully open reproduction of DeepSeek-R1
State of open video generation models in Diffusers
TinyZero: Reproduce DeepSeek R1-Zero for $30
AI Maps Titan’s Methane Clouds in Record Time
Mastering Long Contexts in LLMs with KVPress
Trading inference-time compute for adversarial robustness