Why we no longer evaluate SWE-bench Verified
Explorar
Noticias de IA
29629 elementos — filtrados, clasificados y sin duplicados
OpenAI announces Frontier Alliance Partners
not much happened today
Our First Proof submissions
GGML and llama.cpp join HF to ensure the long-term progress of Local AI
Train AI models with Unsloth and Hugging Face Jobs for FREE
Gemini 3.1 Pro: A smarter model for your most complex tasks
Advancing independent research on AI alignment
Gemini 3.1 Pro: 2x 3.0 on ARC-AGI 2
Introducing OpenAI for India
IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST
A new way to express yourself: Gemini can now create music
Evals: Your Bridge From AI Experimentation To Confident Production Deployments
not much happened today
One-Shot Any Web App with Gradio's gr.HTML
Introducing EVMbench
Accelerating discovery in India through AI-powered science and education
Claude Sonnet 4.6: clean upgrade of 4.5, mostly better with some caveats
LWiAI Podcast #234 - Opus 4.6, GPT-5.3-Codex, Seedance 2.0, GLM-5
Import AI 445: Timing superintelligence; AIs solve frontier math proofs; a new ML researc…
Qwen3.5-397B-A17B: the smallest Open-Opus class, very efficient model
Last Week in AI #335 - Opus 4.6, Codex 5.3, Gemini 3 Deep Think, GLM 5, Seedance 2.0
GPT-5.2 derives a new result in theoretical physics
Introducing Lockdown Mode and Elevated Risk labels in ChatGPT
Beyond rate limits: scaling access to Codex and Sora
Scaling social science research
MiniMax-M2.5: SOTA coding, search, toolcalls, $1/hour
Custom Kernels for All from Codex and Claude
Gemini 3 Deep Think: Advancing science, research and engineering
Introducing GPT-5.3-Codex-Spark