Introducing EVMbench
Explorar
Noticias de IA
21861 elementos — filtrados, clasificados y sin duplicados
Import AI 445: Timing superintelligence; AIs solve frontier math proofs; a new ML researc…
GPT-5.2 derives a new result in theoretical physics
OpenEnv in Practice: Evaluating Tool-Using Agents in Real-World Environments
Accelerating Mathematical and Scientific Discovery with Gemini Deep Think
Import AI 444: LLM societies; Huawei makes kernels with AI; ChipBench
GPT-5 lowers the cost of cell-free protein synthesis
Training Design for Text-to-Image Models: Lessons from Ablations
Import AI 443: Into the mist: Moltbook, agent ecologies, and the internet in transition
MoltBook takes over the timeline
Project Genie: Experimenting with infinite, interactive worlds
not much happened today
We Got Claude to Build CUDA Kernels and teach open models!
Architectural Choices in China's Open-Source AI Ecosystem: Building Beyond DeepSeek
Alyah ⭐️: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs
Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective
Categories of Inference-Time Scaling for Improved LLM Reasoning
AssetOpsBench: Bridging the Gap Between AI Agent Benchmarks and Industrial Reality
not much happened today
not much happened today
D4RT: Teaching AI to see the world in four dimensions
Import AI 440: Red queen AI; AI regulating AI; o-ring automation
not much happened today
Import AI 439: AI kernels; decentralized training; and universal representations
not much happened today
LLM Research Papers: The 2025 List (July to December)
Google's year in review: 8 areas with research breakthroughs in 2025
Evaluating chain-of-thought monitorability
The Open Evaluation Standard: Benchmarking NVIDIA Nemotron 3 Nano with NeMo Evaluator
Evaluating AI’s ability to perform scientific research tasks