Common Corpus: 2T Open Tokens with Provenance
Explorar
Noticias de IA
21861 elementos — filtrados, clasificados y sin duplicados
BitNet was a lie?
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Creating a LLM-as-a-Judge
Introducing SimpleQA
Universal Assisted Generation: Faster Decoding with Any Assistant Model
Expert Support case study: Bolstering a RAG app with LLM-as-a-Judge
s{imple|table|calable} Consistency Models
A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality
Simplifying, stabilizing, and scaling continuous-time consistency models
CinePile 2.0 - making stronger datasets with adversarial refinement
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
Evaluating fairness in ChatGPT
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Faster Assisted Generation with Dynamic Speculation
Introducing the Open FinLLM Leaderboard
🇨🇿 BenCzechMark - Can your LLM Understand Czech?
not much happened today
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
Learning to reason with LLMs
Answering quantum physics questions with OpenAI o1
Decoding genetics with OpenAI o1
Economics and reasoning with OpenAI o1
not much happened this weekend
Scaling robotics datasets with video encoding
Nvidia Minitron: LLM Pruning and Distillation updated for Llama 3.1
A failed experiment: Infini-Attention, and why we should keep trying?
Introducing SWE-bench Verified
Tool Use, Unified
GPT-4o System Card