DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
Evaluating fairness in ChatGPT
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Faster Assisted Generation with Dynamic Speculation
Introducing the Open FinLLM Leaderboard
🇨🇿 BenCzechMark - Can your LLM Understand Czech?
not much happened today
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
Learning to reason with LLMs
Answering quantum physics questions with OpenAI o1
Decoding genetics with OpenAI o1
Economics and reasoning with OpenAI o1
not much happened this weekend
Scaling robotics datasets with video encoding
Nvidia Minitron: LLM Pruning and Distillation updated for Llama 3.1
A failed experiment: Infini-Attention, and why we should keep trying?
Introducing SWE-bench Verified
Tool Use, Unified
GPT-4o System Card
Introducing TextImage Augmentation for Document Images
Memory-efficient Diffusion Transformers with Quanto and Diffusers
AlphaProof + AlphaGeometry2 reach 1 point short of IMO Gold
LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?
Improving Model Safety Behavior with Rule-Based Rewards
Docmatix - a huge dataset for Document Visual Question Answering
Prover-Verifier Games improve legibility of language model outputs
SciCode: HumanEval gets a STEM PhD upgrade
Microsoft AgentInstruct + Orca 3
We Solved Hallucinations
FlashAttention 3, PaliGemma, OpenAI's 5 Levels to Superintelligence