Titans: Learning to Memorize at Test Time
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
Visual Document Retrieval Goes Multilingual
not much happened today
CO₂ Emissions and Models Performance: Insights from the Open LLM Leaderboard
PRIME: Process Reinforcement through Implicit Rewards
not much happened to end the year
Deliberative alignment: reasoning enables safer language models
Evaluating Audio Reasoning with Big Bench Audio
Genesis: Generative Physics Engine for Robotics (o1-2024-12-17)
Benchmarking Language Model Performance on 5th Gen Xeon at GCP
LeMaterial: an open source initiative to accelerate materials discovery and research
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
How good are LLMs at fixing their mistakes? A chatbot arena experiment with Keras and TPUs
Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard
Investing in Performance: Fine-tune small models with LLM insights - a CFM case study
You could have designed state of the art positional encoding
Letting Large Models Debate: The First Multilingual LLM Debate Competition
Faster Text Generation with Self-Speculative Decoding
Judge Arena: Benchmarking LLMs as Evaluators
Common Corpus: 2T Open Tokens with Provenance
BitNet was a lie?
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Creating a LLM-as-a-Judge
Introducing SimpleQA
Universal Assisted Generation: Faster Decoding with Any Assistant Model
Expert Support case study: Bolstering a RAG app with LLM-as-a-Judge
s{imple|table|calable} Consistency Models
A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality
Simplifying, stabilizing, and scaling continuous-time consistency models
CinePile 2.0 - making stronger datasets with adversarial refinement