Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
How good are LLMs at fixing their mistakes? A chatbot arena experiment with Keras and TPUs
Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard
Investing in Performance: Fine-tune small models with LLM insights - a CFM case study
You could have designed state of the art positional encoding
Faster Text Generation with Self-Speculative Decoding
Letting Large Models Debate: The First Multilingual LLM Debate Competition
Judge Arena: Benchmarking LLMs as Evaluators
Common Corpus: 2T Open Tokens with Provenance
BitNet was a lie?
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Creating a LLM-as-a-Judge
Introducing SimpleQA
Universal Assisted Generation: Faster Decoding with Any Assistant Model
Expert Support case study: Bolstering a RAG app with LLM-as-a-Judge
s{imple|table|calable} Consistency Models
A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality
Simplifying, stabilizing, and scaling continuous-time consistency models
CinePile 2.0 - making stronger datasets with adversarial refinement
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
Evaluating fairness in ChatGPT
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Faster Assisted Generation with Dynamic Speculation
Introducing the Open FinLLM Leaderboard
🇨🇿 BenCzechMark - Can your LLM Understand Czech?
not much happened today
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
Learning to reason with LLMs
Economics and reasoning with OpenAI o1
Decoding genetics with OpenAI o1