How NuminaMath Won the 1st AIMO Progress Prize
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
Test-Time Training, MobileLLM, Lilian Weng on Hallucination (Plus: Turbopuffer)
Preference Optimization for Vision Language Models
Problems with MMLU-Pro
Qdrant's BM42: "Please don't trust us"
GraphRAG: The Marriage of Knowledge Graphs and RAG
Accelerating Protein Language Model ProtST on Intel Gaudi 2
Finding GPT-4’s mistakes with GPT-4
Shazeer et al (2024): you are overpaying for inference >13x
Consistency Models
Data Is Better Together: A Look Back and Forward
Improved Techniques for Training Consistency Models
A Holistic Approach to Undesired Content Detection in the Real World
Is this... OpenQ*?
BigCodeBench: The Next Generation of HumanEval
Hybrid SSM/Transformers > Pure SSMs/Pure Transformers
Putting RL back in RLHF
Francois Chollet launches $1m ARC Prize
Extracting Concepts from GPT-4
Contextual Position Encoding (CoPE)
Somebody give Andrej some H100s already
Clémentine Fourrier on LLM evals
Anthropic's "LLM Genome Project": learning & clamping 34m features on Claude Sonnet
Unlocking Longer Generation with Key-Value Cache Quantization
Introducing the Open Arabic LLM Leaderboard
LMSys advances Llama 3 eval analysis
Kolmogorov-Arnold Networks: MLP killers or just spicy MLPs?
$100k to predict LMSYS human preferences in a Kaggle contest
Evals: The Next Generation
OpenAI's Instruction Hierarchy for the LLM OS