Mixture of Depths: Dynamically allocating compute in transformer-based language models
Explorar
Noticias de IA
21813 elementos — filtrados, clasificados y sin duplicados
ReALM: Reference Resolution As Language Modeling
AdamW -> AaronD?
Evals-based AI Engineering
Andrew likes Agents
Pollen-Vision: Unified interface for Zero-Shot vision models in robotics
not much happened today
Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval
Welcome /r/LocalLlama!
Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Mode…
GaLore: Advancing Large Model Training on Consumer-grade Hardware
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
DeepMind SIMA: one AI, 9 games, 600 tasks, vision+language ONLY
Introducing ConTextual: How well can your Multimodal model jointly reason over text and i…
The Era of 1-bit LLMs
TTS Arena: Benchmarking Text-to-Speech Models in the Wild
Ring Attention for >1M Context
🪆 Introduction to Matryoshka Embedding Models
Karpathy emerges from stealth?
Introducing the Open Ko-LLM Leaderboard: Leading the Korean LLM Evaluation Ecosystem
Synthetic data: save money, time and carbon with open source
Video generation models as world simulators
AI gets Memory
The Core Skills of AI Engineering
NPHardEval Leaderboard: Unveiling the Reasoning Abilities of Large Language Models throug…
Constitutional AI with Open LLMs
The Hallucinations Leaderboard, an Open Effort to Measure Hallucinations in Large Languag…
RIP Latent Diffusion, Hello Hourglass Diffusion
Preference Tuning LLMs with Direct Preference Optimization Methods
1/16/2024: TIES-Merging