Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generat…
Communication styles and reader preferences of LLM- and human-authored COVID-19 informati…
Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinfor…
Temporally-Grounded Language Generation: Towards Real-Time Vision-Language Models
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Rec…
A Very Big Video Reasoning Suite
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Conte…
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
How LLMs Distort Our Written Language
MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and a…
MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Predic…
MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on …
Pixel Wised Lesion Prediction on COVID-19 CT Imagery: A Comparative Analysis of Automated…
Unsupervised Post-Training of Foundation Models: A Survey
GameWAM: A World Action Model for Video Games
A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeti…
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Con…
The BS-meter: Detecting Politics and Labour through ChatGPT's Language
AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air
What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Un…
LoRA as Oracle
When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Me…
ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One…
Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from …