Xetrieval: Mechanistically Explaining Dense Retrieval
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions vi…
ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for …
When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role…
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression
Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomi…
Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Fra…
Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Gener…
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Ev…
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
Dataset-Driven Channel Masks in Transformers for Multivariate Time Series
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
Steering Language Models Before They Speak: Logit-Level Interventions
GPIC: A Giant Permissive Image Corpus for Visual Generation
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous S…
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Mode…
GroundAct: Can LLM Agents Ground Actions in Environmental States?
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-ground…
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Cryst…
Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teache…
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajecto…
GrepSeek: Training Search Agents for Direct Corpus Interaction
Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Models