In-Context Examples Suppress Scientific Knowledge Recall in LLMs
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Lea…
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pre…
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inferen…
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
MemWM: Memory-Augmented Text-Based World Model
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Pr…
How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid …
FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabul…
Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems
SkillEval: Decomposing Agent Skill Quality into Interpretable Signals
GeoUniPR: A Geometry-Consistent Unified Framework for Cross-Modal Place Recognition
SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision
Evo-Bench: Can Language Models Improve Agent Harness?
Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reaso…
Security and Privacy Taxonomy Generation from Mobile App Reviews
Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Ba…
Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning
A Tight Lower Bound for Smooth Nonconvex Stochastic Optimization with Bounded Gradient No…
DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Si…
Automating and Scaling Behavioral Scientific Research on AI Agents
SoftMCC: An MCC-Brier Calibration Bridge for Threshold-Free Model Selection under Class I…
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based …
Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient…