Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Browse
AI News
18308 items — filtered, classified, deduplicated
Bayesian and Motivated Reasoning in AI Agents
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their I…
Neuro-Evolved Heuristics for Variable Gapped Common Subsequence Identification
SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented a…
Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LL…
TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LL…
F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Lang…
Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucina…
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
Modeling Social Dynamics with an LLM-Enabled Agent Based Network-Dynamic (LAND) Model
More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness
Memory Reward Inflation in Self-Improving LLM Agents
RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learn…
Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale
RF-HOI: Recognize Human-Object Interaction with Radio Frequency Signals
Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization
Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates
Motif-Mamba: network motif improved mamba for long-range sequence modeling
TRACE-TS: Attribution-Grounded and Traceable Sensor-Language Reasoning for Human Activity…
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Mode…
HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive …
The Bayesian Reflex: A Predictive Coding Engine for Artificial Intelligence
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimizat…
Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmar…
AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
Assuming You Knew: Fixing an Epistemic Semantics for Flow Policies Using Agentic AI