Browse

AI News

30934 items — filtered, classified, deduplicated

arXiv cs.AI Research & Papers
Fundamental Limitation in Explaining AI
arXiv cs.AI Research & Papers
When Mean CE Fails: Median CE Can Better Track Language Model Quality
arXiv cs.AI Research & Papers
Agent-Facing Information Design in LLM Tool Registries
arXiv cs.AI Research & Papers
Document Classification Pattern Recognition via Information Fusion: A Systematic Review o…
arXiv cs.AI Research & Papers
LETS Forecast: Learning Embedology for Time Series Forecasting
arXiv cs.AI Research & Papers
From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
arXiv cs.AI Research & Papers
Retrying vs Resampling in AI Control
arXiv cs.AI Research & Papers
Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qu…
arXiv cs.AI Research & Papers
Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with…
arXiv cs.AI Research & Papers
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
arXiv cs.AI Research & Papers
TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval
arXiv cs.AI Research & Papers
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
arXiv cs.AI Research & Papers
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
arXiv cs.AI Research & Papers
Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending
arXiv cs.AI Research & Papers
Certified Robustness from Approximate Gaussian Mixture Structures in Pretrained Latent Sp…
arXiv cs.AI Research & Papers
Generative structure search for efficient and diverse discovery of molecular and crystal …
arXiv cs.AI Research & Papers
Dynamic Dual-Granularity Skill Bank for Agentic RL
arXiv cs.AI Research & Papers
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
arXiv cs.AI Research & Papers
JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architec…
arXiv cs.AI Research & Papers
Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajecto…
arXiv cs.AI Research & Papers
AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent
arXiv cs.AI Research & Papers
CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
arXiv cs.AI Research & Papers
Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning
arXiv cs.AI Research & Papers
Positivity in classical enumerative geometry: a case study in synchronized AI-assisted ma…
arXiv cs.AI Research & Papers
Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC…
arXiv cs.AI Research & Papers
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
arXiv cs.AI Research & Papers
On the Epistemic Uncertainty of Overparametrized Neural Networks
arXiv cs.AI Research & Papers
By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They…
arXiv cs.AI Research & Papers
Hide to Guide: Learning via Semantic Masking
arXiv cs.AI Research & Papers
Remote sensing data imputation using deep learning for multispectral imagery