Explorar

Noticias de IA

21271 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
arXiv cs.AI Research & Papers
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality…
arXiv cs.AI Research & Papers
On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach
arXiv cs.AI Research & Papers
CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intellige…
arXiv cs.AI Research & Papers
Beyond Questions: Evaluating What Large Language Models (Actually) Know
arXiv cs.AI Research & Papers
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
arXiv cs.AI Research & Papers
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
arXiv cs.AI Research & Papers
Yes, Q-learning Helps Offline In-Context RL
arXiv cs.AI Research & Papers
Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection
arXiv cs.AI Research & Papers
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Dat…
arXiv cs.AI Research & Papers
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
arXiv cs.AI Research & Papers
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
arXiv cs.AI Research & Papers
Mechanized Foundations of Structural Governance: Machine-Checked Proofs for Governed Inte…
arXiv cs.AI Research & Papers
Algebraic Semantics of Governed Execution: Monoidal Categories, Effect Algebras, and Cote…
arXiv cs.AI Research & Papers
Many Logics, One Methodology: A Plea for Logical Pluralism in Formalised Reasoning (prepr…
arXiv cs.AI Research & Papers
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
arXiv cs.AI Research & Papers
Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models
arXiv cs.AI Research & Papers
LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop P…
arXiv cs.AI Research & Papers
How Chain-of-Thought Works? Tracing Information Flow from Decoding, Projection, and Activ…
arXiv cs.AI Research & Papers
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
arXiv cs.AI Research & Papers
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting …
arXiv cs.AI Research & Papers
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
arXiv cs.AI Research & Papers
Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V
arXiv cs.AI Research & Papers
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Des…
arXiv cs.AI Research & Papers
MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering via Miscoverage Ri…
arXiv cs.AI Research & Papers
Grounding Text Embeddings in Stakeholder Associations
arXiv cs.AI Research & Papers
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical No…
arXiv cs.AI Research & Papers
From Feasible to Practical: Pareto-Optimal Synthesis Planning
arXiv cs.AI Research & Papers
GraphMind: From Operational Traces to Self-Evolving Workflow Automation
arXiv cs.AI Research & Papers
Declarative Data Services: Structured Agentic Discovery for Composing Data Systems