Explorar

Noticias de IA

37834 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
arXiv cs.AI Research & Papers
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
arXiv cs.AI Research & Papers
From Fact Overwriting to Knowledge Evolution: Causal Editing via On-Policy Self-Distillat…
arXiv cs.AI Research & Papers
Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains
arXiv cs.AI Research & Papers
Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refi…
arXiv cs.AI Research & Papers
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
arXiv cs.AI Research & Papers
You Live More Than Once: Towards Hierarchical Skill Meta-Evolving
arXiv cs.AI Research & Papers
GONDOR to the Rescue: Satisficing Planning with Low Memory
arXiv cs.AI Research & Papers
GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven…
arXiv cs.AI Research & Papers
Cultural Binding Heads in Language Models
arXiv cs.AI Research & Papers
A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enha…
arXiv cs.AI Research & Papers
Auditable Decision Models with Learned Abstention and Real-Time Steering
arXiv cs.AI Research & Papers
Constrained Auto-Bidding via Generative Response Modeling
arXiv cs.AI Policy & Regulation
Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensi…
arXiv cs.AI Research & Papers
Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting
arXiv cs.AI Research & Papers
OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Int…
arXiv cs.AI Security & Safety
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
arXiv cs.AI Research & Papers
When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?
arXiv cs.AI Research & Papers
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
arXiv cs.AI Research & Papers
PetroBench: A Benchmark for Large Language Models in Petroleum Engineering
arXiv cs.AI Research & Papers
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
arXiv cs.AI Research & Papers
MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents
arXiv cs.AI Research & Papers
FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted …
arXiv cs.AI Research & Papers
Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Bette…
arXiv cs.AI Research & Papers
Dr-CiK: A Testbed for Foresight-Driven Agents
arXiv cs.AI Research & Papers
DiagramRAG: A Lightweight Framework to Retrieve Scientific Diagram for Figure Generation
arXiv cs.AI Research & Papers
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
arXiv cs.AI Research & Papers
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization …
arXiv cs.AI Research & Papers
On the Origin of Synthetic Information by Means of Steganographic Inheritance
arXiv cs.AI Research & Papers
DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM…