Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
From Fact Overwriting to Knowledge Evolution: Causal Editing via On-Policy Self-Distillat…
Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains
Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refi…
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
You Live More Than Once: Towards Hierarchical Skill Meta-Evolving
GONDOR to the Rescue: Satisficing Planning with Low Memory
GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven…
Cultural Binding Heads in Language Models
A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enha…
Auditable Decision Models with Learned Abstention and Real-Time Steering
Constrained Auto-Bidding via Generative Response Modeling
Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensi…
Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting
OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Int…
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
PetroBench: A Benchmark for Large Language Models in Petroleum Engineering
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents
FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted …
Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Bette…
Dr-CiK: A Testbed for Foresight-Driven Agents
DiagramRAG: A Lightweight Framework to Retrieve Scientific Diagram for Figure Generation
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization …
On the Origin of Synthetic Information by Means of Steganographic Inheritance
DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM…