Dr-CiK: A Testbed for Foresight-Driven Agents
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Bette…
FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted …
MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
PetroBench: A Benchmark for Large Language Models in Petroleum Engineering
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit As…
Show, Don't TELL: Explainable AI-Generated Text Detection
When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?
OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Int…
Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting
Constrained Auto-Bidding via Generative Response Modeling
Auditable Decision Models with Learned Abstention and Real-Time Steering
A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enha…
Cultural Binding Heads in Language Models
GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven…
GONDOR to the Rescue: Satisficing Planning with Low Memory
You Live More Than Once: Towards Hierarchical Skill Meta-Evolving
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refi…
Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains
From Fact Overwriting to Knowledge Evolution: Causal Editing via On-Policy Self-Distillat…
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential …
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scena…
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback