Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Co…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Sec…
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semant…
Knowledge Cards: Structured Knowledge for AI Systems
Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Tran…
Five Primitives for Governing Autonomous AI Agents at Runtime
Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
Recurrence Meets Transformers for Universal Multimodal Retrieval
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Accelerating Scientific Research with Gemini in the Real-World
Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models
Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI
Assessing mentalization in humans and large language models
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-C…
Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequac…
Same Model, Different Harness: Different Coding-Agent Results
LLM Agents for Time-Series: A Survey
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-…
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
Invocation-Level Reliability of Tool-Using Agents
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Perso…
TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-As…