Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive …
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Servi…
Reducing Hallucinations in LLM-based Scientific Literature Analysis Using Peer Context Ou…
When Can One Obtain Certificates of Optimality Using Positivstellensaetze?
Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction
What Does an LLM-Agent Leaderboard Rank Actually Compare?
Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliabi…
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
A radiographic world model for clinical reasoning and evidence generation
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems
Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors
SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis …
Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware …
Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier
Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Froze…
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persi…
Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supe…
Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection
Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models
PAGR: Proof-Carrying Algebraic-Geometric Retrieval: A Quiver-, Provenance-, and Sheaf-The…
RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavior…
CIT-CAD: Constraint Intent Tree-based CAD Code Generation and Verification
The Internal Anatomy of Strategic Choice in Large Language Models
Human-like moral judgments conceal divergent motive attributions in large language models
ViT3Flow: A Test-Time Training Transformer MeanFlow for Postoperative Radiograph Synthesi…
Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement …
Qiushi Engine on AstaBench E2E-Bench-Hard
SkillAlign: Aligning Skill Interfaces for LLM-based Agents
Weakly supervised neural network: segmentation of complex structures in X-ray microCT