From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI …
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate
Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study
ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Dele…
WorldAgen: Unified State-Action Prediction with Test-Time World Model Training
From LLM-Generated Specifications to Learned Quadruped Locomotion
Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval
Topology Obstructs Pure Foundation Neural Quantum States
CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows
Support Topology and Gradient Mixing in Sinkhorn Layers
PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations
From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining
Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval
Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Servi…
Mapping the Emerging Social Science of Large Language Models
Thermodynamic Cyclic Processes with Markov Samplers in Bayesian Inference
IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monito…
Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampli…
Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
A radiographic world model for clinical reasoning and evidence generation
A Tool-Augmented, GPT-4 Chatbot for Real-Time Repository Data Analysis
Latent-to-Latent Flow for Volumetric Stochastic Segmentation
SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Inter…
What Does an LLM-Agent Leaderboard Rank Actually Compare?
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization
RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavior…