VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Efficient Exploration Is Enough
Support Topology and Gradient Mixing in Sinkhorn Layers
CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information
CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows
We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation
Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection
Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem
Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantize…
Simulating the Marginal Green Contribution of AI Modules in a Smart-Agriculture Platform:…
Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool Actio…
Monte Carlo-Based Ex-Ante Assessment of the Green Benefits of an AI-Driven Smart Agricult…
Formation of structural attractors in neuromorphic systems
Latent-to-Latent Flow for Volumetric Stochastic Segmentation
From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning C…
Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specif…
TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tabl…
We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consc…
NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagatio…
Federated Binary Gating with Server-Side Vision-Language Inference for Surveillance Anoma…
Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-…
Generating Instance Generators in PDDL Planning
SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores
Generator-Independent Runtime Assurance under Partial Observation
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts
Learning transferable human physiology from two million hours of sleep with SleepFM-2
AutoKD: Autonomous Knowledge Discovery