Learning-Driven Adaptive Audit Scheduling: A Sequential Decision Approach to Off-Chain Da…
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents
FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision …
Decision Making Needs Uncertainty Quantification [Lecture Notes]
AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Effici…
Human Grounded Evaluation of Large Language Models for Optical Network Automation
A Benchmark for Early-stage Parkinson's Disease Detection from Speech
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Vid…
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Fo…
DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decou…
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LL…
ClawBench: Can AI Agents Complete Everyday Online Tasks?
A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents
HiCI: Hierarchical Construction-Integration for Long-Context Attention
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
ConceptTree: Bringing Semantic Transparency to Black-Box Decision Making for Robotic Mani…
FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Mo…
Integrating High-Level Requirements to Low-Level Tests with Machine-Readable V&V Specific…
Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Sus…
Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool…
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand For…
TypiCore: A Hybrid Active Query Strategy for Class-Incremental Learning on Time Series
Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Iden…
Predictive Training with Latent Imagination for Visual Quadruped Navigation
Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration