FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Finan…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic …
How People Evaluate AI-, Expert-, and Peer-Style Financial Advice
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
Learning an Interior Layout Policy in a Domain Specific Language Action Space
Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasonin…
ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
Deep probabilistic logic programming for diagnostic reasoning from incomplete information…
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Lang…
A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scali…
A foundation model of numerical intelligence with cross-disciplinary generalization
FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Seri…
VTO: Visual Tool Orchestration for Video Anomaly Detection
Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via…
El Agente Gr\'afico: A Semantic Execution Runtime for Scientific Agents
MOSAIC: Adversarial Co-evolution of Specialist Heuristics and Problem Instances for LLM-b…
AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification
Biologically Informed Representation Learning for Robust Cross-Center Generalization of M…
$\texttt{DisMorph}$: learning to disentangle technical distortions from true biological c…
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-I…
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Mar…
IntelliAudit: Using Large Language Models to Evaluate Audit Controls
A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges
Generalizing deep reinforcement learning across cable-driven parallel robot configuration…