Herculean: An Agentic Benchmark for Financial Intelligence
Explorar
Noticias de IA
29655 elementos — filtrados, clasificados y sin duplicados
From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Mac…
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Ru…
Ethical Hyper-Velocity (EHV): A Hardware-Rooted Zero-Trust Runtime Enforcement Architectu…
LLM-Guided Communication for Cooperative Multi-Agent Reinforcement Learning
Towards a General Intelligence and Interface for Wearable Health Data
c-TPE: Tree-structured Parzen Estimator with Inequality Constraints for Expensive Hyperpa…
Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles …
DeepIPCv2: LiDAR-powered Robust Environmental Perception and Navigational Control for Aut…
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
Implicit Regularization for Multi-label Feature Selection
Self-supervised Monocular Depth and Pose Estimation for Endoscopy with Latent Priors
Introduction to Graph Neural Networks for Machine Learning Engineers
ShapeLib: Designing a library of programmatic 3D shape abstractions with Large Language M…
Efficient LLM Moderation with Multi-Layer Latent Prototypes
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred…
Ideas in Inference-time Scaling can Benefit Generative Pre-training Algorithms
T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
MARFT: Multi-Agent Reinforcement Fine-Tuning
A Survey of 3D Reconstruction with Event Cameras
Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
GFlowGR: Fine-tuning Generative Recommendation Frameworks with Generative Flow Networks
Hyperspherical Variational Autoencoders Using Efficient Spherical Cauchy Distribution
FedS2R: One-Shot Federated Domain Generalization for Synthetic-to-Real Semantic Segmentat…
Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representatio…
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignme…
Towards a Physics Foundation Model
T-POP: Test-Time Personalization with Online Preference Feedback