InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Se…
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Mode…
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error…
Evaluating LLM Generated Detection Rules in Cybersecurity
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal …
Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Ag…
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Hori…
AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search
HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Sy…
Harnessing agent memory to build lifelong AI partners for materials scientists
AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection
Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problem…
VQ-bench: A Composable Vector Quantization Framework
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance A…
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operat…
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: …
Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Con…
Causal inference for group-contaminated structured outcomes: observable quotients, lossle…
Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-t…
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Qu…
Instruction Alignment for Binary Code Representation Learning