Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal St…
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
Recursive Harness Self-Improvement
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning
A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance
Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime St…
When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
An Exam for Active Observers
When Does Muon Help Agentic Reinforcement Learning?
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcemen…
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workf…
A Formally Grounded ODRL Evaluator: Implementation and Comparison
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tu…
Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous V…
How Does Empowering Users with Greater System Control Affect News Filter Bubbles?
On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous L…
Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario…
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding
Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly D…
In-context learning of closed form solution to simple linear regression task using transf…
Agentic Synthesis against Counterexample-Supplemented Sketches
Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detecti…
AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adap…
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Constructi…
Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain …
Cura 1T: Specialized Model for Agentic Healthcare
CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Age…
Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chas…