SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative…
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
Analysis of Prompt Engineering for Drug Toxicity Prediction
Rethinking World Models for Safety-Critical Embodied Systems
Environment Evolution for Terminal Agents
The Attention Triangle in Audio-Video Models
Counterfactual Routing Using Integer Programming with Constraint Generation
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Re…
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Netw…
What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking P…
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respirat…
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GU…
Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
AutoGraphForge: Towards Automated Graph Theory Discovery
Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice …
A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-L…
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection
</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination
How Far Can Synthetic Data Take Thai OCR?
FailBench: How Reliable are VLMs at Judging Robot Task Success?