AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Tra…
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in C…
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
Governed Shared Memory for Multi-Agent LLM Systems
A specialized reasoning large language model for accelerating rare disease diagnosis: a r…
On the Smallness of the Large Language Models Scaling Exponents
The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-imag…
LemonHarness Technical Report
Tractable Reasoning and Conjunctive Query Answering for Defeasible DL-Lite under Rational…
Probing the Misaligned Thinking Process of Language Models
Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach
Exploring the relationship between human-centric AI and firm idiosyncratic risks
MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and F…
QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Scie…
Navigating User Behavior toward Personalized Multimodal Generation
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserst…
T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Lay…
OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility
BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding
Female-RHINO: A Real-Time Scanner-Integrated Framework for Automated Quantitative Uterine…
Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-I…
Entity Resolution via Batched Oracle Queries
Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million…
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Underst…
RetiSEM: Generalising Causal Models for Fragmented Biomedical Data
Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality With…
Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storyte…