AI Can Be Easily Persuaded in Clinical Decision Making
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative…
Towards Cognitive Process-Aware Proactive Writing Support
Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA …
An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecast…
AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning
DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Groun…
Reverse N-Wise Output-Oriented Testing for AI/ML and Quantum Computing Systems
PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC
CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding
A rigor-matched audit of periodic-step layer skipping for efficient llm inference: confla…
Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks
mRNA Design and Optimization with Deep Knowledge-Infused Approach
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Ur…
Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management
Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Rea…
FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations
An Agentic Retrobiosynthesis Framework with Learned Frontier Selection
The reach of a verification tool decides its value: A controlled study of verification su…
AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud Agent
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research …
Learning Simple Test-Time Environments for LLM Web Agents
Masked Distillation: Internalizing the Chain-of-Thought in Language Models
post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding …
Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
On the Prospects of Dynamic LLM Conversations in Software Development
Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large …
The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys