PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Explorar
Noticias de IA
22065 elementos — filtrados, clasificados y sin duplicados
Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessme…
From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representat…
Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners
Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
ScalableRAG: High-Quality RAG at Zero Ingestion Cost
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Whi…
Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive …
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical …
Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios
Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for D…
scMIR: a vision-language foundation model for single-cell light microscopy image represen…
RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Predi…
The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the P…
EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
Deep Delta Learning
Localizing Persona Representations in LLMs
Do Models Fake Alignment Without Clear Consequences?
GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
OmniQEC: discovering practical quantum error-correcting codes by an AI scientist
Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume …
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
Measuring the State of Open Science in Transportation Using Large Language Models