CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs
Explorar
Noticias de IA
30353 elementos — filtrados, clasificados y sin duplicados
From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation
Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Networ…
Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric A…
Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective P…
Specula: Scaling formal specifications for autonomous model checking of system code
Hybrid Analysis for Secure MCP Tool Use in LLM Agents
ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields
AnnoBench: A Benchmark for Visualization Annotation Generation
Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governa…
Three Sides of Retrieval: Factorial Evidence for Document-Side, Query-Side, and Answer-Si…
DDSNet: Dual-domain Symmetry-aware Network for PCSEL Property Prediction
Grounded in Consensus, In Step With Emerging Science: A Consensus-Anchored Multi-Corpus C…
Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Crit…
Automated Numerical Stability Analysis of Deep Learning Operators
Real-time Spatial Retrieval Augmented Generation for Urban Environments
Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for D…
LLM-generated personalized nudges for improving pro-environmental behavior: Field evidenc…
Eliminating Propagation Delay: Attention-Based Spatial-Temporal Fusion Graph Convolution …
Multiclass Classification without Labels via Posterior Simplex Geometry
Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM…
Where Steering Signals Come From: Activation Source Selection in Activation Steering
LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledg…
"We'll have to see how it works": An interview study to understand collaborative practice…
COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations
On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Revie…
JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search a…
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities