Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure …
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science S…
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration
Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection
A Two-Stage Learning PINN Approach for Solving the Inverse Problem of the 1D Porous Mediu…
RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Gr…
NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representatio…
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG…
GATTA: Graph Active Learning with Test-Time Augmentation
Beyond Direct Access: Resource Hijacking in LLM Agents
Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World M…
MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Sol…
Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference
LLMs Get Smarter from Targeted Synthetic Multilingual Data
Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal A…
Orbital AI Computing: Carbon Tradeoffs Across Satellite Scale
Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telem…
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared …
DriveCache: Action-Aware Caching for Driving World Model Inference
A Policy Algebra for Trust-Preserving Agentic AI Execution
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence