Before You Poll with LLMs: A Deliberative Diagnostic Framework
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
The Misery of Mechanistic Interpretability: A Formal Perspective
IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Traje…
Specifying Reward Functions for RL Without Environment Sampling
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
Privacy-enhanced federated learning via asynchronous aggregation and local differential p…
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evalu…
Privileged observations enable rapid and reliable policy discovery directly in the physic…
FICAug: Feature-Informed Clustering and Augmentation for Facial-Expression-Based Parkinso…
$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
Estimating Uncertain Spatial Relationships in Robotics
Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks
Positioning manuscripts in the scientific landscape with agentic AI
Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm …
Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language …
Modality-Guided Mixture of Structured Experts with Entropy-Triggered Routing for Multimod…
Echo-CoPilot: A Multiple-Perspective Agentic Framework for Reliable Echocardiography Inte…
Synthetic Data in Marketing Research: How to Evaluate and When to Trust
Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning
ClinicalReTrial: Clinical Trial Redesign with Self-Evolving Agents
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
Question's Gambit: The First Move Matters in Agentic Deep Search
DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis
CRAMER: Control via Request-Aware Masking for Editing Recommenders