Hierarchical Latent Prediction for Language Models
Explorar
Noticias de IA
29347 elementos — filtrados, clasificados y sin duplicados
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
Shapes from Examples: Foundations of Shape Learning in Recursive SHACL
Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Age…
Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
Why the Third Axis Is Freedom
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architect…
CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New S…
D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation
Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for …
Challenges for Musical Education in the Age of AI and Digital Transformation
Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborat…
Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Se…
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
BaKron: Efficient Quantization with Kronecker-Factored Hessians
Failing Gracefully: Mitigating Impact of Inevitable Robot Failures
An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
Agentic Software Issue Resolution with Large Language Models: A Survey
CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical…
A note on conditional PAC-efficient reasoning in large language model routing
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implication…
IMMENSE: Inductive Multi-perspective User Classification in Social Networks
The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. con…
Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn V…