When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with…
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in …
Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work
Environmental Slow AI: Design Principles for Generative Systems
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their O…
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
Volumetric Radiology AI in the Era of Multimodal Large Language Models
Dual-Cache Latent Space Communication between Heterogeneous Language Models
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical …
SENTRY: Deterministic, Intelligent Risk Assessment for IT Change Management
ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure
When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a R…
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic …
Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Ap…
STCO: Conditional Neural Operators for Time-Dependent PDEs
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification
SDAD: Spec-Driven Agentic Development for the AI-Native SDLC
Who Delegates to AI? Evidence from 53,000 Agent Configurations
Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariat…
TRACE: Training-time Report-guided and Clinically Ordered Concept Editing
Scaling Muon for Diffusion Transformers
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intell…
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models