Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking
Explorar
Noticias de IA
30329 elementos — filtrados, clasificados y sin duplicados
ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
Energy-Structured Low-Rank Adaptation for Continual Learning
MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryM…
Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels
Trinity: Unifying Class-Agnostic Terrain and Semantic Segmentation for Unstructured Outdo…
On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note
From Detection to Mechanism: Cross-Attention Graph Neural Networks Enable Drug-Drug Inter…
Fine-Tuned LLM as a Complementary Predictor Improving Ads System
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level En…
AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems
HEAL: Resilient and Self-* Hub-based Learning
Debate Helps Weak Judges Reward Stronger Models
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Sys…
The Future of Facts: Tracing the Factual Generation-Verification Gap
Not All NVFP4 QAT Recipes Are Equal: How Architecture and Scale Shape Model Quality for A…
Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancie…
High-Fidelity Industrial Crash Dynamics Prediction via Geometry-Aware Operator Learning w…
ChildEval: When large language models meet children's personalities
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervis…
SPAR: Support-Preserving Action Rectification
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
Geometry-Correct Diffusion Posterior Sampling with Denoiser-Pullback Curvature Guidance a…
Learning Compositional Latent Structure with Vector Networks
StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment