JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
FLUIDSPLAT: Reconstructing Physical Fields from Sparse Sensors via Gaussian Primitives
CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intellige…
Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts …
Constructing Industrial-Scale Optimization Modeling Benchmark
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Pos…
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
Governed Metaprogramming for Intelligent Systems: Reclassifying Eval as a Governed Effect
Post-training makes large language models less human-like
JobBench: Aligning Agent Work With Human Will
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and…
AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compressio…
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Proce…
What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global De…
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?
High-Quality Synthetic Financial Time-Series using a GAN-Diffusion Framework
A Sharper Picture of Generalization in Transformers
The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypot…
Maat: The Agentic Legal Research Assistant for Competition Protection
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluati…
GEM: Geometric Entropy Mixing for Optimal LLM Data Curation
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents