NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs
Explorar
Noticias de IA
30334 elementos — filtrados, clasificados y sin duplicados
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Mode…
MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clini…
Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion
In-Context Reward Adaptation for Robust Preference Modeling
On Language Generation in the Limit with Bounded Memory
Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback
PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game?
LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Langua…
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Da…
Benchmarking at the Edge of Comprehension
Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangle…
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approxi…
Recurrent Structural Policy Gradient for Partially Observable Mean Field Games
DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization
When Models Learn to Ask Why: Adaptive Causal Reasoning for Trustworthy Medical Vision-La…
MediHive: A Decentralized Agent Collective for Medical Reasoning
GroundAct: Can LLM Agents Ground Actions in Environmental States?
S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-ground…
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Cryst…
Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teache…