Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task …
The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybr…
The Compaction Cliff in Long-Running AI Agent Memory
Search Broadly, Seek Evidence on Both Sides, Decide Narrowly: Evidence-Admissible GraphRA…
Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classi…
Molecular LLM Agents: From Architectural Design to Scientific Autonomy
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When…
Complexity Induction: Compositional Generalization via Structured Label Distortion
Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual …
ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD
Diagnosing Capability Preservation and Task Sensitivity in Memory Augmented Document Clas…
Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents
Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?
A Survey Instrument to Assess Students' AI and Generative AI Knowledge
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resourc…
Correcting Variable Importance Scored by Random Forests
PepLLM: ESM-Guided Llama for Structured Protein-Peptide Binding Interface Analysis
Prime Agent: A Self-Improving RLM Harness
ReWorld: An Interactive World Model with Long-Horizon Memory
Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, a…
Bulbul: A Dataset for Dialectal Arabic Speech Recognition
MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interact…
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models
Walking on the DARKSIDE
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
AI emotional support is better only when chosen, but shifts preferences even when it is n…