How Proper Scoring Rules Shape LLM Forecasting
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failu…
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Mo…
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement a…
CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action
MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded…
SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question Answeri…
A Framework for Object-Centric Predictive Monitoring of Collaborative Processes
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
A comprehensive and trustworthy benchmark of AI methods for change detection in Earth obs…
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Rethinking Vacuity for OOD Detection in Evidential Deep Learning
See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs
SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts
Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Exper…
I accidentally turned LLM memory into program analysis
An Anthropic researcher just gave us a peek at self-improving AI
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Pr…
Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation
Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Sear…
Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detecti…
Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Mo…
A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Ans…
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
Fine-Tuning of Transformer models with Frames