Explorar

Noticias de IA

27413 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
How Proper Scoring Rules Shape LLM Forecasting
arXiv cs.AI Research & Papers
Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failu…
arXiv cs.AI Research & Papers
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Mo…
arXiv cs.AI Research & Papers
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv cs.AI Research & Papers
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
arXiv cs.AI Research & Papers
If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement a…
arXiv cs.AI Research & Papers
CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action
arXiv cs.AI Research & Papers
MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
arXiv cs.AI Research & Papers
When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded…
arXiv cs.AI Research & Papers
SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question Answeri…
arXiv cs.AI Research & Papers
A Framework for Object-Centric Predictive Monitoring of Collaborative Processes
arXiv cs.AI Research & Papers
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
arXiv cs.AI Research & Papers
A comprehensive and trustworthy benchmark of AI methods for change detection in Earth obs…
arXiv cs.AI Research & Papers
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
arXiv cs.AI Research & Papers
Rethinking Vacuity for OOD Detection in Evidential Deep Learning
arXiv cs.AI Research & Papers
See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs
arXiv cs.AI Research & Papers
SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts
arXiv cs.AI Research & Papers
Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Exper…
Hacker News (AI filter) Research & Papers
I accidentally turned LLM memory into program analysis
TechCrunch AI Research & Papers
An Anthropic researcher just gave us a peek at self-improving AI
arXiv cs.AI Research & Papers
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Pr…
arXiv cs.AI Research & Papers
Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
arXiv cs.AI Research & Papers
Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
arXiv cs.AI Research & Papers
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation
arXiv cs.AI Research & Papers
Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Sear…
arXiv cs.AI Research & Papers
Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detecti…
arXiv cs.AI Research & Papers
Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Mo…
arXiv cs.AI Research & Papers
A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Ans…
arXiv cs.AI Research & Papers
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
arXiv cs.AI Research & Papers
Fine-Tuning of Transformer models with Frames