Agent Seer: Synthesizing Scenarios from Specification Understanding
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
The Open ASR Leaderboard Adds Its First Global South Language
Piloting the world's first double-blind AI evaluations
Better answers, broader thinking: What students gain from ChatGPT and critical-thinking t…
From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers
When LLM judges agree, should we believe them?
Preparing data for supervised fine-tuning Part 2: Advanced data strategies
Preparing data for supervised fine-tuning Part 1: Formatting and quality
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Profe…
Learning never stops: How AI makes learning continuous
AI models flub these intelligence tests. Can you fare any better?
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Visi…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device La…
CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support
Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audi…
Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnost…
Method, Mind, and Morality: How People Make Sense of Artificial Intelligence
FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human O…
ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Callin…
ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory For…
ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Le…
PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understandi…
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Ba…
When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and th…
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
Quasar: A Programming Language Specialized for LLM Code Actions
Robust Motion Generation using Part-level Reliable Data from Videos
PatientHub: A Unified Framework for Patient Simulation
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving