OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
Browse
AI News
30934 items — filtered, classified, deduplicated
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Dat…
Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection
Yes, Q-learning Helps Offline In-Context RL
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
Beyond Questions: Evaluating What Large Language Models (Actually) Know
On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Condit…
CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intellige…
On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach
Generating Robust Portfolios of Optimization Models using Large Language Models
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality…
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
CogAdapt: Transferring Clinical ECG Foundation Models to Wearable Cognitive Load Assessme…
TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL
Searching the Internet for Challenging Benchmarks at Scale
How Reliable are LLMs for Reasoning on the Re-ranking task?
The Two Boundaries: Why Behavioral AI Governance Fails Structurally
LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene…
ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Fe…
Identifiable Token Correspondence for World Models
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
EHRSummarizer: A Privacy-Aware, FHIR-Native Reference Architecture for Source-Grounded EH…
ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driv…
AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito
Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language …
Assessing Per-Sample Membership Inference Vulnerability without Retraining
Learning When to Think While Listening in Large Audio-Language Models
Deep-layer limit and stability analysis of the basic forward-backward-splitting induced n…