Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Under…
Browse
AI News
27413 items — filtered, classified, deduplicated
(How) Do MLLMs Report Bistable Images Like Humans?
From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collabor…
A Conservative OCR-Enabled Workflow for R214 Sodium Screening of South African Packaged F…
Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LL…
From Voice to Value: Leveraging AI to Enhance Spoken Online Reviews on the Go
SkillAtlas: An Attack Trace Library for Agent Skills
Multi-Agent Empowerment and Emergence of Complex Behavior in Groups
Self-Evolving Memory for Generative Recommendation
Predictive audio representations for early detection and tracking of hidden dynamic objec…
Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal S…
Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering
Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Coll…
Zonal RL-RRT: Integrated RL-RRT Path Planning with Collision Probability and Zone Connect…
Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Di…
SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection
Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the …
Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision
Thought without systematicity? Evaluating reasoning models on rule induction tasks
From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Imag…
BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds
Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning
LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents
Talking to Me or Someone Else? Rethinking Talk-to-Me Detection in Egocentric Videos
Utility-Guided Agent Orchestration for Efficient LLM Tool Use
RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself