AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
Explorar
Noticias de IA
30321 elementos — filtrados, clasificados y sin duplicados
VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Contex…
LNN-PINN: A Unified Physics-Only Training Framework with Liquid Residual Blocks
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for…
StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation
Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer
Preference-Shaped Expected Hypervolume and R2 Improvement: Exact Computation and Monotoni…
Do Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
Tell Me a Story! Narrative-Driven XAI with Large Language Models
Text-Only Data Synthesis for Vision Language Model Training
Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional…
Aligning Language Model Benchmarks with Pairwise Preferences
FinTexTS: Financial Text-Paired Time-Series Dataset via Semantic-Based and Multi-Level Pa…
Towards Rigorous Explainability by Feature Attribution
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning
Can We Formally Verify Neural PDE Surrogates? SMT Compilation of Small Fourier Neural Ope…
Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
Quantum Machine Learning-based 6G edge Network: Enabling Adaptive Communication and Model…
Measuring Massive Multitask Chinese Understanding
Sinc Kolmogorov-Arnold network and its application for solving PDEs with singularities
Manboformer: Learning Gaussian Representations via Spatial-temporal Attention Mechanism
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an An…
Clinical Validation of the Melanoscope AI Mobile Dermoscopy Clinical Decision Support Sys…
Supervised Distributional Reduction via Optimal Transport and Dependence Maximization