Human-AI Collaboration for Estimating Scientific Replicability
Explorar
Noticias de IA
30334 elementos — filtrados, clasificados y sin duplicados
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Ge…
Detection Without Correction: A Two-Parameter Decomposition of Multi-Stage LLM Pipelines
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
Learning after COVID-19 and the ICT career aspirations: Are students entering the AI era …
Personalized Observation Normalization for Federated Reinforcement Learning in Simulation…
I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech …
The Computational Boundary of Inference: Capability Internalization, Training, and the Tu…
CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models
AlphaTransit: Learning to Design City-scale Transit Routes
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reason…
DEPART: DEcomposing PARiTy across Multilingual LLMs
Architecture-driven Shift: towards a lightweight selector for capturing the trends of log…
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of…
SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping
TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-…
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verific…
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy…
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation…
The Ethics of LLM Sandbox and Persona Dynamics
Continual Model Routing in Evolving Model Hubs
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Hybrid Neural World Models
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Act…
Improving Evaluation of Recombination-based Cartesian Genetic Programming