A Unified and Reproducible Experimentation Framework for Speech Understanding
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
ReasonOps: Operator Segmentation for LLM Reasoning Traces
Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Ar…
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
VikingMem: A Memory Base Management System for Stateful LLM-based Applications
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaborati…
Unlocking the Working Memory of Large Language Models for Latent Reasoning
Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured …
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Ad…
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
Sustainable Metal-Organic Framework Water Harvesters in the Artificial Intelligence Era
Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federa…
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Super…
FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Stateme…
TIMEGATE: Sustainable Time-Boxed Promotion Gates for Continual ML Adaptation Under Resour…
Planning with the Views via Scene Self-Exploration
Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Arti…
Benchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Models
CB-SLICE: Concept-Based Interpretable Error Slice Discovery
Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search…
Mind-Omni: A Unified Multi-Task Framework for Brain-Vision-Language Modeling via Discrete…
Differentiable Belief-based Opponent Shaping
BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
Personalized Turn-Level User Conversation Satisfaction Benchmark
DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Age…
Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verifi…