The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probabili…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models
HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representati…
PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents
Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms
Partial Identification under Causal Orders by Linear Programming
Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Proce…
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Visi…
CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support
Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnost…
FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audi…
From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human O…
Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watche…
Method, Mind, and Morality: How People Make Sense of Artificial Intelligence
Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable …
ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory For…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device La…
Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target obser…
ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Callin…
ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Le…
PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understandi…
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Ba…
Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Spa…
Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tie…
SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Percepti…
Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems