RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
Browse
AI News
30934 items — filtered, classified, deduplicated
Recursive Flow Matching
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentat…
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Langua…
The ATOM Report: Measuring the Open Language Model Ecosystem
Governed Capability Evolution: Lifecycle-Time Compatibility Checking and Rollback for AI-…
SenBen: Sensitive Scene Graphs for Explainable Content Moderation
Cryptographic Registry Provenance: Structural Defense Against Dependency Confusion in AI …
Understanding the Challenges in Iterative Generative Optimization with LLMs
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Ques…
Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argum…
StreamSplit: Continuous Audio Representation Learning via Uncertainty-Guided Adaptive Spl…
Algorithmic Monocultures in Hiring
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagati…
Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinfo…
APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented…
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-…
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
Geometrically Constrained Outlier Synthesis
GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User Hist…
Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts …
Constructing Industrial-Scale Optimization Modeling Benchmark
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Pos…
Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Sp…