The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologicall…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with…
EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL
CRAWO: Custom Resources for Adaptive Workload Orchestration
OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraini…
Attention-based Experience Replay Framework for Continual Learning of Agnostic Time Serie…
Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under A…
From Errors to Rules: Iterative Prompt Optimization for Text Classification
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge …
ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthes…
FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Inv…
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization
Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
Representation Robustness Under Executable Reasoning Constraints in Large Language Models…
Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Interv…
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP S…
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
StrideDiffusion: Accelerating Diffusion Models for Time-series Generation
CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement…
AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optim…
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
Can an AI System Be Creative? A Critical Perspective from Art and Engineering
Profiling Lightweight Large Language Models
Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Conv…