EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL
Explorar
Noticias de IA
30675 elementos — filtrados, clasificados y sin duplicados
Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with…
The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologicall…
Tractable Hierarchical Control of Autoregressive Language Models
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source…
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detecti…
Semi-Supervised Text-Attributed Graph Distillation
Benchmarking the Personalization Capabilities of Large Language Models
Robust Critics: Defending LLMs Against Multi-Turn Attacks
PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs
DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligen…
RUMBA: Russian User Memory Benchmark
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-su…
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
OPOD: On-Policy Omni Distillation
Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
DynaMark: A Reinforcement Learning Framework for Dynamic Watermarking in Industrial Machi…
Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driv…
Loss-Complexity Landscape and Model Structure Functions
Generative AI and Agency in Education: A Critical Scoping Review and Thematic Analysis
Representative Sets in Propositional Abduction
Synthetic minority data is redundant or invalid: a data-dependent validity theory and a d…
Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro