Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing
Explorar
Noticias de IA
30294 elementos — filtrados, clasificados y sin duplicados
Tensor-Coord: Algebraic Decomposition of Joint Plan Tensors for Conflict-Free Multi-Agent…
When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk thr…
Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions
Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assist…
Limited Marginal Benefit of Reasoning-Heavy LLM Deployment in ESG Narrative Scoring: A 4-…
Exploiting Search in Symbolic Numeric Planning with Patterns
AdaSTORM: Scaling LLM Reasoning on Dynamic Graphs via Adaptive Spatio-Temporal Multi-Agen…
Architectural Wisdom: A Framework for Governing Optimization in AI Systems
SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthe…
Latent Thought Flow: Efficient Latent Reasoning in Large Language Models
Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Fore…
AI Pluralism and the Worlds It Misses
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reason…
Symbolic Informalization: Fluent, Productive, Multilingual
GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
ACC: Compiling Agent Trajectories for Long-Context Training
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO an…
AgentFairBench: Do LLM Agents Discriminate When They Act?
From Affect Prediction to Affect Forecasting: Evidence for Distinct Information Sources i…
MR-GVNO: A Geometry-Aware Variational Physics-Informed Neural Operator for Mindlin-Reissn…
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event M…
OSGuard: A Benchmark for Safety in Computer-Use Agents
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability