not much happened today
**Harness engineering** is emerging as the key differentiator for coding agents, emphasizing the stack of **model + harness + eval loop** over just stronger base models. **DeepSeek** is building a harness team to optimize interaction and verification loops, while **Google's Gemini Managed Agents** and **LangChain** formalize harness concepts like context governance and dynamic skill routing. New benchmarks like **DeepSWE** align closely with real developer experience, with **Qwen3.7 Max** and **Claude Opus 4.6** showing strong agentic coding performance. **Anthropic** introduced a security-guidance plugin for **Claude Code** reducing security PR comments by 30–40%, and **OpenAI** highlighted **GPT-5.5** in Codex for improved document parsing. In research, **Claude Mythos** solved Erdős problem #90 with a cleaner proof path than previous models, showing latent capabilities unlocked by appropriate harnesses. The paper "Language Models Need Sleep" proposes a sleep-like consolidation phase for long-horizon memory, addressing bottlenecks in persistent context storage. Open research agents like **QUEST** (2B–35B parameters) advance long-horizon fact-seeking and citation grounding, while the **CUSP benchmark** from Sakana/Stanford/Oxford/AI2 evaluates current model capabilities in science.