← Todas las noticias

not much happened today

**Prime Intellect's `prime-rl` v0.6.0** advances agentic reinforcement learning infrastructure supporting **1 trillion parameter MoE models** with sub-5-minute step times and a **131k context GLM-5 agentic setup**. The release includes optimizations in inference, training, and rollout orchestration, supporting models like **GLM5, Kimi, Nemotron**. **Anthropic's Claude Tag** exemplifies the shift to persistent, asynchronous agents embedded in organizations, already writing **65% of the product team's code** and operating as background watchers and proactive task executors in workflows. The ecosystem features innovations like **StarAgent**, **Self-Harness**, **Hermes Agent**, and **Executor's MCP gateway** for operational agent fleets. **GLM-5.2** gains momentum as a leading open model, especially for coding and agentic workflows, raising security concerns about enabling private offensive workflows without API logging. This highlights a broader trend of agent training becoming an infrastructure challenge, with emphasis on open post-training stacks, verifiable environments, and task-specific rollouts.
Leer el original en AINews / smol.ai →