Self‑Evolving System Handles Multiple Tasks with Self‑Optimizing Agent
agents
| Source: ArXiv | Original article
Researchers propose a self‑evolving harness that enables language‑model agents to serve as their own optimizer across multiple tasks.
A new pre‑print on arXiv (2609.38372v1) introduces a “self‑evolving harness” that lets a language‑model agent act as its own optimizer across a range of tasks. The paper defines a harness as the surrounding code that structures prompts, calls tools, manages context and steers execution. While earlier work let a separate proposer tweak a human‑crafted harness for each benchmark, the authors propose a single agent that recursively refines its own harness through multi‑task pre‑training and continual learning, without external hand‑crafted scaffolding.
The development matters because harnesses have become a bottleneck in scaling agent capabilities. By collapsing the proposer and solver into one self‑modifying system, the approach promises faster adaptation to new workloads and the ability to share improvements across deployments. The idea builds on recent collaborative‑evolution methods such as EvolveNet, which broadcast a shared harness to data‑local agents that each evolve it on their own workload before merging updates. It also echoes Jiaxin Zhang’s “outer‑loop search” concept, where the LLM itself searches over prompts, code and skills rather than relying on gradient‑based optimisation. ByteDance’s Seed team has already demonstrated a related framework, HarnessDev, showing that LLMs can construct and continuously improve their own harnesses based on task feedback.
What to watch next is whether the self‑evolving harness can be integrated into existing agent platforms such as the AREX‑2 self‑improving agents we covered earlier, and how it performs on real‑world coding and reasoning benchmarks that have exposed limits in models like Gemini 4. Follow‑up studies will likely explore scaling the continual‑training loop, measuring cross‑task transfer, and assessing safety implications of agents that rewrite their own execution logic.
Sources
Back to AIPULSEN