Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
What changed
ByteDance Seed and research partners introduced HarnessDev, a test that evaluates the working setups AI models build to complete tasks, rather than only judging their final answers. Across five tests and 2,207 tasks, the setups matched human references in writing and machine-learning experiments but performed worse in coding and search; only 34 of 64 improvements also worked on new tasks.
What this means for you
The findings suggest AI models can improve the tools and instructions they use for some tasks, but those improvements may not transfer reliably to unfamiliar work. HarnessDev is a research benchmark, not a new consumer product, so it offers evaluation rather than immediate access to a new capability.