Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

AITopTools Editorial TeamSeptember 11, 2026

What changed

ByteDance Seed and research partners introduced HarnessDev, a test that evaluates the working setups AI models build to complete tasks, rather than only judging their final answers. Across five tests and 2,207 tasks, the setups matched human references in writing and machine-learning experiments but performed worse in coding and search; only 34 of 64 improvements also worked on new tasks.

What this means for you

The findings suggest AI models can improve the tools and instructions they use for some tasks, but those improvements may not transfer reliably to unfamiliar work. HarnessDev is a research benchmark, not a new consumer product, so it offers evaluation rather than immediate access to a new capability.

Related AI tools

Explore directory listings connected to the products, companies, and workflows in this story.

Related AI news

Read Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
Ideas & discoveriesSep 12, 2026

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

This is mainly a research result rather than a ready-to-use product: the evidence suggests that adding biological wiring does not automatically improve a language model. The code is MIT-licensed and can be run locally, but the reported benefit is small and the graph-based design performed worse than its simpler control.

MarkTechPostSee why it matters
Read Anthropic CEO says it’s time to pump the brakes on AI
Safety & rulesSep 12, 2026

Anthropic CEO says it’s time to pump the brakes on AI

Independent evaluators may get a stronger role in checking whether Anthropic follows its safety commitments. This is a proposed approach rather than a new consumer feature, so its practical impact will depend on how the evaluations work and what Anthropic does with their findings.

The Verge AISee why it matters