Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

AITopTools Editorial TeamSeptember 11, 2026

What changed

ByteDance Seed and research partners introduced HarnessDev, a test that evaluates the working setups AI models build to complete tasks, rather than only judging their final answers. Across five tests and 2,207 tasks, the setups matched human references in writing and machine-learning experiments but performed worse in coding and search; only 34 of 64 improvements also worked on new tasks.

What this means for you

The findings suggest AI models can improve the tools and instructions they use for some tasks, but those improvements may not transfer reliably to unfamiliar work. HarnessDev is a research benchmark, not a new consumer product, so it offers evaluation rather than immediate access to a new capability.

Related AI news

Read Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
New AI featuresSep 15, 2026

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

The models are available now in Google’s Gemini API and AI Studio, allowing developers to build voice applications at $0.005 per minute for audio input. Generated audio includes Google DeepMind’s SynthID watermark, which identifies it as AI-created.

MarkTechPostSee why it matters
Read Meta now lets AI agents handle the boring parts of WhatsApp Business setup
AI toolsSep 15, 2026

Meta now lets AI agents handle the boring parts of WhatsApp Business setup

Developers can use tools such as Claude, Cursor, Codex, and ChatGPT to reduce the manual work involved in launching WhatsApp Business messaging. The feature is aimed at developers, and the feed does not specify pricing or broader access details.

TechCrunch AISee why it matters