Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

AITopTools Editorial TeamSeptember 11, 2026

What changed

Anthropic has published a plugin-evaluation workflow for Claude Code. The new `claude plugin eval` command tests plugins with realistic prompts, grades Claude’s results, and compares them with results from a run without the plugin. It helps developers measure whether a plugin activates and improves responses.

What this means for you

Plugin developers can use these checks to catch problems before release and add them to continuous integration, which automatically tests changes. The feature is aimed at Claude Code users building or maintaining plugins, rather than everyday users.

Related AI tools

Explore directory listings connected to the products, companies, and workflows in this story.

Related AI news

Read Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
New AI featuresSep 15, 2026

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

The models are available now in Google’s Gemini API and AI Studio, allowing developers to build voice applications at $0.005 per minute for audio input. Generated audio includes Google DeepMind’s SynthID watermark, which identifies it as AI-created.

MarkTechPostSee why it matters
Read Meta now lets AI agents handle the boring parts of WhatsApp Business setup
AI toolsSep 15, 2026

Meta now lets AI agents handle the boring parts of WhatsApp Business setup

Developers can use tools such as Claude, Cursor, Codex, and ChatGPT to reduce the manual work involved in launching WhatsApp Business messaging. The feature is aimed at developers, and the feed does not specify pricing or broader access details.

TechCrunch AISee why it matters
Read The AI graveyard: a running list of projects and startups that didn’t make it
AI toolsSep 15, 2026

The AI graveyard: a running list of projects and startups that didn’t make it

The list helps users, creators, and businesses separate announced AI plans from products that are actually available and working as expected. It also highlights the risk of relying on projects that may be delayed, changed, or abandoned.

TechCrunch AISee why it matters