·8 min readAI Skills

I Tested All 36 of Matt Pocock’s AI Skills

A breakdown of which skills actually deliver, which ones are thin wrappers, and what is missing for production.

Look, I know you saw that repo with 135K stars and thought “damn, this must be good.” I thought the exact same thing. So I tore apart all 36 skills in there and tested whether they actually do what they promise. Here is what I found: no fluff, just real breakdown.


Wait, What Even Are “AI Skills”?

Before we dive in, let’s get the core concept straight.

Imagine you hire a sharp assistant. But when you tell them “build me a website,” they immediately start typing code without asking you a single question. They don’t ask what colors you want, who your users are, or whether you need an auth system. They just build blindly.

That’s an AI without skills.

Now imagine you hand that same assistant a playbook before they touch the keyboard. The playbook says:

  • First, interview the user about what they actually need.
  • Then write tests before writing any implementation code.
  • Review your own work before presenting it.

That’s an AI with skills. It’s simply a markdown document that instructs the AI on how to operate for a specific workflow.

Matt Pocock created a collection of these playbooks. Some are genuinely brilliant. Some are, well, like buying a sports car and finding out it has no engine inside.

36 Skills: The Big Picture

rich skills


Why Did I Test These?

Before getting carried away by star counts, there is one question worth answering: why analyze them in the first place?

Because popular does not mean bulletproof.

Matt’s repository has over 135,000 stars. Thousands of developers use these skills daily. That is real community proof, and these skills clearly offer value.

But star counts don’t tell you which skills do the heavy lifting, which ones are along for the ride, or whether they hold up under real development pressures.

So I put them through an evaluation framework similar to what enterprise teams use for AI agents:

  1. Tier 1 (Static Analysis): Does the structure look solid under the hood?
  2. Tier 2 (Duplicate Detection): Is it an original workflow or just a thin wrapper?
  3. Tier 3 (Workflow Testing): Does the skill actually produce the promised outcome?

Here is how Matt’s skills scored.


The Audit Numbers

What Passed (The Good News)

Check Result What This Means In Practice
Formatting & Structure 36/36 passed Clean, well-structured markdown files
Safe Commands 36/36 clean No dangerous destructive scripts
Secrets & Keys 36/36 clean No leaked API credentials or keys
Link Integrity 36/36 passed Internal cross-references resolve cleanly

The skills are properly formatted and safe to run. That is like a restaurant having clean plates: necessary, but you still care about the food.

What’s Missing (The Reality Check)

Check Result What This Means In Practice
Automated test cases 0/36 No automated tests to catch regressions
Benchmark reports 0/36 No documented with-vs-without metrics
Negative trigger tests 0/36 No guardrails against triggering at the wrong time
Version locking 0/36 Changes to base skills can break dependent skills silently

Translation: These skills work, and thousands of developers vouch for them. But there is no automated test harness. They are battle-tested by human trial and error.

That means:

  • When they work, it is because users figured out the nuances manually.
  • When they break, you only find out while actively coding.
  • There is no automated CI pipeline confirming that an update didn’t regress an existing workflow.

The Crown Jewels: Skills That Actually Change How You Work

crown_skills

1. grilling: The Structured Interviewer

What it does: Before touching code, the AI interviews you. Not like an unstructured chat where it repeats random questions, but in a strict dependency sequence.

Why it’s brilliant: If you are planning a trip, an undisciplined AI asks “beach or mountains?” before checking if you have a passport or a budget. grilling forces the AI to resolve prerequisite questions first. It follows a clean dependency chain (a frontier algorithm for conversations).

The numbers: Quality 92/100 (A-). Estimated workflow lift: +35%. Completely self-contained.

The verdict: Without this, the AI starts writing code immediately and often builds the wrong feature. With this, it spends 5 minutes clarifying requirements first, saving hours of refactoring. This is the strongest skill in the entire repo.


2. tdd: Test-Driven Development That Teaches

What it does: Forces the AI to write failing tests before implementing features (Red -> Green -> Refactor).

Why it’s special: Most generic TDD prompts simply say “write tests first.” That is too vague. Matt’s skill goes deeper by teaching the AI what bad tests look like:

  1. Implementation-coupled tests: Tests that break during routine refactoring even when user-facing behavior never changed.
  2. Tautological tests: Tests that merely duplicate the implementation logic instead of asserting behavior.
  3. Horizontal slicing: Writing all tests upfront without validating individual feedback loops.

The numbers: Quality 88/100 (B+). Estimated workflow lift: +30%.

The verdict: The AI writes tests that assert behavior instead of fragile implementation details. When you refactor later, the tests remain reliable.


3. code-review: Dual-Track Reviewing

What it does: Evaluates code changes across two independent perspectives simultaneously:

  • Standards track: Code cleanliness, project conventions, and smell detection.
  • Spec track: Verifies whether the change actually fulfills the requested feature ticket.

Why it matters: Most reviews catch either code formatting or logic errors, rarely both. This prompt runs two evaluation paths so they don’t bias each other.

The numbers: Quality 88/100 (B+). Estimated workflow lift: +25%.


The Thin Wrappers: Skills That Do Almost Nothing

Now let’s talk about the weak spots in the collection.

Some skills in the repository are essentially one-line redirects:

Skill File Size What It Actually Does If The Core Skill Breaks
grill-me 157 bytes “Call the grilling skill” Completely broken
grill-with-docs 247 bytes “Call grilling + domain-modeling” Completely broken
implement 433 bytes “Use TDD, then code-review, then commit” Severely degraded

At 157 bytes, grill-me is just a pointer note. For comparison, grilling is nearly 2KB of detailed instructions that do all the heavy lifting.

Because there is no version locking, if grilling changes its format or naming, these wrapper skills can fail silently without warning.


The Over-Engineered One: improve-codebase-architecture

This skill sounds great on paper. It claims to scan your codebase, identify structural issues, and output an interactive HTML report.

The reality:

  • It injects 24,500 tokens into context before reading any of your project code (nearly 20% of a standard context window).
  • Over 6KB of the skill is dedicated to styling HTML, Mermaid diagrams, and SVGs.
  • The model often spends more compute formatting presentation markup than reasoning about system design.

The verdict: Skip the bloated HTML report. Ask for a clean Markdown architecture review instead. You get the same technical insights at a fraction of the token budget.


What Is Missing for Production Readiness

Missing Pieces

The production readiness evaluation across the 36 skills.

  1. Zero Automated Evaluation Suites: None of the 36 skills include machine-readable eval sets to test for prompt regressions across model updates.
  2. Missing Negative Triggers: Model-invoked skills lack explicit conditions for when not to activate (e.g., triggering a full TDD cycle for a one-line typo fix).
  3. Implicit Dependency Chains: Skills reference other skills by name without pinned versions.

Which Skills Should You Actually Use?

Start Here

Recommended adoption order.

For Individual Developers

  1. /grill-me (or grilling): Use this to nail down requirements before writing code.
  2. /tdd: Enforces strict test-first development with clean assertions.
  3. /code-review: Catches spec discrepancies and stylistic issues before committing.
  4. /domain-modeling: Great for establishing clear naming conventions across features.

For Teams

  • Recommended: grilling, tdd, code-review, domain-modeling, codebase-design.
  • Use with caution: wayfinder (useful for deep exploration, but context-heavy).
  • Skip: improve-codebase-architecture (too token-heavy) and thin wrapper skills (invoke base skills directly).

Hype vs Reality

Hype vs Reality

  • High Quality & High Impact: grilling, tdd, code-review.
  • High Impact via Delegation: grill-me, grill-with-docs (work well only because they call grilling).
  • Niche Reference Material: codebase-design, writing-for-agents.
  • Low Value: implement, research, handoff.

Cost vs Value

Cost vs Value

  • Best Value: grilling (2.5K tokens, ~35% quality lift).
  • Solid ROI: tdd (12K tokens) and code-review (11K tokens).
  • Poor ROI: improve-codebase-architecture (24.5K tokens for only marginal improvement over simple prompting).

The Scorecard

Skill Score Recommendation Key Takeaway
grilling 92 (A-) Yes The standout skill. Structured prerequisite interview.
tdd 88 (B+) Yes Teaches test quality and anti-patterns.
code-review 88 (B+) Yes Evaluates both specification accuracy and code health.
codebase-design 90 (A-) Yes Strong architectural vocabulary and reference guides.
domain-modeling 88 (B+) Yes Clarifies domain entities and aligns naming.
writing-for-agents 90 (A-) Yes Excellent meta-guide for writing agent prompts.
wayfinder 85 (B) Yes Good for navigating complex architectural decisions.
diagnosing-bugs 85 (B) Yes Disciplined reproduction and debugging workflow.
prototype 80 (B-) Yes Rapid throwaway exploration.
grill-me 45 (F) Skip Thin wrapper. Use grilling directly.
grill-with-docs 45 (F) Skip Wrapper. Use grilling + domain-modeling.
implement 50 (F) Skip Wrapper. Call tdd and code-review directly.
improve-codebase-architecture 78 (C+) Skip Heavy token footprint for HTML formatting.

Summary

Matt Pocock’s skills repository contains genuinely strong ideas. The grilling interview flow and tdd anti-pattern guides are among the best prompt architectures available publicly.

While they lack the automated evaluation suites and version locking required for rigorous production CI pipelines, the core skills (grilling, tdd, code-review) will immediately improve your daily AI coding workflows.