Agent Skills: Judging Test Quality

Grade and review an existing spec suite for kill power before adding more. Use when you must find whether specs are real or cosmetic — each gets STRONG/ACCEPTABLE/WEAK/KILL plus one mutation it survives. Triggers on "audit these tests", "how strong is this", "mutation score", "green but buggy".

UncategorizedID: ANcpLua/ancplua-claude-plugins/judging-test-quality

Install this agent skill to your local

pnpm dlx add-skill https://github.com/ANcpLua/ancplua-claude-plugins/tree/HEAD/plugins/mutation-minded-testing/skills/judging-test-quality

Skill Files

Browse the full folder contents for judging-test-quality.

Download Skill

Loading file tree…

plugins/mutation-minded-testing/skills/judging-test-quality/SKILL.md

Skill Metadata

Name
judging-test-quality
Description
Grade and review an existing spec suite for kill power before adding more. Use when you must find whether specs are real or cosmetic — each gets STRONG/ACCEPTABLE/WEAK/KILL plus one mutation it survives. Triggers on "audit these tests", "how strong is this", "mutation score", "green but buggy".

Judging Test Quality

Core principle

A test's value is not "it passes". A test's value is "it would fail if the code were broken in a realistic way". Coverage is not evidence. Green is not evidence. The only evidence is mutation resistance.

When to invoke this skill

  • Before writing new tests in an area — know the existing baseline first.
  • After a bug ships that the suite "should have caught".
  • When a test file has 100% coverage but users still report regressions.
  • When onboarding to a module and the suite feels suspicious.

The rubric

For every test, assign exactly one verdict:

| Verdict | Criterion | |---------|-----------| | STRONG | Would die under multiple plausible mutations. Asserts exact semantic effects. Tests a contract. | | ACCEPTABLE | Would die under at least one plausible mutation. Assertions are specific enough to have meaning. | | WEAK | Survives common mutations. Assertions are vague, happy-path-only, or implementation-coupled. | | KILL | Actively misleading. Asserts nothing meaningful or only internal mechanics. Delete or rewrite from scratch. |

No "borderline" verdicts. When unsure between two, pick the lower.

The mutation checklist

For each test, ask: would this test fail if any of these mutations were applied to the production code?

  • >>=, <<=, ==!=
  • &&||
  • truefalse as a literal
  • Remove a throw; return null / default instead
  • Return empty collection instead of the real result
  • Silently swallow an error (catch {})
  • Skip a side effect (no reload after save, no selection update after delete)
  • Invert a guard
  • Off-by-one in a slice or range
  • Return the input instead of the computed result
  • Reorder two side effects that must run in a specific order

Automatic downgrades

These apply across the runners mutation testing targets — jest and vitest in the JavaScript/TypeScript world, pytest in Python — since the same vacuous assertions show up in every one. A test earns at most WEAK if any of these apply:

  • Primary assertion is toBeTruthy(), toBeDefined(), toBeFalsy(), or not.toBeNull() with no follow-up shape check.
  • Primary assertion is array.length === N with no content check.
  • Primary assertion is mock.toHaveBeenCalled() without toHaveBeenCalledWith(specific) and an external state or output check.
  • Snapshot without a hand-crafted structural assertion alongside it.
  • Test name is "should work" / "should be defined" / "should render".
  • Only the happy path; no failure-path sibling exists for the behavior.
  • Asserts on private state, internal caches, or impl-only helpers.

The contrast is concrete. A weak vitest assertion survives the mutation that returns the input unchanged; a strong one names the exact computed value:

// WEAK — survives "return input instead of computed result"
const outWeak = applyDiscount(cart, "SAVE10");
expect(outWeak).toBeDefined();
expect(outWeak.items.length).toBe(cart.items.length);

// STRONG — dies under arithmetic, rounding, and off-by-one mutations
const outStrong = applyDiscount(cart, "SAVE10");
expect(outStrong.total).toBe(90);
expect(outStrong.discountApplied).toBe(10);

The same shape holds in pytest: assert result is weak; assert result.total == 90 dies the moment the discount math is mutated.

Negative-space rule

A behavior is not fully tested unless at least one test asserts what must NOT happen:

  • save with invalid input → no HTTP write fires.
  • load error → existing selection is not silently overwritten.
  • delete non-selected item → current selection unchanged.
  • null input → no network call.
  • malformed data → no crash; typed fallback returned.

A behavior with a positive-path test and no negative-space counterpart is incomplete, even if the positive test is ACCEPTABLE.

Output shape

For each file:

## `path/to/file.spec.ts`

### Summary
Total: N · STRONG: a · ACCEPTABLE: b · WEAK: c · KILL: d
Behaviors missing negative-space: [list]

### Verdicts
<one entry per test with: verdict, surviving mutation, recommendation>

End with a suite-level triage table and one sentence naming the single weakest behavior — the one where a realistic production bug would ship undetected.

What this skill will NOT do

  • Suggest specific new assertions — hand those to improving-weak-tests.
  • Write replacement tests — hand those to the expressive-verifier-improver agent.
  • Celebrate high coverage without a mutation-resistance check.
  • Grade tests against implementation details ("does this test cover every branch in the function").

Related

  • For the upstream cause of weak tests: see reviewing-testability.
  • For rewriting specific weak tests: see improving-weak-tests.
  • For end-to-end coverage closure: see mutation-resistant-coverage.