Verification Techniques
Scope: Hypothesis testing, root cause analysis, and verification Load if: Bug reported, test failure, proving correctness, root cause analysis Prerequisites: @smith-guidance/SKILL.md
Foundation: Based on PDSA's Study phase (Deming) and Popper's Falsification - understanding WHY something works or doesn't, not just IF it works.
When to use: Debugging, testing hypotheses, validating solutions, proving correctness.
Hypothesis Testing
Strong Inference
Rapid progress through multiple competing hypotheses:
- Devise multiple hypotheses - Not just one, but several alternatives
- Design crucial experiments - Tests that exclude one or more hypotheses
- Execute experiments - Run tests to eliminate hypotheses
- Iterate - Refine remaining hypotheses, repeat
Key insight: Science advances fastest when we actively try to disprove hypotheses, not confirm them.
For debugging:
- Bug: "Login fails intermittently"
- H1: Session storage full
- H2: Race condition in token refresh
- H3: Network timeout on auth server
- Crucial test: Check if failures correlate with session count (tests H1)
Falsification Principle (Popper)
A theory is scientific only if it can be proven false:
- Design tests that could disprove your hypothesis
- Seek evidence that contradicts, not confirms
- One counterexample disproves a universal claim
Anti-pattern: Only running tests you expect to pass Good practice: Actively try to break your own code
Falsify a workaround before presenting it as the solution: when proposing a fix or workaround that depends on external system behavior (MCP/OAuth/API/CLI/feature support), first search the issue tracker for known failures of that exact mechanism. Never present an untested mechanism in a confident voice — say "unverified — let me check" and check. (Triggered 2026-06: proposed two Slack-MCP OAuth setups as if they'd work; both failed; a 30-second search would have found the closed-as-not-planned regression that made the whole route impossible.)
Bugfix Discipline: Trace the Real Path, Reproduce First
Before writing ANY bugfix:
- Trace the real execution path — from the failing input, follow the source to the EXACT branch that input actually takes. Read the code; follow the return/artifact types. A symptom (error string, stack frame) names a place, not the branch. Don't pattern-match a fix (often "reuse this existing helper") from the symptom alone. Never assume the error path — the input may take a success branch instead (e.g. a function returns an empty success artifact, not an error).
- Reproduce with the real input — run the actual failing case and observe the failure before touching code. Static reading is a hypothesis, not a diagnosis.
- Put the fix on the branch the real input hits — confirm by re-running the repro that it now passes.
The test-masking trap (2026-06): a fix placed in a branch the real
input never enters, paired with a unit test that mocks an input to force that
branch → green test, live bug. The test fit the fix instead of reproducing the
bug. Guard both ends: @smith-tests/SKILL.md (never mock the branch/unit under
test; reproduce the bug as a failing test first) and @smith-subagents/SKILL.md
(audit the execution path of a delegated diff, not just its style).
Anti-Workaround Policy
- Only add
# noqa,// NOLINT, or similar inline suppressions when the exception criteria below are met - Only increase timeouts after diagnosing root cause
- Remove the actual dead code rather than merely using a
_prefix to suppress unused-variable warnings - Only disable warnings with documented justification
When lint or test failures occur:
- Apply 5 Whys to find root cause first
- Fix the underlying issue, not the symptom
- Suppressions allowed ONLY when all criteria are met:
- Reason (at least one):
- External library false positive (document which)
- Verified false positive (document why)
- Explicit user approval (cite the approval)
- Mechanism: prefer tool config (ruff.toml, .flake8) for repo-wide patterns; inline comments only for isolated cases with reason on the same line
- Reason (at least one):
Timeout changes require:
- Profiling evidence showing actual duration
- Diagnosis of why the operation is slow
- User approval before increasing
Root Cause Analysis
5 Whys (Toyota)
Root cause analysis through iterative questioning:
- State the problem
- Ask "Why did this happen?"
- Repeat for each answer (typically 5 times)
- Stop when you reach an actionable root cause
Example:
- Bug: Users logged out unexpectedly
- Why? Session expired
- Why? Token refresh failed
- Why? Refresh endpoint returned 401
- Why? Clock skew between servers
- Root cause: NTP not configured on auth server
Caution: Don't stop at symptoms. "Why?" should reach systemic causes.
Explanation Techniques
Rubber Duck Debugging
Explain code line-by-line aloud; when explanation doesn't match code, you've found the bug.
For AI agents: When stuck, explain the problem step-by-step before proposing solutions.
Feynman Technique
Explain simply to reveal gaps: Choose concept → Explain to child → Identify gaps → Review.
If you can't explain it simply, you don't understand it well enough.
Systematic Isolation
Delta Debugging
Minimize failing input: split in half, test each, recurse on failing half until minimal.
Use when: Large input crashes, many files break tests, config changes fail.
Scientific Debugging (TRAFFIC)
Track → Reproduce → Automate → Find origins → Focus → Isolate → Correct
Work backward: Failure → Propagation → Infection → Defect.
Version Control Debugging
Git Bisect
Binary search through commit history:
Usage:
git bisect start
git bisect bad
git bisect good abc1234
git bisect good
git bisect reset
Mark current as bad, known-good commit, then test each checkout (good/bad) until culprit found.
Automated:
git bisect run ./test.sh
Exit codes: 0 = good, 1-127 = bad, 125 = skip
Complexity: O(log n) - tests ~7 commits for 100 commit range
When to use:
- Regression appeared, unknown when
- Automated test can detect the bug
- Need to find exact commit that broke something
Coverage-Based Localization
Spectrum-Based Fault Localization (SBFL)
Use test coverage data to locate bugs:
Concept: Statements executed by failing tests but not passing tests are more suspicious.
Ochiai Formula (most effective):
suspiciousness(s) = failed(s) / sqrt(total_failed * (failed(s) + passed(s)))
Practical application:
- Run test suite with coverage
- Note which tests fail
- Rank statements by how often they appear in failing vs passing tests
- Inspect highest-ranked statements first
For AI agents: When multiple tests fail, identify code paths common to failures but not successes.
Before You Finish
When debugging or validating:
- Use Strong Inference: devise multiple hypotheses before testing
- Apply 5 Whys to find root cause, not symptoms
- Use Git Bisect for regressions (binary search ~7 commits for 100-commit range)
- Run tests with coverage; inspect code paths common to failures
- Bugfix? Trace to the real branch and reproduce real input BEFORE fixing
Claude Code Plugin Integration
When pr-review-toolkit is available:
- silent-failure-hunter agent: Detects silent failures, inadequate error handling
- Analyzes catch blocks, fallback behavior, missing logging
- Trigger: "Check for silent failures" or use Task tool
Ralph Loop Integration
Debugging = Ralph iteration: hypothesis → test → eliminate → iterate until <promise>ROOT CAUSE FOUND</promise>.
See @smith-ralph/SKILL.md for full patterns.
Related
- @smith-guidance/SKILL.md - Anti-sycophancy, HHH framework, exploration workflow
@smith-analysis/SKILL.md- Reasoning patterns, problem decomposition@smith-clarity/SKILL.md- Cognitive guards, logic fallacies@smith-tests/SKILL.md- Reproduce-first; never mock the branch under test@smith-subagents/SKILL.md- Audit a delegated diff's execution path