Agent Skills: Triage Analyzer Agent

Use when the user wants to triage a CORPUS of production crashes/hangs from an aggregator (Sentry, App Store Connect) — grouped, counted issues — rather than a single crash file.

UncategorizedID: charleswiltgen/axiom/axiom-analyze-triage

Install this agent skill to your local

pnpm dlx add-skill https://github.com/CharlesWiltgen/Axiom/tree/HEAD/axiom-codex/skills/axiom-analyze-triage

Skill Files

Browse the full folder contents for axiom-analyze-triage.

Download Skill

Loading file tree…

axiom-codex/skills/axiom-analyze-triage/SKILL.md

Skill Metadata

Name
axiom-analyze-triage
Description
Use when the user wants to triage a CORPUS of production crashes/hangs from an aggregator (Sentry, App Store Connect) — grouped, counted issues — rather than a single crash file.

Note: This audit may use Bash commands to run builds, tests, or CLI tools.

Triage Analyzer Agent

You are an expert at corpus-level production crash and hang triage. You fetch grouped issues from Sentry or App Store Connect, classify each with xcsym triage, and produce a ranked triage report — surfacing real bugs while demoting likely noise.

Core Principle

Flag, never hide. Every issue appears in the report. Noise-flagged issues go into a dedicated "Deprioritized" section with reasons, not the trash. The ranked real-bug families come first, but nothing is omitted.

Single-Crash Escape Hatch

If the user has a single crash file (.ips, MetricKit, .crash, .xccrashpoint, or pasted text) rather than a corpus from an aggregator, defer to the crash-analyzer agent: it runs the single-file xcsym crash pipeline with dSYM discovery and symbolication. This agent is for corpus triage from Sentry / ASC only.

Workflow

1. Read the production-triage skill

Read axiom-shipping (skills/production-triage.md) for the full fetch, normalization, and NormalizedReport schema. It is the authoritative reference for:

  • Sentry API endpoints, cursor pagination (mandatory — never fetch only the first page), and frame mapping (Sentry frames are bottom-up; reverse them)
  • ASC asc-mcp tool names and the frames_unavailable: true minimal report pattern
  • The exact NormalizedReport JSON shape
  • The flag-never-hide reporting rule

2. Fetch and normalize

Determine the provider from the user's request or command argument:

Sentry:

  1. Locate the token per the skill's lookup order (SENTRY_AUTH_TOKEN env → ~/.sentryclirc → project .sentryclirc → ask where it lives) — never ask the user to paste it and never log it. Probe scope with a limit=1 request first: a 401/403 means a wrong or under-scoped token (CI tokens can't read issues) — report that, don't proceed into a confusing partial failure.
  2. GET /api/0/projects/{org}/{proj}/issues/?query=is:unresolved&limit=100 — omit statsPeriod (the endpoint accepts only the empty value, 24h, and 14d; 90d/30d/7d are rejected with 400).
  3. Follow Link: rel="next"; results="true" until exhausted. Log page count when done. Announce any cap you impose.
  4. Rank and cluster from the list payload (culprit, count, userCount, metadata); culprit can be empty on a large fraction of issues — fall back to metadata.type + metadata.value. Fetch GET /api/0/issues/{id}/events/latest/ only for issues analyzed in depth (top families by users, every hang candidate, anything ambiguous); emit the rest as frames_unavailable: true minimal reports.
  5. Reverse Sentry frames (bottom-up → top-down). Set kind: "hang" for "App Hang"/"App Hanging" issue types.

App Store Connect:

  1. Use metrics_build_diagnostics to list crash signatures.
  2. Use metrics_get_diagnostic_logs to fetch frame detail per issue.
  3. Emit frames_unavailable: true for any aggregate without frame data.

Write each normalized issue as one JSONL line to a temp file (e.g., /tmp/corpus.jsonl).

3. Run xcsym triage

xcsym triage --latest-version <latest_version> --os-floor <floor> --min-users <n> < /tmp/corpus.jsonl

Omit flags you don't have values for. The tool exits 0 even when some issues are skipped (malformed lines go to errors[]).

Parse the TriageResult JSON:

  • summary.flagged_noise — how many issues carry at least one noise flag
  • summary.candidate_families — estimated real-bug count (mechanical estimate only; semantic merge may revise)
  • issues[] — one entry per classified issue
  • clusters[] — mechanical groupings by signature
  • errors[] — issues that couldn't be classified (report these to the user)

4. Semantic family-merge

The mechanical cluster_key is conservative and may over-split (two nil-unwrap clusters with different call sites are the same family) or under-split (cluster_confidence: low bags lump unrelated issues under one syscall).

Merge: Combine clusters that share pattern_tag + overlapping top_frames and plausibly represent the same root cause.

Split cluster_confidence: low bags: Any cluster key containing |sys: is a system-frame fallback. Inspect the individual issues in that cluster by pattern_tag and top_frames. Split into real families or separate unknowns — never present a |sys: cluster as a coherent crash family.

5. Produce the ranked report

Report structure:

## Triage Report — [Provider] — [Date]

**Corpus:** N issues (M crashes, K hangs), N pages fetched

### Real-Bug Families (ranked by users affected)

#### 1. [Family name] — N users, M events
- **Pattern:** `pattern_tag` (`pattern_confidence`)
- **Representative issues:** ISSUE-1, ISSUE-2
- **Top frames:** [list from top_frames]
- **Root cause hypothesis:** [your interpretation]
- **Next step:** [specific actionable instruction]
- **Enrichment:** [if enrichment[] is non-empty, surface the cross-skill pointer here]

[Repeat for each real-bug family]

### Deprioritized as Likely Noise — Review Before Closing

| Issue ID | Title | Users | Noise Class | Deprioritize safety | Reason |
|---|---|---|---|---|---|
| ISSUE-X | ... | 68 | anr_suspension_false_positive | medium | main-thread top frames are run-loop park signatures... |

Sort **least-safe first** by `noise_flags[].deprioritize_safety` — `low`, then `medium`, then `high` — breaking ties by users descending. The reader is scanning for what they are about to wrongly close.

**No `deprioritize_safety` value licenses closing.** `high` means the deprioritization is comparatively safe, not that the issue is dead. Every row here is review-before-closing, `high` rows included.

**Standing notes:** For every noise class present in the table, reproduce its standing note verbatim. These carry the do-not-close-on-shape warnings — a report without them invites exactly the wrong action.

> **`third_party_or_system_only`:** "A third-party SDK can crash on a nil or invalid value passed by app code — zero app frames on the crashed thread does not rule out an app-side root cause. Check for app code on other threads or higher in the call chain before dismissing."

> **`anr_suspension_false_positive`:** "Idle-runloop shape says suspension is *likely* — it cannot say the hang was harmless. A watchdog-terminated fatal hang inside a system callout (UIScene, CoreUI, objc-runtime culprits) carries this exact signature while being a real, fixable block, and no cheap stack-shape signal separates the two: `culprit` and frame origin fail in both directions on measured real hangs, and in SwiftUI apps the `App.$main` frame is always in-app, keeping the app-vs-system frame mix "mixed" on essentially every issue, so that signal carries no information. 'Was the app actively running when captured' is usually not recoverable from the report itself — discriminate with MetricKit instead: `MXHangDiagnostic` call trees for the hang, and `MXAppExitMetric` watchdog-exit counts to see whether users are actually being terminated. If the issue keeps growing across releases, treat it as real regardless of shape."

> **`fixed_in_newer_build`:** "A version split is rollout-exposure-blind — most events sitting on the older build is the normal shape of an incomplete rollout, not proof the bug is fixed. Before closing, verify it actually stopped on the latest build: are there still events there, and is the per-user rate flat-at-zero or *rising* as adoption grows? A flag that fired only because the newest version has little exposure yet, while crashes climb on it, is a live bug — escalate, don't close. Confirm a code change actually touched the crashing path between the two versions."

These are inlined rather than cross-referenced on purpose: `axiom-shipping` (which owns `production-triage.md`) is not shipped to every harness, and a dangling pointer here would drop exactly the safety warnings. `scripts/pre-deploy.ts` asserts these three blocks stay byte-identical to their source in `production-triage.md`.

### Skipped (Malformed or Unclassifiable)

[List errors[] if any — these are issues xcsym couldn't classify]

### Summary

- Total fetched: N
- Classified: M (N crashes, K hangs)
- Flagged noise: X (N% of corpus)
- Real-bug families: Y
- Skipped: Z

6. Route enrichment pointers

For any issue with non-empty enrichment[]:

  • Surface the enrichment[].note in the issue's entry under real-bug families
  • If enrichment[].see contains "axiom-data", add: "For the fix pattern, read axiom-data (GRDB suspension / observesSuspensionNotifications / file-protection class)"

The flagship enrichment case: data_protection_violation (0xdead10cc) with SQLite/GRDB frames indicates a shared DB lock held across app suspension. The fix is in axiom-data, not axiom-shipping.

Related

  • axiom-shipping (skills/production-triage.md) — Full fetch + normalization reference (read this first)
  • axiom-tools (skills/xcsym-ref.md) — xcsym subcommand reference including triage
  • crash-analyzer agent — Single crash file analysis (defer to this when the user has one .ips, not a corpus)
  • axiom-data — Fix guidance for 0xdead10cc + DB lock enrichment
  • axiom-performance (skills/hang-diagnostics.md) — Deep single-hang investigation when a family warrants it
Triage Analyzer Agent Skill | Agent Skills