Agent Skills: calibrating-thresholds-and-baselines
Check that a score-threshold-flag pipeline actually separates anything before you trust or tune it — reference constants that are hand-authored but formatted as measurements, a computed feature whose units do not match the constant it is compared against, per-class fire rates that expose a flag firing on nearly every item of both classes, rate features confounded by item length, confidence scores where one feature family contributes once per sub-feature, threshold sweeps hardcoded to one comparison direction and one statistic, and separating features that carry the decision from features that only populate a warning list. Use when writing or reviewing an anomaly detector, a quality or risk score, a fraud/abuse heuristic, a lint-style flag set, or any pipeline that z-scores features against stored baseline constants and emits flags.
Install this agent skill to your local
Skill Files
Browse the full folder contents for calibrating-thresholds-and-baselines.
Loading file tree…
Select a file to preview its contents.