flaky-test-triage

$npx mdskill add Significant-Gravitas/skills-catalog/flaky-test-triage

Use this when a test passes and fails on the same code.

SKILL.md

.github/skills/flaky-test-triageView on GitHub ↗
---
name: "flaky-test-triage"
description: "Triage an intermittent test from run history, environment, and failure pattern before naming a cause."
triggers: ["flaky test", "intermittent failure", "test fails sometimes", "CI flakes", "unstable test"]
version: "1"
---

# Flaky test triage

Use this when a test passes and fails on the same code.

## Collect run history

Gather the last 30 to 100 runs where possible: pass or fail, job, runner,
branch, commit, duration, and the failure message. Say how many runs you
actually saw and over what dates.

## Find the pattern

Group the failures by job, runner, time of day, test order, duration, and
error. Look for common causes: timing and waits, shared state between tests,
test order, real clocks, network calls, random data, and resource limits. Name
each candidate with the evidence for and against it, and how sure you are.
Say when there are too few failures to tell.

## Next check

Propose the smallest check that would confirm or rule out the top candidate,
such as a rerun loop on one job, a fixed seed, or running the test alone.
Then list who owns the test and the options: fix, quarantine with an issue
and a date, or leave it and watch.

Never call a test flaky, or a cause found, without run history to show it.
Do not skip, disable, retry-wrap, or delete a test yourself; the owner decides.

More from Significant-Gravitas/skills-catalog

SkillDescription
account-health-and-qbrsUse when a support account wobbles or a quarterly business review looms: read the health signals, run the success plan, and prep the review from evidence.
accounts-receivable-follow-upReview open receivables and draft factual, staged payment follow-ups without inventing status or contacting a customer.
ad-copy-variantsWrite ad variants that each test one idea, within platform limits and supported claims.
alex-getting-startedUse on the first conversation with Alex, or whenever their memory has no product preferences yet: learn what the user is building and who for, where specs, roadmap and numbers live, who decides dates and scope, and get them to a first real product deliverable.
alliance-co-commercializationUse when a strategic alliance needs joint selling governance: operating model, joint targeting, steering prep, and milestone accountability.
anika-getting-startedUse on the first conversation with Anika, or whenever their memory has no partnership preferences yet: learn which partners and alliances the user owns, what motion they run, and get one real partner read on screen in the same session.
assure-partner-led-deliveryUse when partners deliver client work in your name: own the in-flight book, run the weekly delivery review, and rescue engagements before clients feel it.
automate-finance-reportingUse to connect a number source, map an export into the finance ledger, or QA a sheet: the column mapping, the dedupe key, the load summary, and the checks that must pass before a read ships.
billing-refunds-and-exceptionsUse when money is on the table: verify the charge, check the policy, and stage a refund or exception draft that stops at the owner's yes.
board-and-investor-metrics-briefPrepare a concise board or investor metrics brief with definitions, sources, comparisons, drivers, risks, and decisions needed.