measured — produced by actually running ShutdownCheck in this environment: a real process, a real SIGTERM, a real report.
from docs — stated in the project's own README, spec, or docs. Tap the chip for the exact source.
illustrative — a diagram or simplification to make the mechanism visible. Never a measurement.
This tool's first design commitment is "never a false PASS". Its marketing gets the same treatment.
REAL LOAD · REAL SIGTERM · NEVER A FALSE PASS
ShutdownCheck terminates your service the way your orchestrator will — under real load — and tells you exactly which stage of shutdown you got wrong, and how to fix it in your framework.
MEASURED · THIS PAGE'S OWN EVIDENCE
A service doing "graceful shutdown" exactly the way the tutorials teach — and dropping every request in flight on every deploy.
THE PROBLEM
Every deploy, every autoscale-down, every node drain, every spot reclaim, every rolling restart does the same thing to your service:
SIGTERM → wait → SIGKILL
If your service mishandles the window in between, every one of those events silently drops production traffic.
It rarely shows up in testing, because the bug is a race: whether a request dies depends on whether it happened to be mid-flight at the microsecond the signal landed. Send the signal to an idle service and everything looks perfect — which is why this survives code review and staging.
Most teams find out from a customer complaint, years later — and never trace it back.
THE ADVICE ALMOST EVERYONE FOLLOWS — AND WHY IT IS WRONG
Search for "graceful shutdown" in any language and you will find that. Under Kubernetes — and behind most load balancers — that is a bug.
THE RACE · WHAT HAPPENS WHEN A POD IS DELETED
Nothing orders these two paths. Nothing waits for the second to finish. A service that closes its listener the instant the signal arrives did graceful shutdown by the book — and still drops traffic on every single deploy.
Fail readiness immediately, keep serving while de-registration propagates, then close the listener and drain.
This is the single most important idea in graceful shutdown, and it is why ShutdownCheck has profiles rather than one fixed definition of "correct": whether closing the listener immediately is a defect depends entirely on what sits in front of your service.
THE MODEL
Between SIGTERM and SIGKILL there are exactly seven things a service must do, in order. Miss one and traffic dies.
S1
The process must react to SIGTERM at all. The usual cause of silence is not carelessness — it is a shell: CMD ./app in a Dockerfile runs /bin/sh -c ./app, so PID 1 is the shell and never forwards the signal. Use the exec form, CMD ["./app"].
S2
The health endpoint starts failing immediately — this starts the clock on de-registration, and every second of delay is a second longer that traffic keeps arriving. Readiness must also latch: a check that flips back to healthy re-adds you to the load balancer, which is worse than never flipping.
S3
Keep serving. This is the stage that contradicts the common advice and the one most services get wrong. The window must cover your infrastructure's propagation delay — seconds, not milliseconds. If you have never measured it, 5 to 15 seconds is a common starting range.
S4
Once the window has elapsed, stop accepting. New connections should be refused cleanly — a socket that is accepted and never answered is worse than a refusal, because the client waits for a timeout instead of failing over immediately.
S5
Keep-alive clients may be holding an idle connection, intending to reuse it. Tell them instead: Connection: close on responses during shutdown, or GOAWAY on HTTP/2. An RST rather than an orderly FIN is a reset most client libraries will not retry.
S6
Every request already accepted must run to completion. The common failure is a drain that runs longer than the grace period — at which point SIGKILL arrives and the drain was pointless. Your drain deadline must be shorter than your orchestrator's.
S7
Exit with status 0 before the grace period expires — and leave nothing behind. A child process that outlives its parent and keeps the listening socket open means the next deploy fails to bind.
Stage text adapted from docs/seven-stages.md and the signature catalogue.
THE EXPERIMENT
ShutdownCheck does not review your code. It kills your service — the same way, under the same conditions — and watches what happens to real requests.
Starts your process, container, or attaches to a PID.
Waits until the readiness endpoint reports healthy.
Auto-tunes load toward the requested concurrency, so work is guaranteed in flight.
Verifies requests were genuinely in flight when the signal landed.
Sends the signal and records the full timeline: traffic, readiness, listener, process.
Named rules over recorded evidence. Deterministic, explainable, no ML.
If it cannot prove the in-flight sample, the result is INCONCLUSIVE — never a pass.
A verification tool that passes when it measured nothing is a liability. Trust is the only asset a tool like this has, so ShutdownCheck would rather tell you it learned nothing than let you ship on an empty experiment.
THE REPORT
This is what a broken shutdown looks like when it is measured instead of guessed. Toggle the anatomy to see what each part of the report is telling you.
TIMELINE T=0 is the signal traffic ████████████████████████████████▓▓▓▓░░░░░░░░░░░░░░░░ ready ─────────────────────────────────···················· listen ─────────────────────────────────────────────────···· process ─────────────────────────────────────────────────╳ exit at +0s3
Provenance line. Version, the profile the run was judged under, and the exact target — every report says what it measured and how it judged it.
Calibration proof. The load was auto-tuned to 77 rps so that 12 requests were provably in flight at the signal. No in-flight sample, no verdict.
The timeline. Traffic, readiness, listener and process plotted against T=0 — the signal. You can see the listener die at the exact moment the signal lands.
The body count. Requests before, during, and after the signal — and what happened to each. 12 in flight, 12 destroyed, all refused.
The verdict. FAIL, with a deterministic 0–100 score derived from named rules over the evidence above. Nothing heuristic, nothing learned.
The gate. Which policy gate tripped and by how much — the number your CI step can fail on.
Named findings. Each defect is a catalogue ID with a count and a one-line mechanism. ✕ is an error; ! is informational under this profile.
The way out. Every finding has an explanation and a fix: shutdowncheck explain SC003.
THE CATALOGUE
Every diagnosis is a named signature with a stage, a mechanism, and framework-specific remediation. No "something went wrong" — the report tells you which thing.
The run did not retain enough trustworthy baseline and in-flight evidence. — PRE-CONDITION
The process showed no reaction to the signal at all. S1
The process was still alive when the grace period expired. S7
Requests that were already being processed failed during shutdown. S6
Connections were destroyed with RST instead of being closed cleanly. S5
The listener kept accepting connections past the de-registration window. S4
The listener closed almost immediately after the signal. S3
The readiness endpoint stayed healthy for the whole shutdown. S2
Readiness took too long to start failing. S2
Responses after the signal did not ask clients to close the connection. S5
Shutdown took longer than the declared budget. S7
The process exited while requests were still in flight. S6
The port was still accepting connections after the main process exited. S7
The process exited with a non-zero status after being asked to stop. S7
Latency rose sharply while draining. S6
Connections were accepted after the signal but never answered. S4
Readiness recovered after starting to fail. S2
Requests were still in flight when SIGKILL landed. S7
INTERACTIVE · MEASUREMENT SEPARATED FROM INTERPRETATION
A run is recorded once and can be re-judged later without repeating it. Whether closing the listener immediately is a defect depends on what sits in front of your service — so pick a profile and watch the same evidence reach different conclusions.
ALL PROFILES
kubernetes — lame-duck required: flip readiness, keep serving, then drain.
lame-duck — as above, for non-Kubernetes load balancers.
standalone — nothing routes to you; closing immediately is fine.
strict — stop accepting immediately; the requirement, not the defect.
docker — docker stop semantics, 10s grace by default.
auto — the default; infers from the target.
Profile semantics from docs/seven-stages.md.
INTERACTIVE · EVERY FINDING HAS A FIX
A verdict tells you what broke. explain tells you why it matters and how to fix it in your framework — Go, ASP.NET Core, Node.js, Python, Java. Pick a signature:
DESIGN COMMITMENTS
If the tool cannot prove correct behaviour, it reports INCONCLUSIVE. Trust is the only asset a verification tool has.
No source access, no library import, no agent, no sidecar. The protocol is stack-independent — the same defects are proven across Go, Node.js and Python conformance fixtures.
Every verdict traces to a named rule over recorded evidence. No heuristics, no ML, no AI.
No runtime, no daemon, no cluster install, no account. One binary, one command.
Your shutdown behaviour is your business. Nothing leaves the machine — there is not even a flag to turn reporting on.
A run can be recorded and re-judged later without repeating it — useful when deciding whether behaviour that is fine standalone would survive behind a load balancer.
SEE IT IN 90 SECONDS
Runs a real check against a deliberately broken service — the one that closes its listener the instant SIGTERM arrives, exactly like the tutorials teach. Nothing is simulated: it is a real process, receiving a real signal, measured by the same code path as any other run.
No part of this tool ever prints a report it did not measure.
Fail → explain → fix → pass: the full arc from the project's docs/demo-script.md. The FAIL and PASS outputs in that script are the measured ones.
TARGETS
01
Attach to a running service by PID: --pid 1234. The tool signals the process it did not start.
02
ShutdownCheck spawns it, waits for readiness, then kills it: shutdowncheck run --url … -- ./bin/my-server --port 8080.
03
Target a running container by name; the probe URL is filled in from its published ports: --docker my-api --url /api/orders.
04
Still to come
Until then, use the kubernetes profile against a process target — the judgement is ready before the targeting is.
On Windows, use the Docker target: the platform has no SIGTERM, and simulating one would produce a verdict about a signal that was never delivered.
IN CI
The GitHub Action runs the check on every pull request. The step fails when a defect is found, writes the report to the job summary, and uploads the JSON — with outputs for verdict, score and exit code so your workflow can decide what "fail" means.
Still working through what it finds? Run in report-only mode while you fix things — gating without blocking.
INSTALL
A note on honesty: the project is alpha, and several install methods below describe a release that has not been tagged yet. The path that works today is go install — everything else is ready for the first tag.
| Method | Command | Status |
|---|---|---|
| Go | go install github.com/Ashutosh-Panda2004/ShutdownCheck/cmd/shutdowncheck@latest | |
| Script | curl -fsSL https://raw.githubusercontent.com/…/install.sh | sh | |
| Homebrew | brew install shutdowncheck/tap/shutdowncheck | |
| Scoop | scoop install shutdowncheck | |
| Docker | docker run --rm ghcr.io/shutdowncheck/shutdowncheck:latest version | |
| Binaries | Releases page — linux, macOS, Windows · amd64 + arm64 |
COMMAND REFERENCE
shutdowncheck run --url http://localhost:8080/api/orders --readiness-url http://localhost:8080/readyz --profile kubernetes --grace-period 30s --trials 3 -- ./bin/my-server --port 8080
--url · endpoint to load (required)--readiness-url · readiness endpoint (strongly recommended)--profile · auto, standalone, strict, lame-duck, kubernetes, docker--grace-period · grace before SIGKILL (default depends on target)--signal · TERM, INT or QUIT--trials · repeat; the worst result across trials is the verdict--ensure-in-flight · requests to hold in flight at the signal--slow-url · deliberately slow endpoint to guarantee in-flight work--rps / --max-rps · fixed rate (disables calibration) / ceiling--prestop-sleep · simulate a Kubernetes preStop hook--docker · target a running container by name or id--pid · attach to an existing process by id--fail-on / --ignore · promote or demote signatures--min-score / --max-inflight-drop-pct / --max-shutdown-time · gates--format · human, json, junit, markdown, ndjson, html--output / --badge · write the report / an SVG score badgeshutdowncheck run … --format ndjson --output run.ndjson
shutdowncheck analyze run.ndjson --profile strict # same evidence, different deployment model
shutdowncheck demo [--profile kubernetes] [--format human] [--no-color]
shutdowncheck explain SC006 — the failure and its fix, per framework.
shutdowncheck validate --config shutdowncheck.yaml — check a config without running anything.
shutdowncheck version — build information.
0 · pass — shutdown behaviour verified as correct1 · fail — one or more defects were detected2 · inconclusive — the run did not prove anything; not a pass3 · usage — the tool was invoked incorrectly4 · target — the target could not be started, reached or signalled5 · internal — a bug in shutdowncheckPOSITIONING
Those are chaos engineering platforms: they inject faults to test resilience, and they need agents, dashboards, often a paid tier. ShutdownCheck is not chaos engineering — it is a pre-deployment verification tool, like a linter for shutdown behaviour. It answers one question: "Will this service drop traffic when the orchestrator terminates it?" — in CI, before you merge.
Meshes handle traffic shifting at the proxy layer, but they do not fix application-level shutdown bugs. If your app closes its listener before de-registration propagates, never fails readiness, or exits with requests in flight, the mesh sees refused connections and broken pipes. ShutdownCheck verifies the application does its part.
You can, but you will not learn why it broke. docker stop gives you a pass/fail. ShutdownCheck gives you which of the seven stages failed, how many requests were affected and what happened to them — refused, reset, timed out — and the exact fix via shutdowncheck explain. The difference between "it broke" and "here is the line to change."
If your framework runs on Linux and handles SIGTERM, yes. Verified with Go (net/http), Node.js and Python conformance fixtures. The seven stages are framework-agnostic — they describe the contract between your app and the orchestrator, not specific APIs.
It is alpha. The core analysis is solid — verified against 14 conformance scenarios in Go — but Kubernetes targeting is not yet implemented and the HTML report is new. Use it to learn, not as a merge gate you depend on. Yet.
ORIGIN
Built in a weekend by Ashutosh Panda because every "graceful shutdown" tutorial was wrong about Kubernetes. The lame-duck pattern is not documented in most frameworks, and the only way to learn it was to cause 502s in production.
This tool exists so you do not have to.