ShutdownCheck on GitHub

REAL LOAD · REAL SIGTERM · NEVER A FALSE PASS

Every deploy sends SIGTERM.
This tells you what breaks.

ShutdownCheck terminates your service the way your orchestrator will — under real load — and tells you exactly which stage of shutdown you got wrong, and how to fix it in your framework.

terminal — quick start
$ go install github.com/Ashutosh-Panda2004/ShutdownCheck/cmd/shutdowncheck@latest
$ shutdowncheck demo # a real check, a real SIGTERM, a real report
$ shutdowncheck run --url http://localhost:8080/api/orders -- ./bin/my-server

MEASURED · THIS PAGE'S OWN EVIDENCE

shutdowncheck — measured output
GET /work   calibrated to 77 rps (12 in flight at the signal)
in flight at signal   12    0 ok   12 failed   (12 refused)
after the signal      58    0 ok   58 failed   (58 refused)
VERDICT: FAIL   score 15/100 (F)
✕ SC003 IN_FLIGHT_DROPPED — 12 of 12 destroyed
✕ SC006 NO_DEREGISTRATION_WINDOW — listener closed 20.07ms after SIGTERM
✕ SC007 READINESS_NOT_FLIPPED

A service doing "graceful shutdown" exactly the way the tutorials teach — and dropping every request in flight on every deploy.

THE PROBLEM

Every deploy is a kill signal

Every deploy, every autoscale-down, every node drain, every spot reclaim, every rolling restart does the same thing to your service:

SIGTERM → wait → SIGKILL

If your service mishandles the window in between, every one of those events silently drops production traffic.

It rarely shows up in testing, because the bug is a race: whether a request dies depends on whether it happened to be mid-flight at the microsecond the signal landed. Send the signal to an idle service and everything looks perfect — which is why this survives code review and staging.

Most teams find out from a customer complaint, years later — and never trace it back.

THE ADVICE ALMOST EVERYONE FOLLOWS — AND WHY IT IS WRONG

"On SIGTERM, immediately stop accepting new connections, finish in-flight requests, then exit."

Search for "graceful shutdown" in any language and you will find that. Under Kubernetes — and behind most load balancers — that is a bug.

THE RACE · WHAT HAPPENS WHEN A POD IS DELETED

kubelet → your process
SIGTERM lands immediately.
endpoints → routing tables
De-registration propagates eventually — hundreds of milliseconds, sometimes seconds later.
your ingress
Traffic still routed to you hits a closed socket → connection refused → your ingress turns it into 502s for real users.

Nothing orders these two paths. Nothing waits for the second to finish. A service that closes its listener the instant the signal arrives did graceful shutdown by the book — and still drops traffic on every single deploy.

the fix — the lame-duck window
1. SIGTERM arrives
2. readiness starts failing at once # de-registration begins
3. keep accepting + serving # cover the propagation delay
4. stop accepting
5. drain what is in flight
6. exit 0, inside the grace period

Fail readiness immediately, keep serving while de-registration propagates, then close the listener and drain.

This is the single most important idea in graceful shutdown, and it is why ShutdownCheck has profiles rather than one fixed definition of "correct": whether closing the listener immediately is a defect depends entirely on what sits in front of your service.

THE MODEL

The seven stages of correct termination

Between SIGTERM and SIGKILL there are exactly seven things a service must do, in order. Miss one and traffic dies.

S1

Signal received

The process must react to SIGTERM at all. The usual cause of silence is not carelessness — it is a shell: CMD ./app in a Dockerfile runs /bin/sh -c ./app, so PID 1 is the shell and never forwards the signal. Use the exec form, CMD ["./app"].

SC001 · SIGTERM_IGNORED

S2

Readiness flipped

The health endpoint starts failing immediately — this starts the clock on de-registration, and every second of delay is a second longer that traffic keeps arriving. Readiness must also latch: a check that flips back to healthy re-adds you to the load balancer, which is worse than never flipping.

SC007 · READINESS_NOT_FLIPPEDSC008 · READINESS_FLIP_SLOWSC016 · READINESS_FLAPPED

S3

The lame-duck window

Keep serving. This is the stage that contradicts the common advice and the one most services get wrong. The window must cover your infrastructure's propagation delay — seconds, not milliseconds. If you have never measured it, 5 to 15 seconds is a common starting range.

SC006 · NO_DEREGISTRATION_WINDOW

S4

Listener closed

Once the window has elapsed, stop accepting. New connections should be refused cleanly — a socket that is accepted and never answered is worse than a refusal, because the client waits for a timeout instead of failing over immediately.

SC005 · LISTENER_OPEN_AFTER_WINDOWSC015 · ACCEPT_WITHOUT_RESPONSE

S5

Connection close signalled

Keep-alive clients may be holding an idle connection, intending to reuse it. Tell them instead: Connection: close on responses during shutdown, or GOAWAY on HTTP/2. An RST rather than an orderly FIN is a reset most client libraries will not retry.

SC004 · ABRUPT_CONNECTION_RESETSC009 · KEEPALIVE_NOT_TERMINATED

S6

In-flight work drained

Every request already accepted must run to completion. The common failure is a drain that runs longer than the grace period — at which point SIGKILL arrives and the drain was pointless. Your drain deadline must be shorter than your orchestrator's.

SC003 · IN_FLIGHT_DROPPEDSC011 · EARLY_EXITSC014 · DRAIN_LATENCY_SPIKE

S7

Clean exit inside budget

Exit with status 0 before the grace period expires — and leave nothing behind. A child process that outlives its parent and keeps the listening socket open means the next deploy fails to bind.

SC002 · SIGKILL_REQUIREDSC010 · SHUTDOWN_BUDGET_EXCEEDEDSC012 · PORT_HELD_AFTER_EXITSC013 · NONZERO_EXIT_CODESC017 · POST_KILL_TRAFFIC_LOSS

Stage text adapted from docs/seven-stages.md and the signature catalogue.

THE EXPERIMENT

How a run works

ShutdownCheck does not review your code. It kills your service — the same way, under the same conditions — and watches what happens to real requests.

1SPAWN

Starts your process, container, or attaches to a PID.

2READINESS

Waits until the readiness endpoint reports healthy.

3CALIBRATE

Auto-tunes load toward the requested concurrency, so work is guaranteed in flight.

4PROVE

Verifies requests were genuinely in flight when the signal landed.

5SIGTERM

Sends the signal and records the full timeline: traffic, readiness, listener, process.

6VERDICT

Named rules over recorded evidence. Deterministic, explainable, no ML.

the honesty rule
# if the in-flight sample cannot be proven:
VERDICT: INCONCLUSIVE # exit code 2 — never a pass
# SC000 · INSUFFICIENT_INFLIGHT
The run did not retain enough trustworthy evidence.

If it cannot prove the in-flight sample, the result is INCONCLUSIVE — never a pass.

A verification tool that passes when it measured nothing is a liability. Trust is the only asset a tool like this has, so ShutdownCheck would rather tell you it learned nothing than let you ship on an empty experiment.

THE REPORT

Anatomy of a FAIL

This is what a broken shutdown looks like when it is measured instead of guessed. Toggle the anatomy to see what each part of the report is telling you.

report — measured
ShutdownCheck v0.1.0-alpha … profile=kubernetes target=./shutdowncheck __demo-server1
GET /work   calibrated to 77 rps (12 in flight at the signal)2
TIMELINE                                    T=0 is the signal
traffic ████████████████████████████████▓▓▓▓░░░░░░░░░░░░░░░░
ready   ─────────────────────────────────····················
listen  ─────────────────────────────────────────────────····
process ─────────────────────────────────────────────────╳    exit at +0s
3
REQUESTS
  before the signal    66   66 ok   0 failed
  in flight at signal  12    0 ok  12 failed  (12 refused)
  after the signal      58    0 ok  58 failed  (58 refused)4
VERDICT: FAIL   score 15/100 (F)5
  gate max_inflight_drop_pct: 100 exceeds the limit of 06
✕ SC003 IN_FLIGHT_DROPPED
12 of 12 in-flight request(s) were destroyed during shutdown.
✕ SC006 NO_DEREGISTRATION_WINDOW
The listener closed 20.07ms after the signal, before de-registration could propagate.
✕ SC007 READINESS_NOT_FLIPPED
The readiness endpoint kept reporting healthy for the entire shutdown.
! SC009 KEEPALIVE_NOT_TERMINATED
1 connection(s) were reused after the signal, but the server never asked clients to close them.
7
  Run `shutdowncheck explain SC003` for how to fix this.8
1

Provenance line. Version, the profile the run was judged under, and the exact target — every report says what it measured and how it judged it.

2

Calibration proof. The load was auto-tuned to 77 rps so that 12 requests were provably in flight at the signal. No in-flight sample, no verdict.

3

The timeline. Traffic, readiness, listener and process plotted against T=0 — the signal. You can see the listener die at the exact moment the signal lands.

4

The body count. Requests before, during, and after the signal — and what happened to each. 12 in flight, 12 destroyed, all refused.

5

The verdict. FAIL, with a deterministic 0–100 score derived from named rules over the evidence above. Nothing heuristic, nothing learned.

6

The gate. Which policy gate tripped and by how much — the number your CI step can fail on.

7

Named findings. Each defect is a catalogue ID with a count and a one-line mechanism. ✕ is an error; ! is informational under this profile.

8

The way out. Every finding has an explanation and a fix: shutdowncheck explain SC003.

THE CATALOGUE

Eighteen ways to die between SIGTERM and exit

Every diagnosis is a named signature with a stage, a mechanism, and framework-specific remediation. No "something went wrong" — the report tells you which thing.

SC000
INSUFFICIENT_INFLIGHT

The run did not retain enough trustworthy baseline and in-flight evidence. — PRE-CONDITION

SC001
SIGTERM_IGNORED

The process showed no reaction to the signal at all. S1

SC002
SIGKILL_REQUIRED

The process was still alive when the grace period expired. S7

SC003
IN_FLIGHT_DROPPED

Requests that were already being processed failed during shutdown. S6

SC004
ABRUPT_CONNECTION_RESET

Connections were destroyed with RST instead of being closed cleanly. S5

SC005
LISTENER_OPEN_AFTER_WINDOW

The listener kept accepting connections past the de-registration window. S4

SC006
NO_DEREGISTRATION_WINDOW

The listener closed almost immediately after the signal. S3

SC007
READINESS_NOT_FLIPPED

The readiness endpoint stayed healthy for the whole shutdown. S2

SC008
READINESS_FLIP_SLOW

Readiness took too long to start failing. S2

SC009
KEEPALIVE_NOT_TERMINATED

Responses after the signal did not ask clients to close the connection. S5

SC010
SHUTDOWN_BUDGET_EXCEEDED

Shutdown took longer than the declared budget. S7

SC011
EARLY_EXIT

The process exited while requests were still in flight. S6

SC012
PORT_HELD_AFTER_EXIT

The port was still accepting connections after the main process exited. S7

SC013
NONZERO_EXIT_CODE

The process exited with a non-zero status after being asked to stop. S7

SC014
DRAIN_LATENCY_SPIKE

Latency rose sharply while draining. S6

SC015
ACCEPT_WITHOUT_RESPONSE

Connections were accepted after the signal but never answered. S4

SC016
READINESS_FLAPPED

Readiness recovered after starting to fail. S2

SC017
POST_KILL_TRAFFIC_LOSS

Requests were still in flight when SIGKILL landed. S7

INTERACTIVE · MEASUREMENT SEPARATED FROM INTERPRETATION

The same recording, judged twice

A run is recorded once and can be re-judged later without repeating it. Whether closing the listener immediately is a defect depends on what sits in front of your service — so pick a profile and watch the same evidence reach different conclusions.

$ shutdowncheck analyze run.ndjson --profile kubernetes
VERDICT: FAIL   score 18/100 (F)
✕ SC003 IN_FLIGHT_DROPPED — 11 of 12 destroyed
✕ SC004 ABRUPT_CONNECTION_RESET
✕ SC006 NO_DEREGISTRATION_WINDOW — listener closed 24.02ms after SIGTERM
✕ SC007 READINESS_NOT_FLIPPED
Behind a load balancer this is 502s on every deploy. Correct verdict: fix it.
$ shutdowncheck analyze run.ndjson --profile standalone
VERDICT: FAIL   score 18/100 (F)
✕ SC003 IN_FLIGHT_DROPPED — 11 of 12 destroyed
✕ SC004 ABRUPT_CONNECTION_RESET
! SC006 NO_DEREGISTRATION_WINDOW — demoted: nothing routes to you
! SC007 READINESS_NOT_FLIPPED — demoted: no load balancer to inform
Nothing sits in front of this service, so closing immediately is fine — but dropping in-flight work is broken under every profile.

ALL PROFILES

kubernetes — lame-duck required: flip readiness, keep serving, then drain.

lame-duck — as above, for non-Kubernetes load balancers.

standalone — nothing routes to you; closing immediately is fine.

 

strict — stop accepting immediately; the requirement, not the defect.

docker — docker stop semantics, 10s grace by default.

auto — the default; infers from the target.

Profile semantics from docs/seven-stages.md.

INTERACTIVE · EVERY FINDING HAS A FIX

shutdowncheck explain

A verdict tells you what broke. explain tells you why it matters and how to fix it in your framework — Go, ASP.NET Core, Node.js, Python, Java. Pick a signature:

$ shutdowncheck explain SC006

DESIGN COMMITMENTS

A verification tool is only as good as its honesty

Never a false PASS.

If the tool cannot prove correct behaviour, it reports INCONCLUSIVE. Trust is the only asset a verification tool has.

Black box.

No source access, no library import, no agent, no sidecar. The protocol is stack-independent — the same defects are proven across Go, Node.js and Python conformance fixtures.

Deterministic and explainable.

Every verdict traces to a named rule over recorded evidence. No heuristics, no ML, no AI.

Single static binary.

No runtime, no daemon, no cluster install, no account. One binary, one command.

No telemetry. Ever.

Your shutdown behaviour is your business. Nothing leaves the machine — there is not even a flag to turn reporting on.

Reproducible by design.

A run can be recorded and re-judged later without repeating it — useful when deciding whether behaviour that is fine standalone would survive behind a load balancer.

SEE IT IN 90 SECONDS

shutdowncheck demo

Runs a real check against a deliberately broken service — the one that closes its listener the instant SIGTERM arrives, exactly like the tutorials teach. Nothing is simulated: it is a real process, receiving a real signal, measured by the same code path as any other run.

No part of this tool ever prints a report it did not measure.

the 90-second script
$ shutdowncheck demo
VERDICT: FAIL · SC006 + SC007 # the tutorial bug, measured
$ shutdowncheck explain SC006
why it matters + the fix, per framework
$ # fix the service: fail readiness, keep serving, then close
$ shutdowncheck run --url … -- ./bin/my-server --port 8080
VERDICT: PASS # zero dropped requests

Fail → explain → fix → pass: the full arc from the project's docs/demo-script.md. The FAIL and PASS outputs in that script are the measured ones.

TARGETS

Whatever you ship, point it at the thing

01

Process

Attach to a running service by PID: --pid 1234. The tool signals the process it did not start.

02

Managed command

ShutdownCheck spawns it, waits for readiness, then kills it: shutdowncheck run --url … -- ./bin/my-server --port 8080.

03

Docker

Target a running container by name; the probe URL is filled in from its published ports: --docker my-api --url /api/orders.

04

Kubernetes

Still to come

Until then, use the kubernetes profile against a process target — the judgement is ready before the targeting is.

On Windows, use the Docker target: the platform has no SIGTERM, and simulating one would produce a verdict about a signal that was never delivered.

IN CI

A merge gate for deploys that 502

The GitHub Action runs the check on every pull request. The step fails when a defect is found, writes the report to the job summary, and uploads the JSON — with outputs for verdict, score and exit code so your workflow can decide what "fail" means.

Still working through what it finds? Run in report-only mode while you fix things — gating without blocking.

.github/workflows/shutdown.yml
- uses: shutdowncheck/shutdowncheck@v1
  with:
    version: v1.0.0
    url: http://localhost:8080/api/orders
    readiness-url: http://localhost:8080/readyz
    command: '["./bin/my-server", "--port", "8080"]'
    profile: kubernetes

INSTALL

One binary. Sixty seconds to your first verdict.

A note on honesty: the project is alpha, and several install methods below describe a release that has not been tagged yet. The path that works today is go install — everything else is ready for the first tag.

MethodCommandStatus
Gogo install github.com/Ashutosh-Panda2004/ShutdownCheck/cmd/shutdowncheck@latest
Scriptcurl -fsSL https://raw.githubusercontent.com/…/install.sh | sh
Homebrewbrew install shutdowncheck/tap/shutdowncheck
Scoopscoop install shutdowncheck
Dockerdocker run --rm ghcr.io/shutdowncheck/shutdowncheck:latest version
BinariesReleases page — linux, macOS, Windows · amd64 + arm64
verifying a download
# every release is cosign-signed (keyless), ships an SBOM, and carries SLSA provenance
$ cosign verify-blob checksums.txt --certificate checksums.txt.pem \
  --signature checksums.txt.sig --certificate-oidc-issuer https://token.actions.githubusercontent.com

COMMAND REFERENCE

Six commands. No daemon.

run — terminate a target under load and report what broke

shutdowncheck run --url http://localhost:8080/api/orders --readiness-url http://localhost:8080/readyz --profile kubernetes --grace-period 30s --trials 3 -- ./bin/my-server --port 8080

--url · endpoint to load (required)
--readiness-url · readiness endpoint (strongly recommended)
--profile · auto, standalone, strict, lame-duck, kubernetes, docker
--grace-period · grace before SIGKILL (default depends on target)
--signal · TERM, INT or QUIT
--trials · repeat; the worst result across trials is the verdict
--ensure-in-flight · requests to hold in flight at the signal
--slow-url · deliberately slow endpoint to guarantee in-flight work
--rps / --max-rps · fixed rate (disables calibration) / ceiling
--prestop-sleep · simulate a Kubernetes preStop hook
--docker · target a running container by name or id
--pid · attach to an existing process by id
--fail-on / --ignore · promote or demote signatures
--min-score / --max-inflight-drop-pct / --max-shutdown-time · gates
--format · human, json, junit, markdown, ndjson, html
--output / --badge · write the report / an SVG score badge

analyze — re-judge a recorded run, optionally under a different profile

shutdowncheck run … --format ndjson --output run.ndjson
shutdowncheck analyze run.ndjson --profile strict # same evidence, different deployment model

demo — run a real check against a deliberately broken service

shutdowncheck demo [--profile kubernetes] [--format human] [--no-color]

explain / validate / version

shutdowncheck explain SC006 — the failure and its fix, per framework.
shutdowncheck validate --config shutdowncheck.yaml — check a config without running anything.
shutdowncheck version — build information.

Exit codes — a pipeline can tell "you configured me wrong" from "your service has a bug"

0 · pass — shutdown behaviour verified as correct
1 · fail — one or more defects were detected
2 · inconclusive — the run did not prove anything; not a pass
3 · usage — the tool was invoked incorrectly
4 · target — the target could not be started, reached or signalled
5 · internal — a bug in shutdowncheck

POSITIONING

What it is not

How is this different from Gremlin, Litmus, or Chaos Mesh?

Those are chaos engineering platforms: they inject faults to test resilience, and they need agents, dashboards, often a paid tier. ShutdownCheck is not chaos engineering — it is a pre-deployment verification tool, like a linter for shutdown behaviour. It answers one question: "Will this service drop traffic when the orchestrator terminates it?" — in CI, before you merge.

What about service-mesh draining — Istio, Linkerd?

Meshes handle traffic shifting at the proxy layer, but they do not fix application-level shutdown bugs. If your app closes its listener before de-registration propagates, never fails readiness, or exits with requests in flight, the mesh sees refused connections and broken pipes. ShutdownCheck verifies the application does its part.

Why not just use docker stop and see what happens?

You can, but you will not learn why it broke. docker stop gives you a pass/fail. ShutdownCheck gives you which of the seven stages failed, how many requests were affected and what happened to them — refused, reset, timed out — and the exact fix via shutdowncheck explain. The difference between "it broke" and "here is the line to change."

Does it work with my framework?

If your framework runs on Linux and handles SIGTERM, yes. Verified with Go (net/http), Node.js and Python conformance fixtures. The seven stages are framework-agnostic — they describe the contract between your app and the orchestrator, not specific APIs.

Is it production-ready?

It is alpha. The core analysis is solid — verified against 14 conformance scenarios in Go — but Kubernetes targeting is not yet implemented and the HTML report is new. Use it to learn, not as a merge gate you depend on. Yet.

ORIGIN

Built in a weekend, out of production 502s

Built in a weekend by Ashutosh Panda because every "graceful shutdown" tutorial was wrong about Kubernetes. The lame-duck pattern is not documented in most frameworks, and the only way to learn it was to cause 502s in production.

This tool exists so you do not have to.