The Homelab Postmortem

Real, dated postmortems from running a homelab — what broke, how it was diagnosed, and the exact fix. Plus the verified scripts that came out of it.

Toolkit

Every script here comes from a specific, dated incident documented on this site — diagnosed on real hardware, fixed, and then generalized into a script that’s safe to hand to a stranger: no silent writes, backups before anything destructive, dry-run modes where it matters.

What’s in it

Also relevant if you’re building your own delivery pipeline: Stripe retries a failed webhook for three days — the idempotency bug this toolkit’s own delivery Worker hit and fixed.

New scripts are added as new incidents happen — this is a living collection, not a one-time release.

Pick the checks you need — $5 each

Each pack answers one question and contains only the scripts for it. Buy the one that matches what you are about to do; you are not paying for checks that do not apply to your machine.

Exposure & firewall — $5

Is anything reachable from your network that your firewall says is blocked?

Two checks. Container ports that answer from the LAN while ufw reports them denied, and whether the iptables your tools call has any kernel support left.

Buy Exposure & firewall →

Backup & restore integrity — $5

Is your backup actually a backup, and will the clone boot?

Four checks. Where your network config really persists before you archive it, the stale PARTUUID that stops a successful-looking clone from booting, the podman cp that copies a different file out of a stopped container, and the podman export that writes every owner in the archive shifted.

Buy Backup & restore integrity →

Persistence across reboot — $5

Will the settings you just applied still be there after a reboot?

Five checks. sysctl, the journal, swap, Docker container logs after a power loss, and whether your Podman containers still have the restart policy you gave them — each a case where the command succeeds, the value reads back, and the next boot disagrees.

Buy Persistence across reboot →

Provisioning verification — $5

Did the machine actually get configured the way you told it to?

Four checks. The cloud-init key that is valid, schema-checked, and silently ignored — in both directions — whether your user-data will be acted on at all or parsed and skipped, the network profile a provisioning run really produced, and whether the WiFi you just set up will come back by itself after one failed handshake.

Buy Provisioning verification →

Local LLM artifacts — $5

Is the model, adapter, or build you just installed the thing it says it is?

Twelve checks. An Ollama quantisation that pulls and runs and returns no working code, a LoRA adapter accepted and applied to zero layers, a build that reports a commit hash from a repository it has never heard of, a response_format the server accepts and then ignores, a previous_response_id that is accepted and dropped, a tool parameter the gemma4 renderer never shows the model, an embedding batch whose positions the library reads past the end of, a prompt cache whose "no limit" setting keeps less than the default, and a decision model whose confidence is sharpened tenfold by a temperature its own library clamps at load time, a VRAM reservation Ollama prints in its log and never passes to the runner that places the layers, model sampling settings the OpenAI-compatible API quietly replaces with 1.0, and an end-of-generation token chosen by its spelling that stops a model mid-sentence with a normal finish.

Buy Local LLM artifacts →

Every script in every pack is also in the complete toolkit below, so there is no reason to buy both.

Get everything — $15

Every script in the toolkit, including the ones that are not in any pack. Buying the 5 packs separately is $25, so this is the cheaper route if you want more than two of them. One-time purchase: the download link is emailed to you immediately, and every script added later is part of the same purchase.

Buy the toolkit →