The Homelab Postmortem

Real, dated postmortems from running a homelab — what broke, how it was diagnosed, and the exact fix. Plus the verified scripts that came out of it.

Homelabs fail in specific, dateable ways. This is a running record of what broke, how it was actually diagnosed, and the exact fix — not generic best-practice advice.

Rootless `podman export` of a `--userns=keep-id` container shifts every file owner in the tarball, and exits 0

A rootless container created with --userns=keep-id stores its files under the intermediate namespace's IDs, and podman export tars that storage without mapping them back. On a Pi with podman 5.4.2, the user's home cam...

Ollama's OpenAI-compatible API replaces your model's temperature and top_p with 1.0. `ollama show` keeps printing the model's values.

A request to /v1/chat/completions that doesn't set temperature or top_p gets 1.0 for both, and that overrides the model's PARAMETER lines. qwen3:0.6b ships temperature 0.6 and top_p 0.95; llama-server ran it at 0.600/...

llama.cpp picks stop tokens partly by their spelling. On PLaMo the closing strikethrough tag is ordinary text, and output stops at it.

llama.cpp adds any vocab entry spelled like an end-of-sequence marker to its stop list and forces it to a control token, even when the model marks it NORMAL. In PLaMo-2 and PLaMo-3 the closing strikethrough tag is exa...

Ollama 0.34 logs your OLLAMA_GPU_OVERHEAD reservation as applied. The llama-server that places the layers never receives it.

Ollama 0.34 runs GGUF models through a bundled llama-server, and with the default num_gpu the layer placement is llama-server's --fit, which keeps its own 1 GiB margin. OLLAMA_GPU_OVERHEAD is read, printed, and subtra...

podman cp on a stopped container ignores the volume's subpath: it reads a different file, writes outside the container's view, and overwrote another container's config. Exit 0 every time.

A named volume mounted with subpath= shows the container one directory of the volume. While the container runs, podman cp respects that. When it is created or stopped, podman resolves the path against the volume's roo...

The Laya checkpoint still ships a 0.1 temperature. laya clamps it when it loads; an ONNX port that reads the file does not, and its 15-option answers come back 0.88 sure when they are wrong.

Laya's English checkpoint carries a softmax temperature of 0.1006 for choice questions with 11 or more options, which sharpens the logits about ten times. laya fixed it on 2026-09-21 by clamping temperatures to [0.5, ...

llama-server's `--cache-ram -1` is documented as "no limit". On a dense model it keeps less than the default; on a hybrid model it has no memory limit at all.

The prompt cache has two limits, bytes and tokens. -1 removes the byte limit, and the token limit, which a positive setting raises to fit the bytes, stays pinned at n_ctx. With 16 distinct prompts on a 4096-token cont...

One failed WPA handshake and a headless Pi on Trixie stays off WiFi until someone types a command — its NetworkManager has no retry for PSK, asks for a key nobody can give, and blocks the profile.

NetworkManager 1.52.1 on Raspberry Pi OS Trixie: fail one 4-way handshake on a profile that has worked for weeks, and 3.6 seconds later NM is 'asking for new key' — the only retry setting is documented as 802.1X-only ...

Send Podman's Docker-compatible API an update that only changes memory, and the container's restart policy is now 'no'. HTTP 200, no warning, and the container stays down after the next crash or reboot.

Podman 5.4.2 (Debian 13), compat API v1.41: POST /containers/{id}/update with any body that omits RestartPolicy — a memory limit, a CpuShares value, an empty {} — returns 200 {"Warnings":null} and quietly sets the res...

Ollama's gemma4 renderer never shows the model a tool parameter named type, description, required, properties or nullable — it stays in the required list, and the model makes a value up.

Same tool schema, temperature 0, only the parameter renamed. Called kind, the model returns the enum value urgent_A7 — a string that exists only in the parameter's definition. Called type, it returns "urgent": a plaus...

llama.cpp reads past your pos array on every embedding batch for an M-RoPE model — the header says n_tokens, the batch code reads four times that, and the logits change with the heap.

Same batch, same context, memory cleared, one thread: token ids in, bitwise identical every run; embedding vectors in, Qwen3.5-0.8B, five runs, five different logits and the argmax moves. It looked like the weight typ...

llama.cpp's batch API says pass pos as NULL and positions are tracked automatically. For M-RoPE embeddings it then reads past its own buffer, and nothing reports it.

An embeddings batch with pos = NULL on a Qwen2.5-VL-style model gets a position vector sized n_tokens, and the batch splitter reads n_tokens × 4 from it. AddressSanitizer on today's master: heap-buffer-overflow, READ ...

Ollama's Responses API accepts previous_response_id, returns 200 and "completed", and starts every turn from nothing.

Plant a secret word in turn 1, ask for it in turn 2 with previous_response_id, and the model guesses. The request costs exactly as many input tokens as a request with no history at all — 41 and 41 — while sending the ...

llama-server accepts the json_schema form its own README documents, returns 200, and generates as if you had sent no response_format at all.

Send {"type":"json_schema","schema":{...}} — the example on the server README — and the reply is byte-identical to sending no response_format. The parser's json_schema branch only ever reads json_schema.schema, the 20...

docker logs stops at the first NUL byte in the file, returns exit 0, and --since after a crash tells you nothing happened.

One record of a json-file log overwritten with NUL bytes — what an abrupt power loss leaves behind — and the full read returns 17 of 50 lines with exit 0 and an empty stderr. --tail works only if its window starts aft...

Raspberry Pi OS ships a meta-data file whose instance-id key cloud-init never reads. Every user-data you write after the first boot is parsed, merged, and then skipped.

The stock /boot/firmware/meta-data sets instance_id with an underscore. cloud-init reads instance-id with a hyphen, so the value is never seen and the id falls back to the constant nocloud. One boot with the stock all...

vLLM accepts a LoRA adapter it has already decided not to apply anywhere, and answers every request with the base model instead.

Restrict a deployment with --lora-target-modules and load an adapter whose modules fall outside that list, and nothing objects. The adapter loads, the LoRA kernels compile, and every response is the unmodified base mo...

An Ollama library tag that pulls, runs, and streams fluent text with no working code in it. The same model at the same quantisation, converted by someone else, is fine.

qwen2.5-coder:3b-instruct-q3_K_M downloads cleanly, loads at normal speed and answers every prompt. It scored 0/3 on trivial coding tasks here. The q4_K_M sibling scored 3/3, and so did the official Qwen GGUF at the s...

llama-cli reports a commit hash that is not in llama.cpp. It belongs to whatever repository you unpacked the source inside.

A release tarball extracted anywhere inside an unrelated git work tree makes CMake stamp that repository's HEAD into the binary. Configure exits 0 with no warning, the dirty flag stays silent because untracked files a...

llama-cli prints the error and exits 0. The same binary exits 1 when the model is missing.

A missing --image or --audio path is printed as an error and then reported as success. It is not that the exit status is unreliable in general: the same binary propagates a missing model correctly. One path prints and...

The same ssh_import_id works in one cloud-init file and does nothing in the next. The difference is one letter.

A top-level ssh_import_id is read only for a user marked default, and a nested one only for a user that is not. Raspberry Pi's documented users: example and Imager's generated user: block sit on opposite sides of that...

iptables says your kernel needs upgrading. Upgrading the kernel is what broke it.

The Pi's 6.18 kernels no longer build the legacy ip_tables modules, but CONFIG_IP_NF_IPTABLES=m is still set — because the symbol that builds them is now called something else. The module directory still looks full, a...

tar backed up your Pi's WiFi config and exited 0. The archive holds one empty directory.

On a stock Raspberry Pi OS Trixie image, /etc/NetworkManager/system-connections is empty — and it is the directory every guide names. The connection you actually care about, WiFi password included, lives in /etc/netpl...

Podman publishes the same port and your firewall still holds. Even as root.

A commenter said rootless Podman avoids the Docker-bypasses-ufw problem because it has no root to write firewall rules with. The conclusion is right and the reason isn't: running Podman as root still leaves ufw in con...

UFW says the port is closed. Docker published it to your whole network anyway.

A stock Raspberry Pi with ufw enabled blocks an unlisted port exactly as advertised. Publish the same port from a container and it answers from anywhere on the LAN, while ufw status keeps reporting deny (incoming). Do...

Your Pi's 2 GB swap file isn't swap, and no swap command will tell you that

Trixie replaced dphys-swapfile with rpi-swap, which defaults to a zram+file hybrid. /var/swap still exists and still costs you 2 GB of disk, but it is zram's writeback target, not a swap area — so swapon, free and /pr...

Your Pi accepts every memory limit you set and enforces none of them

The memory cgroup controller is off by default on Raspberry Pi OS, so docker --memory, systemd MemoryMax and k3s limits are all accepted and silently ignored. The firmware injects cgroup_disable=memory ahead of your c...

A new NetworkManager property broke a decade of headless WiFi scripts

nmcli 1.52 added autoconnect-ports, which made the long-standing abbreviation autoconnect-p ambiguous. Provisioning scripts using it now abort with exit 2, and because nmcli connection modify is all-or-nothing, the SS...

sysctl -p says it worked. The next reboot disagrees.

On Raspberry Pi OS Trixie, /etc/sysctl.conf does not exist and is not read at boot — systemd-sysctl only reads the .d directories. The procps sysctl CLI still reads it, so sysctl -p applies your change and reports suc...

Raspberry Pi OS Trixie throws away your logs on reboot, and /var/log/journal exists anyway to reassure you it doesn't

Stock Trixie ships a vendor drop-in setting Storage=volatile, so the journal lives in RAM and dies on reboot. The empty /var/log/journal directory sitting there makes it look like persistence is already on. And the fi...

Stripe retries a failed webhook for three days, and I found out by getting the same email three times

A webhook that failed once during setup kept being retried in the background. Once the underlying bug was fixed, every one of those retries succeeded — and every one of them sent a real customer the same email again.

Undervoltage doesn't look like a power problem. It looks like bad Wi-Fi.

A Raspberry Pi 4B refused to hold a Wi-Fi connection, wouldn't finish a login, and behaved like it had a driver fault. The actual cause was a power supply that couldn't hold 5V — and the firmware had been saying so in...

Your Discord bot is ignoring you, and the status output is lying about why

A self-hosted agent bot connected fine, posted fine, and silently discarded every message aimed at it. The status line blamed a Discord intent. The actual cause was an unset applicationId — and the logs said so all al...

The rpi-clone PARTUUID trap: why your Pi 4B silently refuses to boot from a cloned SSD

The unmaintained rpi-clone tries to fix the root=PARTUUID reference in the destination's cmdline.txt — but looks for it at the pre-Bookworm path, finds a placeholder file, and silently does nothing. Flip the boot orde...