From d028fb4e1bad4939d800e626530f4964cb7a7460 Mon Sep 17 00:00:00 2001 From: ginnoir Date: Wed, 5 Aug 2026 13:05:31 -0500 Subject: [PATCH] docs: add Jellyfin migration implementation plan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 25 tasks across 7 phases, written to be executed in Cursor rather than Claude Code — every action is a literal shell command, file edit or named UI step, with no dependency on this repo's Claude Code skills or MCP servers. Phase 0 is a hard gate: the root disk is at 97% with /config on it, so nothing deploys until space is reclaimed and free-space alerting exists. Plex runs untouched through Task 22, and rollback stays a single command until the 30-day archive expires. Co-Authored-By: Claude Opus 5 --- .../plans/2026-08-05-jellyfin-migration.md | 1816 +++++++++++++++++ 1 file changed, 1816 insertions(+) create mode 100644 docs/superpowers/plans/2026-08-05-jellyfin-migration.md diff --git a/docs/superpowers/plans/2026-08-05-jellyfin-migration.md b/docs/superpowers/plans/2026-08-05-jellyfin-migration.md new file mode 100644 index 0000000..3795046 --- /dev/null +++ b/docs/superpowers/plans/2026-08-05-jellyfin-migration.md @@ -0,0 +1,1816 @@ +# Jellyfin Migration Implementation Plan + +> **For agentic workers:** This plan is executed in **Cursor**, not Claude Code. It deliberately +> does **not** use this repo's Claude Code skills (`homelab-apply`, `homelab-ssh`) or MCP servers +> (portainer, uptime-kuma, authentik). Every action is a literal shell command, file edit, or +> named UI step. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Replace host-native Plex on valhalla with containerised Jellyfin, preserving curated +artwork, collections and per-user watch history, then decommission Plex. + +**Architecture:** A new Portainer-managed stack `stacks/streaming/` (jellyfin, jellystat + +postgres, wizarr, jellyplex-watched) on the shared `edge` network behind Caddy. Bulk data +(metadata, trickplay, cache, transcodes) binds to `/storage1`; only SQLite config lands on the +97%-full root disk. Plex runs untouched until Task 23. + +**Tech Stack:** Docker Compose, Portainer (Gitea git-polling), Caddy, Jellyfin 10.11, +PostgreSQL 16, systemd timers, Uptime Kuma push monitors. + +**Spec:** `docs/superpowers/specs/2026-08-05-jellyfin-migration-design.md` + +--- + +## A note on "tests" in this plan + +This is deployment configuration, not application code, so there is no test suite to write +first. The TDD discipline is preserved in a different form: **every task states its verification +command and expected output before the change is made**, and no task is complete until that +command produces that output. Treat a failed verification exactly as you would a failing test — +stop and diagnose, do not proceed. + +## Conventions this repo enforces + +Violating any of these breaks a live system. They are not stylistic. + +1. **The repo is canonical.** Editing files here changes nothing until pushed to Gitea. + Portainer polls `https://gitea.ginnoir.com/ginnoir/homelabstack` on `main` every 5 minutes. +2. **No `${VAR}` interpolation in compose, ever.** Every service gets `env_file: stack.env`, and + variable names in `stack.env` are exactly what the container reads. +3. **Portainer UI environment variables must be empty** for application stacks. All config comes + from `stack.env`. +4. **LF line endings only.** `.gitattributes` enforces this. CRLF in `stack.env` appends stray + `\r` to values and breaks containers silently. +5. **`stack.env` is committed intentionally.** Secrets are versioned in this private repo. Do not + gitignore or scrub them. +6. **Infra images are version-pinned** with `com.centurylabs.watchtower.enable=false`. Apps run + `:latest` under Watchtower. + +## Host facts you will need + +| Fact | Value | +|---|---| +| SSH | `ssh -o BatchMode=yes ginnoir@valhalla` | +| valhalla LAN IP | `192.168.1.69` | +| Plex | host-native systemd unit `plexmediaserver`, data at `/var/lib/plex` → `/storage1/labdata/plex` | +| Plex API | `http://192.168.1.69:32400` | +| Root disk | `/dev/sda2`, 194G, **97% full at plan time** | +| Libraries | `/storage1/Movies`, `/storage1/TV`, `/storage1/Anime` | +| Free ports | 8096, 5690, 3009 all confirmed unused | + +--- + +# PHASE 0 — Disk remediation + +**Nothing in Phase 1 may start until Task 4 is complete.** The root disk is at 97% and `/config` +lives on it. This phase is a hard gate. + +--- + +### Task 1: Reclaim Docker disk space + +**Files:** none — live host operation. + +- [ ] **Step 1: Record the baseline** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h / | tail -1; docker system df" +``` + +Write the output into the task notes below. Expected at plan time: `/dev/sda2 194G 179G 5.7G 97%` +and roughly `Images 124 / 123.5GB / 51.47GB reclaimable`. + +- [ ] **Step 2: Verify no container depends on what you are about to remove** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker ps -a --filter status=exited --format '{{.Names}}\t{{.Image}}'" +``` + +Expected: a short list, or empty. Any **exited container you still want** keeps its image, so +note them. `docker image prune -a` removes only images not referenced by *any* container, +including stopped ones — it will not break a stopped-but-wanted service. + +- [ ] **Step 3: Prune** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker image prune -af && docker builder prune -f" +``` + +- [ ] **Step 4: Verify space was reclaimed** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h / | tail -1" +``` + +Expected: `Avail` is now **at least 40G** and `Use%` is at or below **80%**. +If it is not, stop. Do not proceed to Task 2 — investigate what is holding the space. + +- [ ] **Step 5: Confirm all containers are still running** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker ps --format '{{.Names}}\t{{.Status}}' | grep -c Up" +``` + +Expected: **71 or more** (the count before the prune). Watchtower re-pulls anything it needs. + +--- + +### Task 2: Audit what else is consuming the root disk + +**Files:** none — live host operation. + +- [ ] **Step 1: Find the top consumers outside Docker** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo du -xh --max-depth=2 / 2>/dev/null | sort -h | tail -25" +``` + +`-x` keeps it on `/dev/sda2` and stops it wandering into the 94T of `/storage1`. + +- [ ] **Step 2: Check the Docker containerd image store specifically** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo du -sh /var/lib/docker /var/lib/containerd /var/log 2>/dev/null" +``` + +This is the directory that caused the 2026-07-07 outage (all Postgres instances crash-looped on +"No space left on device"). It should now be well under 60G. + +- [ ] **Step 3: Record findings** + +Append the output to this task as a comment. You are looking for anything **structurally +growing** — not one-off large files. If you find a second unbounded growth source, add a task +to this plan to bound it before continuing. + +- [ ] **Step 4: Commit the findings** + +```bash +git add docs/superpowers/plans/2026-08-05-jellyfin-migration.md +git commit -m "docs(plan): record root disk audit findings" +``` + +--- + +### Task 3: Add root-disk free-space alerting + +Uptime Kuma cannot inspect disk space directly. The pattern is a **push monitor**: valhalla +reports in on a timer, and reports `down` when free space is low. If the timer stops, Kuma also +alerts — so this catches both "disk full" and "monitoring broke". + +**Files:** +- Create on host: `/usr/local/bin/root-disk-check.sh` +- Create on host: `/etc/systemd/system/root-disk-check.service` +- Create on host: `/etc/systemd/system/root-disk-check.timer` + +> These are host-managed files, not repo files. valhalla has **no crontab** — systemd timers only +> (this matches the existing `docker-prune.timer`). + +- [ ] **Step 1: Create the Uptime Kuma push monitor and capture its token** + +In the Uptime Kuma UI at `https://uptime.ginnoir.com`: +1. **Add New Monitor** +2. Monitor Type: **Push** +3. Friendly Name: `valhalla root disk` +4. Heartbeat Interval: `300` seconds +5. Retries: `1` +6. Notifications: tick the existing **ntfy** notification +7. Save + +The monitor page then shows a Push URL of the form +`https://uptime.ginnoir.com/api/push/XXXXXXXXXX`. **Copy the token** (the `XXXXXXXXXX` part). + +Record it here before continuing — the next step needs it: + +``` +PUSH_TOKEN = ____________________ +``` + +- [ ] **Step 2: Write the check script on the host** + +Substitute your real token for `PUSH_TOKEN_HERE`. + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo tee /usr/local/bin/root-disk-check.sh >/dev/null <<'EOF' +#!/bin/sh +# Reports root-disk headroom to Uptime Kuma. Alerts below MIN_FREE_PCT. +set -eu +MIN_FREE_PCT=15 +TOKEN=PUSH_TOKEN_HERE + +USED_PCT=\$(df --output=pcent / | tail -1 | tr -dc '0-9') +FREE_PCT=\$((100 - USED_PCT)) +AVAIL=\$(df -h --output=avail / | tail -1 | tr -d ' ') + +if [ \"\$FREE_PCT\" -ge \"\$MIN_FREE_PCT\" ]; then + STATUS=up +else + STATUS=down +fi + +curl -fsS --max-time 10 \\ + \"https://uptime.ginnoir.com/api/push/\${TOKEN}?status=\${STATUS}&msg=root%20\${AVAIL}%20free%20(\${FREE_PCT}%25)\" \\ + >/dev/null +EOF +sudo chmod 755 /usr/local/bin/root-disk-check.sh" +``` + +- [ ] **Step 3: Run it once by hand and verify Kuma turns green** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo /usr/local/bin/root-disk-check.sh && echo OK" +``` + +Expected: `OK`, and the `valhalla root disk` monitor in Kuma flips to **Up** within a few +seconds with a message like `root 45G free (23%)`. + +If it prints a curl error instead, the token is wrong — fix it before continuing. + +- [ ] **Step 4: Create the systemd service and timer** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo tee /etc/systemd/system/root-disk-check.service >/dev/null <<'EOF' +[Unit] +Description=Report root disk headroom to Uptime Kuma + +[Service] +Type=oneshot +ExecStart=/usr/local/bin/root-disk-check.sh +EOF +sudo tee /etc/systemd/system/root-disk-check.timer >/dev/null <<'EOF' +[Unit] +Description=Run root-disk-check every 5 minutes + +[Timer] +OnBootSec=2min +OnUnitActiveSec=5min +AccuracySec=30s + +[Install] +WantedBy=timers.target +EOF +sudo systemctl daemon-reload && sudo systemctl enable --now root-disk-check.timer" +``` + +- [ ] **Step 5: Verify the timer is scheduled** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "systemctl list-timers root-disk-check.timer --no-pager" +``` + +Expected: one row showing `root-disk-check.timer` with a `NEXT` time under 5 minutes away. + +- [ ] **Step 6: Verify the alert path actually fires** + +Temporarily raise the threshold so the script reports `down`: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo sed -i 's/^MIN_FREE_PCT=15/MIN_FREE_PCT=99/' /usr/local/bin/root-disk-check.sh && sudo /usr/local/bin/root-disk-check.sh" +``` + +Expected: the Kuma monitor flips to **Down** and an ntfy notification arrives. Then restore it: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo sed -i 's/^MIN_FREE_PCT=99/MIN_FREE_PCT=15/' /usr/local/bin/root-disk-check.sh && sudo /usr/local/bin/root-disk-check.sh" +``` + +Expected: monitor returns to **Up**. An untested alert is not an alert. + +--- + +### Task 4: Correct the Plex documentation in CLAUDE.md + +`CLAUDE.md` documents a Plex data path that does not exist. Anyone (human or agent) following it +will operate on the wrong directory. + +**Files:** +- Modify: `CLAUDE.md` (the `### Plex Media Server` section) + +- [ ] **Step 1: Verify the true paths one more time** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "ls -ld /var/lib/plex; readlink -f /var/lib/plex; ls -d /var/lib/plexmediaserver 2>&1" +``` + +Expected: `/var/lib/plex -> /storage1/labdata/plex`, and +`ls: cannot access '/var/lib/plexmediaserver': No such file or directory`. + +- [ ] **Step 2: Replace the incorrect line** + +In `CLAUDE.md`, find: + +```markdown +- **App data / metadata / DB:** `/var/lib/plexmediaserver/Library/Application Support/Plex Media Server/` +``` + +Replace it with: + +```markdown +- **App data / metadata / DB:** `/var/lib/plex/Plex Media Server/` — note `/var/lib/plex` is a + **symlink to `/storage1/labdata/plex`**, so Plex's 383G of app data lives on the ZFS array, not + the root disk. (`/var/lib/plexmediaserver/` does not exist.) +``` + +- [ ] **Step 3: Also correct the media library paths in the same section** + +Find: + +```markdown +- **Media libraries:** `/storage1/Media/{Movies,TV,Anime}` — same ZFS pool bind-mounted into the *arr containers. +``` + +Replace with: + +```markdown +- **Media libraries:** `/storage1/{Movies,TV,Anime}` — `/storage1/Media/*` are symlinks to these. + Same ZFS pool bind-mounted into the *arr containers. +``` + +- [ ] **Step 4: Verify no other stale Plex paths remain** + +```bash +grep -rn "plexmediaserver/Library" . --include="*.md" +``` + +Expected: **no output**. If the string appears in `.claude/skills/` or `.agents/skills/`, fix +those too. + +- [ ] **Step 5: Commit** + +```bash +git add CLAUDE.md +git commit -m "docs: correct Plex data and library paths in CLAUDE.md + +/var/lib/plexmediaserver does not exist; the real path is /var/lib/plex, +itself a symlink to /storage1/labdata/plex. Media libraries are +/storage1/{Movies,TV,Anime} with /storage1/Media/* as symlinks to them." +``` + +**PHASE 0 GATE:** Do not continue until `df -h /` shows at least 40G available and the +`valhalla root disk` Kuma monitor is green. + +--- + +# PHASE 1 — Deploy the streaming stack + +--- + +### Task 5: Create the host directory tree + +Directories must exist and be owned by UID/GID 1000 **before** the containers start, or Jellyfin +will fail to write and the bind mounts will be created root-owned by Docker. + +**Files:** none — live host operation. + +- [ ] **Step 1: Create the directories** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "mkdir -p /storage1/labdata/jellyfin/{metadata,cache,trickplay,transcodes} && sudo mkdir -p /config/{jellyfin,jellystat,jellystat-db,wizarr} && sudo chown -R 1000:1000 /config/jellyfin /config/jellystat /config/wizarr /storage1/labdata/jellyfin && echo CREATED" +``` + +`/config/jellystat-db` is deliberately **not** chowned — the postgres image manages its own +ownership as UID 999. + +- [ ] **Step 2: Verify ownership and placement** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "ls -ld /config/jellyfin /config/jellystat /config/wizarr /storage1/labdata/jellyfin/*; df -h /config/jellyfin /storage1/labdata/jellyfin | tail -3" +``` + +Expected: `/config/*` owned by `1000 1000` and on `/dev/sda2`; `/storage1/labdata/jellyfin/*` +owned by `1000 1000` and on `storage1` (virtiofs). **If the `/storage1` paths report `/dev/sda2`, +stop — `/storage1` is not mounted** and you would be writing into the placeholder directory that +caused the 2026-07-06 incident. + +--- + +### Task 6: Write the streaming stack compose file + +**Files:** +- Create: `stacks/streaming/docker-compose.yml` + +- [ ] **Step 1: Generate the secrets you will need in Task 7** + +```bash +echo "POSTGRES_PASSWORD=$(openssl rand -hex 24)" +echo "JWT_SECRET=$(openssl rand -hex 32)" +``` + +Keep this output — Task 7 pastes it into `stack.env`. + +- [ ] **Step 2: Create `stacks/streaming/docker-compose.yml`** + +```yaml +# streaming stack — Jellyfin and its satellites. +# +# Storage split is deliberate: SQLite config on the root SSD, all bulk data +# (metadata, trickplay, cache, transcodes) on /storage1. The root disk is only +# 194G and has filled twice; nothing here may grow unbounded on it. +# +# Trickplay gets an explicit sub-bind because Jellyfin 10.11 moved it to +# /config/data/trickplay and it is NOT relocatable from the UI. +# +# Transcodes are disk-backed, not tmpfs: jellyfin#16608 means Remux and +# DirectStream jobs finish before the segment cleaner's first check, orphaning +# segments. Plex already writes transcodes to virtiofs at a 48% transcode rate. +# +# Env convention: every service gets stack.env via env_file; var names there are +# exactly what the containers read. No ${VAR} interpolation, no Portainer UI +# env vars. + +services: + jellyfin: + container_name: jellyfin + image: lscr.io/linuxserver/jellyfin:latest + restart: unless-stopped + networks: [streaming, edge] + env_file: + - stack.env + environment: + - JELLYFIN_PublishedServerUrl=https://jellyfin.ginnoir.com + volumes: + - /etc/localtime:/etc/localtime:ro + - /config/jellyfin:/config + - /storage1/labdata/jellyfin/metadata:/config/metadata + - /storage1/labdata/jellyfin/cache:/config/cache + - /storage1/labdata/jellyfin/trickplay:/config/data/trickplay + - /storage1/labdata/jellyfin/transcodes:/transcode + - /storage1/Movies:/data/movies:ro + - /storage1/TV:/data/tv:ro + - /storage1/Anime:/data/anime:ro + ports: + - "8096:8096" + + jellystat-db: + container_name: jellystat-db + image: postgres:16-alpine + restart: unless-stopped + networks: [streaming] + env_file: + - stack.env + labels: + - "com.centurylabs.watchtower.enable=false" + volumes: + - /config/jellystat-db:/var/lib/postgresql/data + + jellystat: + container_name: jellystat + image: cyfershepard/jellystat:latest + restart: unless-stopped + networks: [streaming, edge] + depends_on: + - jellystat-db + env_file: + - stack.env + volumes: + - /config/jellystat:/app/backend/backup-data + ports: + - "3009:3000" + + wizarr: + container_name: wizarr + image: ghcr.io/wizarrrr/wizarr:latest + restart: unless-stopped + networks: [streaming, edge] + env_file: + - stack.env + volumes: + - /config/wizarr:/data/database + ports: + - "5690:5690" + + jellyplex-watched: + container_name: jellyplex-watched + image: luigi311/jellyplex-watched:latest + restart: unless-stopped + networks: [streaming, edge] + env_file: + - stack.env + +networks: + streaming: + name: streaming + driver: bridge + edge: + name: edge + external: true +``` + +- [ ] **Step 3: Verify the YAML parses** + +```bash +docker compose -f stacks/streaming/docker-compose.yml config --quiet && echo "YAML OK" +``` + +Expected: `YAML OK`. A non-zero exit means a syntax error — fix before committing. + +> If `stack.env` does not exist yet this may warn about a missing env file. Complete Task 7 and +> re-run this step. + +--- + +### Task 7: Write the streaming stack environment file + +**Files:** +- Create: `stacks/streaming/stack.env` + +- [ ] **Step 1: Get the Plex token** + +The token is in Plex's preferences on the host: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo grep -o 'PlexOnlineToken=\"[^\"]*\"' '/var/lib/plex/Plex Media Server/Preferences.xml'" +``` + +Expected: `PlexOnlineToken="xxxxxxxxxxxxxxxxxxxx"`. Record the value. + +- [ ] **Step 2: Create `stacks/streaming/stack.env`** + +Substitute the two secrets generated in Task 6 Step 1 and the Plex token from Step 1. +`JELLYFIN_TOKEN` is filled in later, in Task 16 — leave it as the empty placeholder for now. + +```bash +# streaming stack config — injected into every service via env_file. +# Var names are exactly what the containers read; no ${VAR} interpolation, +# no Portainer UI env vars needed. + +# --- shared --- +TZ=America/Chicago +PUID=1000 +PGID=1000 + +# --- jellystat-db (postgres:16-alpine) --- +POSTGRES_DB=jfstat +POSTGRES_USER=jellystat +POSTGRES_PASSWORD=REPLACE_WITH_GENERATED_POSTGRES_PASSWORD + +# --- jellystat --- +POSTGRES_IP=jellystat-db +POSTGRES_PORT=5432 +JWT_SECRET=REPLACE_WITH_GENERATED_JWT_SECRET + +# --- jellyplex-watched --- +# Plex is host-native, not containerised, so it is reached on the LAN IP. +PLEX_URL=http://192.168.1.69:32400 +PLEX_TOKEN=REPLACE_WITH_PLEX_TOKEN +JELLYFIN_BASEURL=http://jellyfin:8096 +JELLYFIN_TOKEN= +SYNC_FROM_PLEX_TO_JELLYFIN=True +SYNC_FROM_JELLYFIN_TO_PLEX=True +# Start in dry-run. Task 16 flips this to False after the mapping is verified. +DRYRUN=True +SLEEP_DURATION=21600 +LOG_LEVEL=INFO +``` + +- [ ] **Step 3: Verify LF line endings** + +CRLF in `stack.env` appends `\r` to every value and breaks containers in ways that are very hard +to diagnose. + +```bash +file stacks/streaming/stack.env +``` + +Expected: `ASCII text` — **not** `ASCII text, with CRLF line terminators`. If it reports CRLF: + +```bash +git add --renormalize stacks/streaming/stack.env +``` + +- [ ] **Step 4: Verify the compose file now resolves cleanly** + +```bash +docker compose -f stacks/streaming/docker-compose.yml config --quiet && echo "COMPOSE OK" +``` + +Expected: `COMPOSE OK`. + +- [ ] **Step 5: Verify Portainer UI env compliance** + +```bash +pwsh -File scripts/check-portainer-stack-env.ps1 +``` + +Expected: no failures reported for `streaming`. This script exists to catch the exact mistake of +leaving UI environment variables set. + +- [ ] **Step 6: Commit** + +```bash +git add stacks/streaming/docker-compose.yml stacks/streaming/stack.env +git commit -m "feat(streaming): add jellyfin stack + +New Portainer stack: jellyfin, jellystat + postgres, wizarr and +jellyplex-watched. Bulk data (metadata, trickplay, cache, transcodes) +binds to /storage1; only SQLite config lands on the root disk. + +Trickplay gets an explicit sub-bind because 10.11 moved it to +/config/data/trickplay, which is not relocatable from the UI. Transcodes +are disk-backed rather than tmpfs because jellyfin#16608 orphans segments +on the Remux/DirectStream jobs that are routine on this server." +``` + +--- + +### Task 8: Add the Caddy site blocks + +**Files:** +- Modify: `Caddyfile` + +`jellyfin` and `wizarr` are **public** — remote users and invite recipients reach them from +anywhere, which is the entire point of leaving Plex. `jellystat` is internal-only and follows the +existing tautulli pattern exactly. + +> Note the `internal_gate` vs `internal_only` distinction: inside a `route {}` block you **must** +> use `internal_gate`, because Caddy sorts `route` ahead of `respond` and an `import internal_only` +> above a catch-all route is dead code. This repo already had to fix that on nine sites. + +- [ ] **Step 1: Add the public streaming blocks** + +Insert after the `requests.ginnoir.com` block (around line 134): + +```caddy +# ============================================================= +# STREAMING — public (replaces Plex's own remote access) +# ============================================================= +jellyfin.ginnoir.com { + reverse_proxy jellyfin:8096 +} + +wizarr.ginnoir.com { + reverse_proxy wizarr:5690 +} +``` + +- [ ] **Step 2: Add the internal-only stats block** + +Insert immediately after the `tautulli.ginnoir.com` block (around line 205): + +```caddy +jellystat.ginnoir.com { + route { + import internal_gate + import authentik_outpost + handle { + import authentik_forward_auth + reverse_proxy jellystat:3000 + } + } +} +``` + +- [ ] **Step 3: Validate the Caddyfile before pushing** + +```bash +docker run --rm -v "${PWD}/Caddyfile:/etc/caddy/Caddyfile:ro" caddy:latest caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile +``` + +Expected: `Valid configuration`. + +> This will warn about the `cloudflare` DNS module being unknown — that module is baked into the +> custom Caddy image on valhalla, not the stock one. Warnings about `acme_dns cloudflare` are +> expected and safe. Any **structural** error is not. + +- [ ] **Step 4: Commit** + +```bash +git add Caddyfile +git commit -m "feat(caddy): add jellyfin, wizarr and jellystat sites + +jellyfin and wizarr are public by design — they replace Plex's own +remote access. jellystat is internal_gate + forward_auth, matching the +tautulli pattern." +``` + +--- + +### Task 9: Register the stack in Portainer and deploy + +New stacks must be registered in Portainer once, by hand. After that, git-polling handles every +subsequent change. + +**Files:** none — push and UI operation. + +- [ ] **Step 1: Push everything so far** + +```bash +git push origin main +``` + +This makes the stack files visible to Portainer. **Caddy reloads on this push too** — the Gitea +Actions runner copies the Caddyfile and runs `caddy reload`. The three new sites will 502 until +the containers exist, which is expected and harmless. + +- [ ] **Step 2: Register the stack** + +In the Portainer UI at `https://portainer.ginnoir.com`: +1. **Stacks → Add stack** +2. Name: `streaming` +3. Build method: **Repository** +4. Repository URL: `https://gitea.ginnoir.com/ginnoir/homelabstack` +5. Repository reference: `refs/heads/main` +6. Compose path: `stacks/streaming/docker-compose.yml` +7. **Authentication: on.** Use the fine-grained Gitea PAT that the other stacks use — the runner + PAT cannot clone. +8. **GitOps updates: on**, Mechanism **Polling**, Interval `5m` +9. **Environment variables: leave completely empty.** This is enforced by + `scripts/check-portainer-stack-env.ps1`. +10. **Deploy the stack** + +> Known quirk: the create call may appear to time out while actually succeeding. If it does, do +> **not** retry — refresh the stack list and confirm `streaming` is present before doing anything +> else, or you will create a duplicate. + +- [ ] **Step 3: Verify all five containers are running** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker ps --filter label=com.docker.compose.project=streaming --format 'table {{.Names}}\t{{.Status}}'" +``` + +Expected: five rows — `jellyfin`, `jellystat`, `jellystat-db`, `wizarr`, `jellyplex-watched` — +all `Up`. + +`jellyplex-watched` may show as restarting or exited: it has no Jellyfin token yet and is in +dry-run. That is expected until Task 16. + +- [ ] **Step 4: Verify Jellyfin is serving** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8096/health" +``` + +Expected: `200`. + +- [ ] **Step 5: Verify the bind mounts landed where intended** + +This is the single most important verification in Phase 1. If trickplay or transcodes end up on +the root disk, the disk will fill again. + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker inspect jellyfin --format '{{range .Mounts}}{{.Source}} -> {{.Destination}} ({{.Mode}}){{\"\n\"}}{{end}}'" +``` + +Expected, exactly these sources: +- `/config/jellyfin -> /config` +- `/storage1/labdata/jellyfin/metadata -> /config/metadata` +- `/storage1/labdata/jellyfin/cache -> /config/cache` +- `/storage1/labdata/jellyfin/trickplay -> /config/data/trickplay` +- `/storage1/labdata/jellyfin/transcodes -> /transcode` +- `/storage1/Movies -> /data/movies (ro)` +- `/storage1/TV -> /data/tv (ro)` +- `/storage1/Anime -> /data/anime (ro)` + +The three library mounts **must** show `ro`. + +- [ ] **Step 6: Verify the public site resolves over HTTPS** + +```bash +curl -s -o /dev/null -w '%{http_code}\n' https://jellyfin.ginnoir.com/health +``` + +Expected: `200`. If you get a TLS error, check that the certificate issued — DNS-01 renewals on +this host have failed before when the split-horizon override shadowed the SOA record. The canary: + +```bash +dig +short SOA ginnoir.com @192.168.1.1 +``` + +Expected: a real SOA record, not empty. + +--- + +### Task 10: Complete Jellyfin setup and add libraries + +**Files:** none — UI operation. + +- [ ] **Step 1: Run the setup wizard** + +Browse to `https://jellyfin.ginnoir.com`. On the wizard: +1. Preferred display language: English +2. Create the admin account. **Use a strong, unique password** — this server is publicly + reachable and Plex's relay is no longer hiding it. +3. Skip adding libraries here; Step 2 does it with the right settings. +4. Set metadata language / country +5. **Allow remote connections: yes.** Leave automatic port mapping **off** — Caddy handles ingress. + +- [ ] **Step 2: Add the three libraries** + +Dashboard → Libraries → Add Media Library. Create each with these settings: + +| Setting | Movies | TV | Anime | +|---|---|---|---| +| Content type | Movies | Shows | Shows | +| Display name | `Movies` | `TV` | `Anime` | +| Folder | `/data/movies` | `/data/tv` | `/data/anime` | +| Preferred metadata language | English | English | English | + +For **all three**, in the library settings: +- **Enable "Save artwork into media folders": OFF**. The library mounts are read-only, and + leaving this on generates a stream of write errors in the logs. +- **"Enable trickplay image extraction": OFF for now.** Task 13 turns it on deliberately, after + the thread cap is configured. Leaving it on here starts an uncapped multi-day job that will + compete with the initial scan. + +For **Anime only**, additionally set metadata downloaders to prefer **AniDB** and **AniList** +above TheMovieDb (these plugins are installed in Task 14 — revisit this after that task). + +- [ ] **Step 3: Trigger and watch the initial scan** + +Dashboard → Scheduled Tasks → **Scan Media Library** → Run. + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker logs --tail 40 jellyfin" +``` + +- [ ] **Step 4: Verify item counts against Plex** + +Once the scan completes (expect a few hours for ~1891 items), check Dashboard → Libraries. + +Expected, approximately: +- Movies: **~1357** +- TV: **~173 series** +- Anime: **~361 series** + +A shortfall over ~2% means naming problems worth investigating before you build anything on top +of this library. + +- [ ] **Step 5: Verify nothing significant landed on the root disk** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h / | tail -1; sudo du -sh /config/jellyfin; du -sh /storage1/labdata/jellyfin/*" +``` + +Expected: `/config/jellyfin` in the low **hundreds of MB to a few GB** (SQLite only), with +`metadata` on `/storage1` holding the artwork. If `/config/jellyfin` is tens of GB, a bind mount +is wrong — go back to Task 9 Step 5. + +- [ ] **Step 6: Regenerate the bookmark files** + +```bash +pwsh -File scripts/gen-bookmarks.ps1 +git add bookmarks-domains.html bookmarks-ports.html +git commit -m "chore: regenerate bookmarks for streaming stack" +git push origin main +``` + +--- + +# PHASE 2 — Parity with Plex + +--- + +### Task 11: Install Plexyfin and import curated artwork and collections + +Plexyfin is the only reason any of Plex's 383G is worth reading. It copies the curated **15G** +`Metadata` investment — posters, backdrops and collections — directly from the live PMS API. + +**This plugin is destructive.** It deletes and replaces artwork wholesale on every run. It is +installed, run once, and then removed. + +**Files:** none — UI operation against the live Plex API. + +- [ ] **Step 1: Confirm Plex is still running and reachable** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "systemctl is-active plexmediaserver; curl -s -o /dev/null -w '%{http_code}\n' 'http://192.168.1.69:32400/identity'" +``` + +Expected: `active` and `200`. Everything in this task reads from the live Plex server. + +- [ ] **Step 2: Add the Plexyfin repository** + +Jellyfin Dashboard → Plugins → Repositories → **Add**: +- Repository Name: `Plexyfin` +- Repository URL: `https://raw.githubusercontent.com/cleverdevil/plexyfin/main/metadata/stable/manifest.json` + +- [ ] **Step 3: Install and restart** + +Dashboard → Plugins → Catalog → **Plexyfin** → Install. Then restart Jellyfin: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker restart jellyfin && sleep 20 && curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8096/health" +``` + +Expected: `200`. + +- [ ] **Step 4: Configure Plexyfin** + +Dashboard → Plugins → Plexyfin: +- Plex Server URL: `http://192.168.1.69:32400` +- Plex Token: the token recorded in Task 7 Step 1 +- Enable **Sync Collections** +- Enable **Sync Artwork** +- Save + +- [ ] **Step 5: Run the sync** + +Use the plugin's **Sync Now** button, then watch: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker logs --tail 60 -f jellyfin" +``` + +Expected: progress lines naming collections and items. Let it finish completely — interrupting +leaves artwork half-replaced. + +- [ ] **Step 6: Verify collections and artwork arrived** + +In the Jellyfin web UI: +- Movies → Collections shows your Plex collections +- Spot-check five movies where you know you set custom posters in Plex; confirm the poster + matches Plex, not TMDB's default + +- [ ] **Step 7: Remove the plugin** + +This is not optional. Left installed, it will destroy any artwork you later customise in +Jellyfin. + +Dashboard → Plugins → Plexyfin → **Uninstall**. Then Dashboard → Plugins → Repositories → +remove the `Plexyfin` repository. Restart Jellyfin: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker restart jellyfin && sleep 20 && curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8096/health" +``` + +Expected: `200`. + +- [ ] **Step 8: Verify it is gone** + +Dashboard → Plugins → My Plugins. Expected: **Plexyfin is not listed.** + +--- + +### Task 12: Configure transcoding — NVDEC decode, CPU encode, segment cleanup + +The P100 cannot encode — GP100 has no NVENC — so this does nothing for live streaming transcode. +It exists for one reason: **trickplay generation in Task 13 is decode-plus-JPEG-encode**, and +NVDEC accelerates exactly that. + +**Files:** +- Modify: `stacks/streaming/docker-compose.yml` + +- [ ] **Step 1: Check whether llama-swap currently holds the GPU** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv" +``` + +If a llama-swap process holds most of the 16G, either wait for it to idle or accept that +trickplay generation will be slower. Note what you found. + +- [ ] **Step 2: Add the CDI device to the jellyfin service** + +In `stacks/streaming/docker-compose.yml`, inside the `jellyfin` service, add after `env_file`: + +```yaml + devices: + - nvidia.com/gpu=0 +``` + +- [ ] **Step 3: Commit and push** + +```bash +git add stacks/streaming/docker-compose.yml +git commit -m "feat(streaming): give jellyfin CDI access to the P100 + +Decode only — GP100 has no NVENC. This is for trickplay generation, +which is decode plus JPEG encode." +git push origin main +``` + +- [ ] **Step 4: Wait for the redeploy, then verify the GPU is visible in the container** + +Portainer polls every 5 minutes. Then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker exec jellyfin nvidia-smi -L" +``` + +Expected: `GPU 0: Tesla P100-PCIE-16GB (UUID: GPU-f23318db-...)`. + +**If this errors**, the CDI device name is wrong. Check the available names and correct the +compose file: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "grep -E '^\s+name:' /etc/cdi/nvidia.yaml" +``` + +- [ ] **Step 5: Enable NVDEC in Jellyfin** + +Dashboard → Playback → Transcoding: +- Hardware acceleration: **NVIDIA NVENC** +- Enable hardware decoding for: **H264**, **HEVC**, **VP9** +- **Leave "Enable hardware encoding" UNCHECKED.** The P100 cannot encode; ticking this makes + every transcode fail. +- Save + +- [ ] **Step 5a: Configure the CPU encoder and segment cleanup** + +On the same Dashboard → Playback → Transcoding page, since encoding is CPU-bound: + +- **Encoding preset: `veryfast`.** The default is slower and this host has no encode offload. +- **Transcode thread count: `8`.** Per session. With a measured peak of 7 concurrent streams, + this bounds the worst case at 56 of 52 threads — deliberately just over-subscribed rather than + unbounded, which is what the default does. +- **Enable throttling: ON.** Pauses ffmpeg once it runs far enough ahead of the player. +- **Enable segment deletion: ON.** +- **Segment keep seconds: `360`.** +- Save + +Segment deletion matters more here than it does on most servers. Per +[jellyfin#16608](https://github.com/jellyfin/jellyfin/issues/16608), the cleaner's first check is +at 20 seconds and the throttler's at 5, but Remux and DirectStream jobs finish in 2-7 seconds — +so on those, neither ever fires. DirectStream (copy video + transcode audio) was **8 of 25** +sampled sessions on this server. This setting does not fix that bug; it bounds the damage on +every *other* job type. The transcode directory is on `/storage1` precisely because this bug +cannot be configured away. + +- [ ] **Step 5b: Verify orphaned segments are actually being cleaned** + +Play something that transcodes, let it run two minutes, stop it, wait five minutes, then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "du -sh /storage1/labdata/jellyfin/transcodes; ls /storage1/labdata/jellyfin/transcodes | wc -l" +``` + +Expected: small size and a low file count. If this directory grows monotonically over the +following days, add a weekly systemd timer to prune files older than 24 hours — the same pattern +as Task 3's timer. + +- [ ] **Step 6: Verify a decode actually uses the GPU** + +Play a HEVC file that requires transcoding, from the web client, then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv | grep -i ffmpeg || echo 'no ffmpeg on GPU'" +``` + +Expected: an `ffmpeg` process listed. If you see `no ffmpeg on GPU`, decoding fell back to CPU — +check Dashboard → Playback logs for the reason. This is not fatal; it only makes Task 13 slower. + +--- + +### Task 13: Generate trickplay + +Plex's 365G of BIF preview thumbnails is **format-incompatible** with Jellyfin's tiled JPEG +trickplay. None of it is reusable. Budget **300-400G on `/storage1`** and a multi-day first pass +over ~1891 items. + +**Files:** none — UI operation plus monitoring. + +- [ ] **Step 1: Confirm you have the disk headroom** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h /storage1 | tail -1" +``` + +Expected: **at least 1T available**. At plan time there was 17T, so this is a formality — but +confirm, because this task writes hundreds of gigabytes. + +- [ ] **Step 2: Configure trickplay with a thread cap** + +Dashboard → Playback → Trickplay: +- Enable trickplay image extraction **during library scan**: **ON** +- **Thread count: `20`.** The host has 52 threads. Leaving this at the default (all cores) starves + live transcoding, which is the thing users actually notice. +- Interval: `10000` ms (the default) +- Width resolutions: `320` (the default) +- Enable hardware decoding: **ON** (this is what Task 12 was for) +- Save + +- [ ] **Step 3: Enable trickplay on each library** + +Dashboard → Libraries. For each of Movies, TV and Anime, edit the library and tick **Enable +trickplay image extraction**. (These were deliberately left off in Task 10 Step 2.) + +- [ ] **Step 4: Start generation** + +Dashboard → Scheduled Tasks → **Generate Trickplay Images** → Run. + +- [ ] **Step 5: Monitor progress and disk growth** + +Check periodically over the following days: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "du -sh /storage1/labdata/jellyfin/trickplay; df -h /storage1 | tail -1; uptime" +``` + +Expected: `trickplay` growing steadily toward the 300-400G range; `/storage1` avail falling +correspondingly; load average high but the host responsive. + +**Abort condition:** if `/storage1` available space drops below **500G**, stop the scheduled task +immediately and reassess — do not let this fill the array. + +- [ ] **Step 6: Verify trickplay works in a client** + +Once the task reports complete, open any movie in the web client and scrub the timeline. +Expected: preview thumbnails appear above the scrubber. + +- [ ] **Step 7: Confirm nothing leaked onto the root disk** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h / | tail -1; sudo du -sh /config/jellyfin" +``` + +Expected: root disk essentially unchanged from Task 10. If it grew by tens of GB, the trickplay +sub-bind is not working — stop and fix it. + +--- + +### Task 14: Install the remaining plugins + +**Files:** none — UI operation. + +- [ ] **Step 1: Install Intro Skipper** + +The original plugin is unmaintained; use the maintained fork. **Verify 10.11 compatibility before +installing** — this plugin patches the web player and a version mismatch breaks playback. + +1. Check the fork's README for a stated 10.11 compatible release: + `https://github.com/Bobby306/intro-skipper` +2. If, and only if, it declares 10.11 support, add its repository under + Dashboard → Plugins → Repositories using the manifest URL from its README, then install. +3. If it does **not** declare 10.11 support, **skip this plugin** and note it here. It is optional + polish, not a migration requirement. + +- [ ] **Step 2: Install the anime metadata plugins** + +Dashboard → Plugins → Catalog. Install from the official catalogue: +- **AniDB** +- **AniList** + +- [ ] **Step 3: Restart and verify** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker restart jellyfin && sleep 25 && curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8096/health" +``` + +Expected: `200`. + +- [ ] **Step 4: Point the Anime library at the anime providers** + +Dashboard → Libraries → Anime → edit. Under metadata downloaders, drag **AniDB** and **AniList** +above **TheMovieDb**. Save, then run **Scan Media Library** limited to the Anime library. + +- [ ] **Step 5: Verify anime metadata improved** + +Spot-check five anime series. Expected: correct season/episode numbering and Japanese-release +metadata rather than TMDB's Western groupings. + +> **Do not install:** `jellyfin-plugin-sso` (browser-only OIDC — Roku, Fire TV and tvOS cannot use +> it, and those are the entire user base), Open Subtitles (Bazarr already writes subtitles to disk +> and does it better), Playback Reporting (Jellystat supersedes it). + +--- + +### Task 15: Connect Sonarr, Radarr and Bazarr to Jellyfin + +New imports must refresh Jellyfin the way they currently refresh Plex. **Leave the existing Plex +connections in place** — both servers stay live until Task 22. + +**Files:** none — UI operation. + +- [ ] **Step 1: Create a Jellyfin API key** + +Jellyfin Dashboard → API Keys → **New API Key**, name it `arr-integration`. Copy the key. + +Record it here — Task 16 needs a Jellyfin key too, and you can reuse this one: + +``` +JELLYFIN_API_KEY = ____________________ +``` + +- [ ] **Step 2: Add the connection in Sonarr** + +At `https://sonarr.ginnoir.com` → Settings → Connect → **+** → **Emby / Jellyfin**: +- Name: `Jellyfin` +- Host: `jellyfin` +- Port: `8096` +- API Key: the key from Step 1 +- Enable **On Import**, **On Upgrade**, **On Rename**, **On Series Delete** +- Update Library: **on** +- **Test**, then Save + +Expected: the Test button returns success. A failure here usually means the containers cannot see +each other — both are on `edge`, so check that first. + +- [ ] **Step 3: Add the connection in Radarr** + +At `https://radarr.ginnoir.com` → Settings → Connect → **+** → **Emby / Jellyfin**. Same host, +port and key. Enable **On Import**, **On Upgrade**, **On Rename**, **On Movie Delete**. Test and +Save. + +- [ ] **Step 4: Add the connection in Bazarr** + +At `https://bazarr.ginnoir.com` → Settings → Subtitles → enable the Jellyfin/Emby integration +with the same host, port and key, so subtitle downloads trigger a refresh. + +- [ ] **Step 5: Verify end to end** + +Trigger a rename or manual import in Radarr on a single film, then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker logs --tail 30 jellyfin | grep -i 'refresh\|library'" +``` + +Expected: a library refresh entry appearing within a minute of the Radarr action. + +--- + +# PHASE 3 — Watch-state sync + +--- + +### Task 16: Bring up two-way watch history sync + +JellyPlex-Watched matches items by filename and provider IDs. It is good but not perfect, so this +task runs it in dry-run first and only enables writes once the mapping is confirmed. + +**Files:** +- Modify: `stacks/streaming/stack.env` + +- [ ] **Step 1: Fill in the Jellyfin token** + +In `stacks/streaming/stack.env`, set `JELLYFIN_TOKEN` to the API key recorded in Task 15 Step 1. +Leave `DRYRUN=True` for now. + +- [ ] **Step 2: Commit, push and wait for the redeploy** + +```bash +git add stacks/streaming/stack.env +git commit -m "feat(streaming): give jellyplex-watched its jellyfin token" +git push origin main +``` + +Portainer redeploys within 5 minutes. + +- [ ] **Step 3: Run a dry-run and read the output carefully** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker restart jellyplex-watched && sleep 90 && docker logs --tail 120 jellyplex-watched" +``` + +Expected: it connects to both servers, lists users, and reports what it *would* sync. **Nothing is +written in dry-run.** + +- [ ] **Step 4: Check for username mismatches** + +In the dry-run output, look for users present on one server but not the other. Plex usernames and +Jellyfin usernames will not match for anyone who chose a different name at signup. + +If there are mismatches, add a mapping to `stacks/streaming/stack.env` using the tool's +`USER_MAPPING` format — a JSON object of `plex_username: jellyfin_username`: + +```bash +USER_MAPPING={"plexname1": "jellyfinname1", "plexname2": "jellyfinname2"} +``` + +Commit, push, wait for redeploy, and re-run Step 3 until the mapping is clean. + +> Early on only your own account exists in Jellyfin, so most Plex users will legitimately have no +> counterpart. That is expected — this mapping gets revisited as people are onboarded in Task 19. + +- [ ] **Step 5: Verify your own account's history would sync correctly** + +In the dry-run output, find your own user. Expected: a plausible list of watched items matching +what you know you have watched in Plex. If it reports zero for a library you have definitely +watched from, path matching is failing — diagnose before enabling writes. + +- [ ] **Step 6: Enable writes** + +In `stacks/streaming/stack.env`: + +```bash +DRYRUN=False +``` + +```bash +git add stacks/streaming/stack.env +git commit -m "feat(streaming): enable jellyplex-watched writes after dry-run verification" +git push origin main +``` + +- [ ] **Step 7: Verify a real sync round-trip** + +After the redeploy, mark an episode watched **in Jellyfin** that is unwatched in Plex. Then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker restart jellyplex-watched && sleep 120 && docker logs --tail 60 jellyplex-watched" +``` + +Check Plex: the episode should now be watched there too. Then do the reverse — mark something +watched in Plex and confirm it appears in Jellyfin after the next cycle. + +Expected: both directions work. `SLEEP_DURATION=21600` means it runs every 6 hours from here on. + +**PHASE 3 GATE:** Jellyfin is now at parity with Plex, with two-way history sync. Everything +after this point involves other people. + +--- + +# PHASE 4 — Pilot + +--- + +### Task 17: Configure Wizarr + +Per the design decision, invites are issued **manually and selectively**. Most Plex users are +dormant. There is no bulk invite generation and no scripted account creation in this plan. + +**Files:** none — UI operation. + +- [ ] **Step 1: Complete Wizarr setup** + +Browse to `https://wizarr.ginnoir.com` and create the admin account. + +- [ ] **Step 2: Connect it to Jellyfin** + +In Wizarr settings, add a media server: +- Type: **Jellyfin** +- URL: `http://jellyfin:8096` +- API Key: the key from Task 15 Step 1 + +- [ ] **Step 3: Restrict which libraries invited users get** + +Grant **Movies**, **TV** and **Anime**. Nothing else exists, but set this explicitly so future +libraries are not exposed by default. + +- [ ] **Step 4: Customise the onboarding wizard** + +Configure the post-invite steps to cover the three platforms that matter here — Roku, Fire TV / +Android TV, and Apple TV / iOS — with the app store link for each. These account for the +overwhelming majority of real playback. + +- [ ] **Step 5: Verify with a throwaway invite** + +Create an invite, open it in a private browser window, and complete the flow with a test account. + +Expected: the account appears in Jellyfin Dashboard → Users with only the three libraries granted. +Delete the test account from both Wizarr and Jellyfin afterwards. + +--- + +### Task 18: Pilot with real users + +**Files:** none — human process. + +- [ ] **Step 1: Pick 3-5 willing users covering every platform** + +You need **at least one each on Roku, Fire TV / Android TV, and Apple TV / iOS**. Platform +coverage matters more than user count — a pilot of five Android users tells you nothing about +Roku. + +- [ ] **Step 2: Send each an individual Wizarr invite** + +- [ ] **Step 3: Add their username mappings to jellyplex-watched** + +As each pilot user creates their Jellyfin account, add their Plex→Jellyfin mapping to +`USER_MAPPING` in `stacks/streaming/stack.env` (see Task 16 Step 4), then commit and push. Their +history then syncs both ways and they can move between servers freely. + +- [ ] **Step 4: Collect feedback on the three things that actually break** + +Ask each pilot user specifically about: playback start time, whether anything stutters mid-stream, +and whether the onboarding wizard was followable. Vague "is it working?" gets vague answers. + +- [ ] **Step 5: Verify transcode behaviour under real load** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker exec jellyfin ps aux | grep -c '[f]fmpeg'; uptime" +``` + +Run this while pilot users are actually watching. Expected: load average well under 52. + +- [ ] **Step 6: Record the pilot outcome** + +Note per platform whether playback worked, and any transcoding observations. This is the input to +the Task 20 hardware decision. + +--- + +# PHASE 5 — Selective migration + +--- + +### Task 19: Onboard users on demand + +**Files:** none — human process, ongoing. + +- [ ] **Step 1: Write a short explainer to accompany invites** + +One that answers, in plain language: why the change is happening, what to install on their device, +that their watch history carries over, and that Plex keeps working until an announced date. + +**This is sent with an invite, not broadcast.** Save it somewhere you can paste from. + +- [ ] **Step 2: Issue invites individually as people ask or as you choose** + +Dormant accounts get nothing. There is no campaign and no deadline pressure at this stage. + +- [ ] **Step 3: Keep the user mapping current** + +Each time someone joins, add their Plex→Jellyfin mapping to `USER_MAPPING` in +`stacks/streaming/stack.env`, then commit and push. Without this their history does not follow +them. + +- [ ] **Step 4: Let it run at least 4 weeks** + +This is a **hard gate on elapsed time**, not on effort. Jellystat needs a meaningful sample before +the Task 20 decision can be made on evidence rather than vibes. Both servers stay live throughout. + +--- + +### Task 20: Decide the transcoding hardware question + +**Files:** none — analysis and a decision. + +- [ ] **Step 1: Set up Jellystat** + +Browse to `https://jellystat.ginnoir.com` (Authentik will challenge you first, and it is LAN or +tailnet only). Create the admin account and connect it to Jellyfin at `http://jellyfin:8096` with +the API key from Task 15. + +> Do this at the **start** of the 4-week window in Task 19, not the end, or it will have nothing +> to report. + +- [ ] **Step 2: Pull the numbers after 4 weeks** + +From Jellystat, record: peak concurrent streams, the direct-play vs transcode ratio, and the most +common transcode reasons. + +- [ ] **Step 3: Compare against the Plex baseline** + +The measured Plex baseline was: **48% transcode, peak 7 concurrent, 1080p SDR, no subtitle burn-in, +no HDR tone-mapping**. + +Jellyfin's Android client direct-plays audio codecs Plex's refuses, and Android was the heaviest +platform at 281 plays — so the transcode rate may well have **fallen**. + +- [ ] **Step 4: Check whether the CPU actually struggled** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "uptime; sudo journalctl -u docker --since '4 weeks ago' | grep -ci 'jellyfin.*oom' || echo 'no OOM events'" +``` + +Also ask users directly whether playback ever stuttered or failed to start. User reports beat +graphs here. + +- [ ] **Step 5: Make and record the decision** + +- **If load is comfortable:** do nothing. Record that CPU-only transcoding was sufficient and + close this task. +- **If the CPU is straining:** buy an **Intel Arc A310 or A380** (~$100-150). No session limit, + AV1, full encode plus decode. This requires a PCIe slot, an R730XD shutdown, and an IOMMU + passthrough to the valhalla VM. **Check the new card's IOMMU group before buying** — only the + P100's group 7 is known clean. Write that up as its own plan; it is not a step in this one. + +--- + +# PHASE 6 — Cutover + +--- + +### Task 21: Convert Seerr from Plex to Jellyfin + +`seerr-team/seerr` supports Jellyfin natively, but converting an existing Plex-configured install +is **unsupported upstream** ([seerr#2522](https://github.com/seerr-team/seerr/issues/2522), open). +Hence the backup first and the documented fallback. + +**Files:** none — live operation. + +- [ ] **Step 1: Back up the Seerr configuration** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo tar czf /storage1/labdata/seerr-preJellyfin-$(date +%Y%m%d).tar.gz -C /config overseerr && ls -lh /storage1/labdata/seerr-preJellyfin-*.tar.gz" +``` + +Expected: a tarball listed with a non-trivial size. **Do not proceed without this.** + +- [ ] **Step 2: Attempt the in-place conversion** + +At `https://requests.ginnoir.com` → Settings → General, change the media server to **Jellyfin**: +- Hostname: `jellyfin` +- Port: `8096` +- API key: the key from Task 15 Step 1 +- Save, then Settings → Jellyfin → **Sync Libraries** and enable Movies, TV and Anime + +- [ ] **Step 3: Verify it works** + +Expected: libraries sync, existing media shows as **Available**, and a test request routes to +Radarr or Sonarr correctly. + +- [ ] **Step 4: If the conversion failed, use the fallback** + +Stop Seerr, move the config aside, and start fresh: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker stop seerr && sudo mv /config/overseerr /config/overseerr.plex-era && sudo mkdir -p /config/overseerr && sudo chown 1000:1000 /config/overseerr && docker start seerr" +``` + +Then reconfigure from scratch at `https://requests.ginnoir.com`: Jellyfin as the media server, +then re-add Radarr and Sonarr with their API keys. **Request history is lost** — this was an +accepted trade-off in the design. + +- [ ] **Step 5: Verify user login** + +Have one migrated user sign in to Seerr with their Jellyfin credentials and submit a request. +Expected: the request appears in Radarr or Sonarr. + +--- + +### Task 22: Identify holdouts and cut the Plex integrations + +**Files:** none — analysis plus UI operations. + +- [ ] **Step 1: Find who is still on Plex and on what** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla 'KEY=$(grep -E "^api_key" /config/tautulli/config.ini | head -1 | sed "s/.*= *//"); curl -s "http://localhost:8181/api/v2?apikey=$KEY&cmd=get_history&length=300" | python3 -c " +import sys,json,collections +d=json.load(sys.stdin)[\"response\"][\"data\"][\"data\"] +c=collections.Counter((r.get(\"friendly_name\"),r.get(\"platform\")) for r in d) +for k,v in c.most_common(30): print(f\"{v:5d} {k[0]:24s} {k[1]}\") +"' +``` + +Expected: a short list. Anyone with a meaningful play count is someone to contact in Step 2. + +- [ ] **Step 2: Contact each remaining active user individually** + +Expected pattern from the baseline: a small number of stragglers, plus **Tizen (Samsung TV)** and +**PlayStation** users — 3 plays out of 600 at plan time. + +- Tizen: point them at the Moonfin client, or a $30 Fire TV / Google TV stick +- PlayStation: there is **no native Jellyfin app**. Offer the same streaming-stick suggestion. + +- [ ] **Step 3: Announce the Plex end date** + +Give at least two weeks' notice to everyone still active on Plex. + +- [ ] **Step 4: Remove the Plex connections from the *arrs** + +- Sonarr → Settings → Connect → delete the **Plex** connection +- Radarr → Settings → Connect → delete the **Plex** connection +- Bazarr → Settings → disable the Plex integration + +Leave the Jellyfin connections from Task 15 in place. + +- [ ] **Step 5: Verify imports still trigger Jellyfin refreshes** + +Trigger a rename in Radarr, then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker logs --tail 30 jellyfin | grep -i 'refresh\|library'" +``` + +Expected: a refresh entry. Nothing should have broken by removing the Plex connections. + +--- + +# PHASE 7 — Decommission + +--- + +### Task 23: Stop Plex and soak + +**Files:** none — live operation. + +- [ ] **Step 1: Final watch-state sync while Plex is still running** + +This is the last chance to capture anything watched on Plex. JellyPlex-Watched needs the live PMS +API. + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker restart jellyplex-watched && sleep 180 && docker logs --tail 80 jellyplex-watched" +``` + +Expected: a clean sync with no errors. + +- [ ] **Step 2: Stop and disable the service** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo systemctl stop plexmediaserver && sudo systemctl disable plexmediaserver && systemctl is-active plexmediaserver; systemctl is-enabled plexmediaserver" +``` + +Expected: `inactive` and `disabled`. + +- [ ] **Step 3: Stop the sync container** + +With Plex down, JellyPlex-Watched has nothing to talk to and will log errors forever. + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker stop jellyplex-watched" +``` + +- [ ] **Step 4: Soak for two weeks** + +The package and all 383G of data remain on disk. **Rollback during this window is a single +command** (`sudo systemctl start plexmediaserver`). + +Do not proceed to Task 24 until two weeks have passed with no one asking for Plex back. + +--- + +### Task 24: Remove Plex and reclaim its storage + +**Files:** +- Modify: `stacks/streaming/docker-compose.yml` (remove `jellyplex-watched`) +- Modify: `stacks/media/docker-compose.yml` (remove `tautulli`) +- Modify: `Caddyfile` (remove the `tautulli` site block) + +- [ ] **Step 1: Archive the Plex data before removing anything** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo tar czf /storage1/labdata/plex-final-$(date +%Y%m%d).tar.gz -C /storage1/labdata plex && ls -lh /storage1/labdata/plex-final-*.tar.gz" +``` + +This takes a long time — 383G of mostly-incompressible thumbnails. Expected: a tarball of roughly +300G+. **Keep it 30 days.** Set a calendar reminder to delete it. + +- [ ] **Step 2: Remove the package** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo pacman -Rns --noconfirm plex-media-server && which plexmediaserver || echo 'plex removed'" +``` + +Expected: `plex removed`. + +- [ ] **Step 3: Reclaim the data directory** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h /storage1 | tail -1 && sudo rm -rf /storage1/labdata/plex && sudo rm -f /var/lib/plex && df -h /storage1 | tail -1" +``` + +Expected: `/storage1` available space increases by roughly **383G**. + +- [ ] **Step 4: Remove jellyplex-watched from the compose file** + +In `stacks/streaming/docker-compose.yml`, delete the entire `jellyplex-watched` service block. It +has no purpose without Plex. + +Also remove these now-dead lines from `stacks/streaming/stack.env`: + +``` +PLEX_URL, PLEX_TOKEN, JELLYFIN_BASEURL, JELLYFIN_TOKEN, +SYNC_FROM_PLEX_TO_JELLYFIN, SYNC_FROM_JELLYFIN_TO_PLEX, +DRYRUN, SLEEP_DURATION, USER_MAPPING +``` + +Keep `TZ`, `PUID`, `PGID`, and every `POSTGRES_*` and `JWT_SECRET` value. + +- [ ] **Step 5: Remove Tautulli** + +Tautulli is Plex-only and now monitors nothing. In `stacks/media/docker-compose.yml`, delete the +`tautulli` service block (lines 110-121 at plan time). In `Caddyfile`, delete the +`tautulli.ginnoir.com` block (lines 196-205 at plan time). + +- [ ] **Step 6: Validate both compose files and the Caddyfile** + +```bash +docker compose -f stacks/streaming/docker-compose.yml config --quiet && docker compose -f stacks/media/docker-compose.yml config --quiet && echo "COMPOSE OK" +docker run --rm -v "${PWD}/Caddyfile:/etc/caddy/Caddyfile:ro" caddy:latest caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile +``` + +Expected: `COMPOSE OK` and `Valid configuration`. + +- [ ] **Step 7: Commit and push** + +```bash +git add stacks/streaming/docker-compose.yml stacks/streaming/stack.env stacks/media/docker-compose.yml Caddyfile +git commit -m "chore: decommission plex + +Plex is removed from valhalla. Drops jellyplex-watched (nothing left to +sync from), tautulli (Plex-only) and the tautulli Caddy site." +git push origin main +``` + +- [ ] **Step 8: Verify the containers are gone** + +Wait 5 minutes for Portainer to poll, then: + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "docker ps -a --format '{{.Names}}' | grep -E 'tautulli|jellyplex' || echo 'both removed'" +``` + +Expected: `both removed`. If they linger, Portainer may need a manual redeploy of the affected +stack from the UI. + +- [ ] **Step 9: Remove the stale Tautulli config** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "sudo rm -rf /config/tautulli && echo removed" +``` + +--- + +### Task 25: Clean up documentation and monitoring + +**Files:** +- Modify: `CLAUDE.md` +- Modify: `plans/homelab-improvements/tasks/TB-027-jellyfin.md` +- Modify: `plans/homelab-improvements/tasks/TB-016-document-plex.md` +- Regenerate: `bookmarks-domains.html`, `bookmarks-ports.html` + +- [ ] **Step 1: Replace the Plex section in CLAUDE.md** + +Delete the entire `### Plex Media Server` block under **External services (not in compose)**. If +no other external service is documented there, remove that heading too. + +Add to the stack table in the **Layout** section: + +``` +│ ├── streaming/ # jellyfin, jellystat + postgres, wizarr +``` + +- [ ] **Step 2: Add a Jellyfin note to CLAUDE.md** + +Under **Known quirks / gotchas**: + +```markdown +- **Jellyfin storage is deliberately split.** SQLite config lives on the root SSD at + `/config/jellyfin`; metadata, cache, trickplay and transcodes bind to + `/storage1/labdata/jellyfin/*`. Trickplay needs an explicit sub-bind on + `/config/data/trickplay` because Jellyfin 10.11 moved it there and it cannot be + relocated from the UI. Do not "simplify" these into one mount — the root disk is + 194G and has filled twice. +- **Jellyfin has no SSO, on purpose.** `jellyfin-plugin-sso` only completes OIDC inside + a browser; Roku, Fire TV and tvOS clients cannot use it. Jellyfin uses native accounts. + This is a deliberate exception to the Authentik standard. +- **The P100 cannot encode.** GP100 has no NVENC. Hardware *decoding* is enabled for + trickplay generation; hardware encoding must stay off or every transcode fails. +``` + +- [ ] **Step 3: Close out the TB tasks** + +In `plans/homelab-improvements/tasks/TB-027-jellyfin.md`, fill in the decision table: + +```markdown +| **Decision** | Do it — completed | +| **Notes** | Migrated to Jellyfin in stacks/streaming. See docs/superpowers/specs/2026-08-05-jellyfin-migration-design.md and docs/superpowers/plans/2026-08-05-jellyfin-migration.md | +| **Date** | (fill in the completion date) | +``` + +In `plans/homelab-improvements/tasks/TB-016-document-plex.md`, mark it obsolete — Plex no longer +exists to document. + +- [ ] **Step 4: Update Uptime Kuma** + +In the Uptime Kuma UI: +- **Delete** the Tautulli monitor and any Plex monitor +- **Add** an HTTP monitor for `https://jellyfin.ginnoir.com/health`, expected status `200`, + interval 60s, with the ntfy notification attached +- **Keep** the `valhalla root disk` push monitor from Task 3 + +- [ ] **Step 5: Regenerate bookmarks** + +```bash +pwsh -File scripts/gen-bookmarks.ps1 +``` + +- [ ] **Step 6: Verify no stale references remain** + +```bash +grep -rni "tautulli\|plexmediaserver\|plex media server" . --include="*.md" --include="*.yml" --include="Caddyfile" | grep -v "docs/superpowers/" +``` + +Expected: **no output**. Matches inside `docs/superpowers/` are historical records of this +migration and should stay. + +- [ ] **Step 7: Commit and push** + +```bash +git add CLAUDE.md bookmarks-domains.html bookmarks-ports.html plans/homelab-improvements/tasks/ +git commit -m "docs: retire plex, document the streaming stack + +Plex is gone from valhalla. Documents the Jellyfin storage split, the +deliberate absence of SSO, and the P100's lack of NVENC. Closes TB-027 +and obsoletes TB-016." +git push origin main +``` + +- [ ] **Step 8: Final verification** + +```bash +ssh -o BatchMode=yes ginnoir@valhalla "df -h / /storage1 | grep -v Filesystem; docker ps --filter label=com.docker.compose.project=streaming --format '{{.Names}}\t{{.Status}}'" +curl -s -o /dev/null -w 'jellyfin: %{http_code}\n' https://jellyfin.ginnoir.com/health +``` + +Expected: root disk healthy, `/storage1` reclaimed by ~383G, four `streaming` containers `Up` +(jellyfin, jellystat, jellystat-db, wizarr), and Jellyfin returning `200`. + +--- + +## Rollback summary + +| Up to and including | How to roll back | +|---|---| +| Task 22 | Plex is running and untouched. Do nothing — users simply keep using it. | +| Task 23 | `sudo systemctl start plexmediaserver` | +| Task 24, within 30 days | `sudo pacman -S plex-media-server`, then extract `/storage1/labdata/plex-final-*.tar.gz` and recreate the `/var/lib/plex` symlink | +| After the tarball expires | No rollback. This is the point of no return. | + +## Deferred, deliberately + +- **Intel Arc GPU purchase and passthrough** — gated on Task 20's evidence, needs its own plan +- **Music and Books libraries** — out of scope per the design +- **Seerr request-history migration** — accepted loss +- **`/storage1` redundancy** — tracked separately