Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
69 changes: 69 additions & 0 deletions .github/workflows/docker-build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -70,6 +70,75 @@ jobs:
cache-from: type=gha
cache-to: type=gha,mode=max

# ── Does the image actually RUN? ────────────────────────────────────────
#
# A green build says the layers assembled. It says nothing about whether the thing
# starts. We published an image that built perfectly and then crash-looped on its
# first run, because a fresh named volume is created root-owned while the container
# runs as uid 10001:
#
# PermissionError: [Errno 13] Permission denied: '/app/cache/_roformer-models'
#
# That is exactly what `docker compose up -d` does, and no amount of building would
# have caught it. So: start the container the way a user does, and require /health
# to answer.
#
# The FRESH NAMED VOLUME is the load-bearing part of this test. Without it the
# container writes to the image's own filesystem, the permission bug never fires,
# and this test would happily pass on the very bug it exists to catch.
#
# SKIP_WARMUP=true: /health binds immediately and answers while the models load;
# warming up would pull ~1.5 GB of weights we don't need in order to answer
# "did the process survive its first write?".
- name: Smoke test — the container must start and serve /health
run: |
set -euo pipefail
docker volume rm -f smoke-cache >/dev/null 2>&1 || true
docker run -d --name smoke \
-e SKIP_WARMUP=true \
-v smoke-cache:/app/cache \
-p 7865:7865 \
"${{ steps.tag.outputs.scan }}"

ok=""
for i in $(seq 1 60); do
if curl -fsS --max-time 3 http://127.0.0.1:7865/health >/dev/null 2>&1; then
ok=1; break
fi
# Fail fast rather than burning two minutes on a container that is already dead
# (or, worse, on one that `restart: unless-stopped` is bouncing in a loop).
state="$(docker inspect -f '{{.State.Status}}' smoke 2>/dev/null || echo gone)"
if [ "$state" != "running" ]; then
echo "::error::container is '$state' - it did not stay up"
break
fi
sleep 2
done

echo "--- container logs ---"
docker logs --tail 60 smoke 2>&1 || true
echo "----------------------"

if [ -z "$ok" ]; then
echo "::error::/health never answered. The image builds but does not run."
docker rm -f smoke >/dev/null 2>&1 || true
docker volume rm -f smoke-cache >/dev/null 2>&1 || true
exit 1
fi

echo "/health responded:"
curl -fsS --max-time 5 http://127.0.0.1:7865/health

# Prove the cache is actually WRITABLE as the unprivileged runtime user - the
# precise thing that was broken. /health answering is necessary but not
# sufficient: the server only touches the cache once a job arrives.
docker exec smoke sh -c 'touch /app/cache/.write-probe && rm /app/cache/.write-probe' \
|| { echo "::error::/app/cache is not writable by the runtime user"; exit 1; }
echo "cache dir is writable by the runtime user"

docker rm -f smoke >/dev/null 2>&1 || true
docker volume rm -f smoke-cache >/dev/null 2>&1 || true

- name: Create output directory
run: mkdir -p sbom/

Expand Down
13 changes: 13 additions & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -99,7 +99,20 @@ RUN out="$(pip install --no-cache-dir --no-deps --only-binary=:all: 'diffq-fixed
COPY . .

# ---- Least privilege runtime user ----
#
# /app/cache MUST exist in the image, owned by appuser, BEFORE the VOLUME below.
#
# Docker initializes a fresh named volume from whatever sits at that mountpoint in the
# image — ownership included. With nothing there, it creates the directory as root:root.
# The container runs as appuser (uid 10001), so the server dies on its first write:
#
# PermissionError: [Errno 13] Permission denied: '/app/cache/_roformer-models'
#
# ...and `restart: unless-stopped` turns that into a crash-loop. Not theoretical: it is
# what the freshly published image did on its very first run, and it breaks the documented
# `docker compose up -d` path exactly as hard as anything else.
RUN useradd --create-home --uid 10001 appuser \
&& mkdir -p /app/cache \
&& chown -R appuser:appuser /app


Expand Down
26 changes: 25 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,10 +190,34 @@ docker run --gpus all -p 7865:7865 slopsmith-demucs-server
# Pull from GHCR and run (CPU)
docker compose up -d

# GPU mode: uncomment runtime: nvidia + NVIDIA_* env vars in compose file
# GPU mode: uncomment `gpus: all` in the compose file (needs nvidia-container-toolkit;
# Linux or Windows/WSL2 only — macOS cannot pass a GPU through at all)
docker compose up -d
```

> #### ⚠️ Ran an image from before 2026-07-12? Delete the cache volume.
>
> Early images created their model-cache volume **owned by root**, while the server runs as
> an unprivileged user (uid 10001). The container would start and then immediately die with
>
> ```
> PermissionError: [Errno 13] Permission denied: '/app/cache/_roformer-models'
> ```
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Outdated
>
> ...which `restart: unless-stopped` turns into a crash-loop.
>
> The image is fixed — but **pulling the new image is not enough on its own**. Docker sets a
> volume's ownership only when it *first creates* it, so a volume made by an older image stays
> root-owned forever and will keep crash-looping on a perfectly good image. Remove it:
>
> ```bash
> docker compose down
> docker volume rm feedback-demucs-cache # or: slopsmith-demucs-server_demucs-cache
> docker compose up -d
> ```
>
> The volume only holds cached model weights — deleting it costs you a re-download, nothing else.

### Persistent model cache

Model weights are stored in `/app/cache` inside the container. The compose file maps this to a persistent volume so weights survive restarts:
Expand Down
Loading