Deployment
Production runs as six Docker containers behind a Cloudflare Tunnel. Nothing listens on a public port: all ingress arrives through the tunnel into nginx, which routes to the backend and frontend. The backend is the "front desk" only: every pipeline (ffmpeg, transcription, the AI) runs in a separate worker container, with Redis as the queue between them. This page covers the topology, the one-time server setup, and how to deploy and update.
Topology
The six services
Defined in docker-compose.prod.yml:
| Service | Image / build | Exposes | Role |
|---|---|---|---|
backend | built from backend/Dockerfile (python:3.11-slim + ffmpeg) | 8000 (internal) | FastAPI via uvicorn, WEB_CONCURRENCY (default 2) worker processes. Answers requests, streams progress, enqueues pipeline jobs. Never runs a pipeline. Has a /health healthcheck. |
worker | the same image | none | python -m arq app.worker.WorkerSettings. Pulls jobs off Redis and runs every pipeline (video, photo, audio, the extra photo pass, PDF import, Contents Capture import), WORKER_CONCURRENCY (default 4) at a time. Scale with --scale worker=N. |
redis | redis:7-alpine | 6379 (internal) | The job queue (ARQ) and the mirror of live job logs and metrics the backend streams to the browser. Persists to its redis_data volume once a minute. |
frontend | built from frontend/Dockerfile (node:20, standalone) | 3000 (internal) | Next.js via node server.js. Waits for backend health. |
nginx | nginx:alpine | 80 (internal) | Reverse proxy: /api/* and /docs to backend, everything else to frontend. |
cloudflared | cloudflare/cloudflared:latest | none | Connects to Cloudflare and routes your domain to nginx:80. |
Key details:
- No host ports are published. Services use
expose(container-to-container) only;cloudflaredis the sole ingress. - Volumes:
uploadsandstorageare named Docker volumes, mounted into both the backend and the worker: the worker reads the upload the backend saved and writes the frames and crops the backend serves. The SQLite database lives instorageand is shared by both too (sameDATABASE_URL).redis_dataholds the queue. REDIS_URLis set on the backend and the worker in the compose file. It is the switch: with it set, every "start processing" in the app enqueues for the worker (app/core/queue.py); without it (local development, the test suite) pipelines run as in-process background tasks exactly as before. If Redis is configured but cannot be reached at that moment (a container still coming up after a deploy, a Docker DNS hiccup, Redis down), the connect gives up after about 3 s and the job runs in-process instead, so the user's click still works; the outage is logged at error level and sent to Sentry ("Queue unavailable, ran ... in-process"), because it means the worker split is silently off for that job. The same applies to the watchdog and purge locks (they run unlocked).- Healthcheck: the backend's
GET /healthgates the frontend's and the worker's startup (depends_on: condition: service_healthy); Redis has its ownredis-cli pingcheck that gates both. - Several processes share one SQLite file.
app/core/database.pysetsjournal_mode=WALand a 15 sbusy_timeouton every connection so readers never block the writer and a brief lock is waited out rather than failing. At boot, every process runs the schema check (init_db) under a lock file next to the database (storage/.init_db.lock, shared through the volume) so they take turns, and retries for up to 30 s if the file is busy anyway; the WAL switch itself is tolerant of another process getting there first. Without that, the first stage deploy had two of three processes crash on "database is locked" and "table already exists" in the same second. This is what makes the split workable on SQLite; the next step past it is Postgres (see what comes next). - Build args: the frontend bakes in
NEXT_PUBLIC_APP_ENV(production / stage), the optional PostHog keys (NEXT_PUBLIC_POSTHOG_KEY/NEXT_PUBLIC_POSTHOG_HOST), and optional commit metadata, at build time. Because the PostHog keys are baked in, turning analytics on or off needs a frontend rebuild (see analytics).NEXT_PUBLIC_APP_ENVmust be set per environment: the stage compose sets it tostageand prod toproduction. It is not cosmetic. It also gates the imgproxy thumbnail CDN, which is a production-only optimization. The CDN fetches each source photo from the origin of the page that asked for it (the site the browser is on), so contentsvision.com and each client's own instance (e.g.stage-sgk.contentsvision.com) point it at their own server with no extra setting. That host must be publicly reachable for the CDN to fetch from it. Stage serves thumbnails raw, exactly like local.
nginx
nginx/nginx.conf is tuned for this workload:
client_max_body_size 2048mand long body timeouts for large video uploads.- Long proxy read/send timeouts and disabled buffering on
/api/so the SSE progress stream flows through immediately (X-Accel-Buffering: noequivalent viaproxy_buffering off). proxy_request_buffering offso large uploads stream straight to the backend.- Upstream retries so a brief container restart does not surface as a 502.
Cloudflare Tunnel
The cloudflared service runs tunnel --no-autoupdate run --token ${CLOUDFLARE_TUNNEL_TOKEN}. Get the token from Cloudflare Zero Trust: Networks → Tunnels → Create tunnel → copy token, point the tunnel's public hostname at nginx:80, and put the token in .env as CLOUDFLARE_TUNNEL_TOKEN.
One-time server setup
setup.sh provisions a fresh Ubuntu 22.04 / 24.04 server:
git clone https://bitbucket.org/adjustsquare/contentsvision.git
cd contentsvision
sudo bash setup.sh
It installs Docker and FFmpeg, configures the UFW firewall to allow only SSH (the tunnel handles all inbound web traffic), prompts for API keys and writes .env, then builds and starts the stack.
Deploying and updating
# from the project root on the server
docker compose -f docker-compose.prod.yml up -d --build
To ship a new version:
git pull
docker compose -f docker-compose.prod.yml up -d --build
This one command builds and starts all six services, including redis and worker, and is the same command as before the worker split. Nothing else changes on the server: no new .env values are required (REDIS_URL is set inside the compose file). The first deploy after the split starts two new containers; check them with docker compose -f docker-compose.prod.yml ps and docker compose -f docker-compose.prod.yml logs -f worker (it prints Starting worker for 7 functions and then one line per job).
Because the database and uploaded files live in named volumes, a rebuild keeps all data. Schema changes are applied automatically at startup by init_db() (see data model); the backend runs them first and the worker waits for its health check before starting.
What happens to a job that is running when you deploy. The worker gets SIGTERM and 30 s of grace (stop_grace_period). ARQ cancels the running pipelines and puts them back on the queue; the new worker picks each up, resets it (partial items cleared, status back to pending, exactly like the Stage 1 requeue) and runs it from the start. Nothing is charged twice (the up-front charge is keyed on the job id) and the user sees "Restarting after an interruption". A job is re-run once; a job interrupted twice, or one whose row goes quiet for 30 minutes, is failed with a refund by the backend's stale-job watchdog (see observability).
Scaling the worker
The worker is where the CPU and the AI calls go, so it is the only service that needs sizing.
WORKER_CONCURRENCY(default 4): pipelines one worker container runs at the same time. Most of a job's clock is spent waiting on the AI, so several overlap well; ffmpeg and image work run on the thread pool. On a small server (2 cores) leave it; on 8 cores or more try 6 to 8.- More containers:
docker compose -f docker-compose.prod.yml up -d --scale worker=2. Each container has its ownOPENROUTER_GLOBAL_CONCURRENCY(24) ceiling, so two containers can have 48 AI calls in flight; watch 429 retries on the admin API Requests page and lowerOPENROUTER_PHOTO_CONCURRENCYor the global cap if they climb. WEB_CONCURRENCY(default 2): uvicorn processes for the backend. Nothing heavy runs there any more, so 2 is plenty until measured otherwise. What is per-process (and so multiplied by this number) is written down in observability; the insured-link rate limits are the one thing that gets looser.JOB_TIMEOUT_S(default 3600): one run of one job is killed after this and re-run once. Only raise it for jobs that legitimately take over an hour.- Queue length:
docker compose -f docker-compose.prod.yml exec redis redis-cli zcard arq:queue. A queue that keeps growing means the worker is the bottleneck: scale it.
What comes next
With the pipeline out of the web server, the next wall is the database: SQLite allows one writer at a time across all processes. WAL and the busy timeout make that workable, but at high job volume the pipelines' progress commits and the users' edits will queue on the same lock. The follow-up is moving DATABASE_URL to Postgres (adding a postgres service to the compose file and switching the URL; the models are plain SQLAlchemy); nothing in this step depends on it.
Environments
The APP_ENV value (production, stage, or local) flows into the frontend as NEXT_PUBLIC_APP_ENV and drives the environment badge in the sidebar (hidden in production, amber for stage, violet for local) and the page title prefix. Staging and production are the same stack with different .env values and tunnels.
Operational checklist
- Confirm
.envhas the AI keys,ADJUSTSQUARE_JWT_SECRET,CLOUDFLARE_TUNNEL_TOKEN, and the Adjust Square base URL. See the configuration reference. - Verify
GET /healthreturns{status: "ok"}(the healthcheck and a quick manual curl through the tunnel). - After every deploy, open
https://<your-domain>/api/health/pipelinein a browser. It says whether jobs are actually running in the worker."mode": "worker"with"worker_alive": trueis healthy;"mode": "in-process-fallback"meansREDIS_URLis set but Redis is unreachable, so the worker split is silently OFF (check theredisandworkercontainers);"mode": "in-process"meansREDIS_URLis not set at all. This needs no server access, so it is the fastest first check. - Watch
docker compose logs -f backend workerduring the first requests after a deploy; the first upload should printEnqueued ARQ job: run_pipelineon the backend and the pipeline lines on the worker. - Back up the
storagevolume (it holds the SQLite database and all per-job files).redis_dataonly holds queued and recently finished jobs; losing it loses at most the jobs that were waiting in the queue (their rows staypendingand the watchdog fails them with a refund after 30 minutes).