Skip to main content

Deployment

Production runs as six Docker containers behind a Cloudflare Tunnel. Nothing listens on a public port: all ingress arrives through the tunnel into nginx, which routes to the backend and frontend. The backend is the "front desk" only: every pipeline (ffmpeg, transcription, the AI) runs in a separate worker container, with Redis as the queue between them. This page covers the topology, the one-time server setup, and how to deploy and update.

Topology​

The six services​

Defined in docker-compose.prod.yml:

ServiceImage / buildExposesRole
backendbuilt from backend/Dockerfile (python:3.11-slim + ffmpeg)8000 (internal)FastAPI via uvicorn, WEB_CONCURRENCY (default 2) worker processes. Answers requests, streams progress, enqueues pipeline jobs. Never runs a pipeline. Has a /health healthcheck.
workerthe same imagenonepython -m arq app.worker.WorkerSettings. Pulls jobs off Redis and runs every pipeline (video, photo, audio, the extra photo pass, PDF import, Contents Capture import), WORKER_CONCURRENCY (default 4) at a time. Scale with --scale worker=N.
redisredis:7-alpine6379 (internal)The job queue (ARQ) and the mirror of live job logs and metrics the backend streams to the browser. Persists to its redis_data volume once a minute.
frontendbuilt from frontend/Dockerfile (node:20, standalone)3000 (internal)Next.js via node server.js. Waits for backend health.
nginxnginx:alpine80 (internal)Reverse proxy: /api/* and /docs to backend, everything else to frontend.
cloudflaredcloudflare/cloudflared:latestnoneConnects to Cloudflare and routes your domain to nginx:80.

Key details:

  • No host ports are published. Services use expose (container-to-container) only; cloudflared is the sole ingress.
  • Volumes: uploads and storage are named Docker volumes, mounted into both the backend and the worker: the worker reads the upload the backend saved and writes the frames and crops the backend serves. The SQLite database lives in storage and is shared by both too (same DATABASE_URL). redis_data holds the queue.
  • REDIS_URL is set on the backend and the worker in the compose file. It is the switch: with it set, every "start processing" in the app enqueues for the worker (app/core/queue.py); without it (local development, the test suite) pipelines run as in-process background tasks exactly as before. If Redis is configured but cannot be reached at that moment (a container still coming up after a deploy, a Docker DNS hiccup, Redis down), the connect gives up after about 3 s and the job runs in-process instead, so the user's click still works; the outage is logged at error level and sent to Sentry ("Queue unavailable, ran ... in-process"), because it means the worker split is silently off for that job. The same applies to the watchdog and purge locks (they run unlocked).
  • Healthcheck: the backend's GET /health gates the frontend's and the worker's startup (depends_on: condition: service_healthy); Redis has its own redis-cli ping check that gates both.
  • Several processes share one SQLite file. app/core/database.py sets journal_mode=WAL and a 15 s busy_timeout on every connection so readers never block the writer and a brief lock is waited out rather than failing. At boot, every process runs the schema check (init_db) under a lock file next to the database (storage/.init_db.lock, shared through the volume) so they take turns, and retries for up to 30 s if the file is busy anyway; the WAL switch itself is tolerant of another process getting there first. Without that, the first stage deploy had two of three processes crash on "database is locked" and "table already exists" in the same second. This is what makes the split workable on SQLite; the next step past it is Postgres (see what comes next).
  • Build args: the frontend bakes in NEXT_PUBLIC_APP_ENV (production / stage), the optional PostHog keys (NEXT_PUBLIC_POSTHOG_KEY / NEXT_PUBLIC_POSTHOG_HOST), and optional commit metadata, at build time. Because the PostHog keys are baked in, turning analytics on or off needs a frontend rebuild (see analytics). NEXT_PUBLIC_APP_ENV must be set per environment: the stage compose sets it to stage and prod to production. It is not cosmetic. It also gates the imgproxy thumbnail CDN, which is a production-only optimization. The CDN fetches each source photo from the origin of the page that asked for it (the site the browser is on), so contentsvision.com and each client's own instance (e.g. stage-sgk.contentsvision.com) point it at their own server with no extra setting. That host must be publicly reachable for the CDN to fetch from it. Stage serves thumbnails raw, exactly like local.

nginx​

nginx/nginx.conf is tuned for this workload:

  • client_max_body_size 2048m and long body timeouts for large video uploads.
  • Long proxy read/send timeouts and disabled buffering on /api/ so the SSE progress stream flows through immediately (X-Accel-Buffering: no equivalent via proxy_buffering off).
  • proxy_request_buffering off so large uploads stream straight to the backend.
  • Upstream retries so a brief container restart does not surface as a 502.

Cloudflare Tunnel​

The cloudflared service runs tunnel --no-autoupdate run --token ${CLOUDFLARE_TUNNEL_TOKEN}. Get the token from Cloudflare Zero Trust: Networks → Tunnels → Create tunnel → copy token, point the tunnel's public hostname at nginx:80, and put the token in .env as CLOUDFLARE_TUNNEL_TOKEN.

One-time server setup​

setup.sh provisions a fresh Ubuntu 22.04 / 24.04 server:

git clone https://bitbucket.org/adjustsquare/contentsvision.git
cd contentsvision
sudo bash setup.sh

It installs Docker and FFmpeg, configures the UFW firewall to allow only SSH (the tunnel handles all inbound web traffic), prompts for API keys and writes .env, then builds and starts the stack.

Deploying and updating​

# from the project root on the server
docker compose -f docker-compose.prod.yml up -d --build

To ship a new version:

git pull
docker compose -f docker-compose.prod.yml up -d --build

This one command builds and starts all six services, including redis and worker, and is the same command as before the worker split. Nothing else changes on the server: no new .env values are required (REDIS_URL is set inside the compose file). The first deploy after the split starts two new containers; check them with docker compose -f docker-compose.prod.yml ps and docker compose -f docker-compose.prod.yml logs -f worker (it prints Starting worker for 7 functions and then one line per job).

Because the database and uploaded files live in named volumes, a rebuild keeps all data. Schema changes are applied automatically at startup by init_db() (see data model); the backend runs them first and the worker waits for its health check before starting.

What happens to a job that is running when you deploy. The worker gets SIGTERM and 30 s of grace (stop_grace_period). ARQ cancels the running pipelines and puts them back on the queue; the new worker picks each up, resets it (partial items cleared, status back to pending, exactly like the Stage 1 requeue) and runs it from the start. Nothing is charged twice (the up-front charge is keyed on the job id) and the user sees "Restarting after an interruption". A job is re-run once; a job interrupted twice, or one whose row goes quiet for 30 minutes, is failed with a refund by the backend's stale-job watchdog (see observability).

Scaling the worker​

The worker is where the CPU and the AI calls go, so it is the only service that needs sizing.

  • WORKER_CONCURRENCY (default 4): pipelines one worker container runs at the same time. Most of a job's clock is spent waiting on the AI, so several overlap well; ffmpeg and image work run on the thread pool. On a small server (2 cores) leave it; on 8 cores or more try 6 to 8.
  • More containers: docker compose -f docker-compose.prod.yml up -d --scale worker=2. Each container has its own OPENROUTER_GLOBAL_CONCURRENCY (24) ceiling, so two containers can have 48 AI calls in flight; watch 429 retries on the admin API Requests page and lower OPENROUTER_PHOTO_CONCURRENCY or the global cap if they climb.
  • WEB_CONCURRENCY (default 2): uvicorn processes for the backend. Nothing heavy runs there any more, so 2 is plenty until measured otherwise. What is per-process (and so multiplied by this number) is written down in observability; the insured-link rate limits are the one thing that gets looser.
  • JOB_TIMEOUT_S (default 3600): one run of one job is killed after this and re-run once. Only raise it for jobs that legitimately take over an hour.
  • Queue length: docker compose -f docker-compose.prod.yml exec redis redis-cli zcard arq:queue. A queue that keeps growing means the worker is the bottleneck: scale it.

What comes next​

With the pipeline out of the web server, the next wall is the database: SQLite allows one writer at a time across all processes. WAL and the busy timeout make that workable, but at high job volume the pipelines' progress commits and the users' edits will queue on the same lock. The follow-up is moving DATABASE_URL to Postgres (adding a postgres service to the compose file and switching the URL; the models are plain SQLAlchemy); nothing in this step depends on it.

Environments​

The APP_ENV value (production, stage, or local) flows into the frontend as NEXT_PUBLIC_APP_ENV and drives the environment badge in the sidebar (hidden in production, amber for stage, violet for local) and the page title prefix. Staging and production are the same stack with different .env values and tunnels.

Operational checklist​

  • Confirm .env has the AI keys, ADJUSTSQUARE_JWT_SECRET, CLOUDFLARE_TUNNEL_TOKEN, and the Adjust Square base URL. See the configuration reference.
  • Verify GET /health returns {status: "ok"} (the healthcheck and a quick manual curl through the tunnel).
  • After every deploy, open https://<your-domain>/api/health/pipeline in a browser. It says whether jobs are actually running in the worker. "mode": "worker" with "worker_alive": true is healthy; "mode": "in-process-fallback" means REDIS_URL is set but Redis is unreachable, so the worker split is silently OFF (check the redis and worker containers); "mode": "in-process" means REDIS_URL is not set at all. This needs no server access, so it is the fastest first check.
  • Watch docker compose logs -f backend worker during the first requests after a deploy; the first upload should print Enqueued ARQ job: run_pipeline on the backend and the pipeline lines on the worker.
  • Back up the storage volume (it holds the SQLite database and all per-job files). redis_data only holds queued and recently finished jobs; losing it loses at most the jobs that were waiting in the queue (their rows stay pending and the watchdog fails them with a refund after 30 minutes).