Troubleshooting runbook
A symptom-first guide for support and operations. Each entry lists what the user sees, the likely cause, and where to look or what to fix.
Sign-in and access
Everyone is suddenly logged out / every request returns 401
- In production:
ADJUSTSQUARE_JWT_SECRETis unset or wrong, so no token validates. The backend logs a loudauth.secret_missingerror on startup. - In local mode: check that
APP_ENV=localandLOCAL_AUTH_USER_EMAILare set correctly in.env. The backend logsauth.local_bypasswhen bypass is active. - Where to look: backend logs for
auth.jwt_error,auth.secret_missing, orauth.local_bypass.
A user sees "Sign-in failed" on the callback page
- Likely cause: the token did not validate (
/api/auth/mereturned 401), often a secret mismatch or an expired token. - Where to look: the callback page shows the error detail outside production; backend logs show the reason.
A user cannot see a claim they expect (403 / not found)
- Likely cause: the claim belongs to a different team. Access is team-scoped.
- Where to look: confirm the user's team (
/api/auth/me) matches the claim's team. This is by design, not a bug.
Uploads and processing
A large video upload fails, hangs, or never creates a job
- Likely cause: the file exceeds a CDN/tunnel request-body cap. Cloudflare's tunnel rejects bodies over ~100 MB on Free/Pro, well below the app's
MAX_UPLOAD_SIZE_MB(2 GB). If no job appears in/api/jobs, the upload never reached the backend. - Mitigation (built in): the frontend uploads video and audio in small resumable chunks (
POST /api/uploads/init, thenPUT .../chunks/{i}, thencomplete), so each request stays under the cap. Confirm the deployed frontend is current. See the chunked upload protocol. - If it still fails: test a small clip to isolate size; check
docker compose logs nginx backendfor413; verifyMAX_UPLOAD_SIZE_MBand nginxclient_max_body_sizeboth allow the size.
A job is stuck or "frozen"
- Check the worker first (production). Pipelines run in the
workercontainer, not the backend.docker compose -f docker-compose.prod.yml psshould show itUp;docker compose -f docker-compose.prod.yml logs -f workershows every job start and finish. A worker that is down leaves jobs atpending("Queued for AI analysis") with the queue growing (exec redis redis-cli zcard arq:queue);up -d workerrestarts it and it drains the queue. A worker that is up but idle while jobs sit atpendingmeans the backend could not reach Redis: look forEnqueued ARQ jobon the backend logs (present means the queue got it) andredis-cli ping. - Self-healing (already in place): a restart used to freeze in-flight jobs at their last progress forever. Now
services/job_recovery.pyand the worker fix this automatically. Locally (noREDIS_URL), on boot every job still in a processing status is requeued and restarts from scratch (partial items are cleared first; the up-front charge is keyed on the job id, so no double-billing), and a watchdog sweeps every minute for jobs whose row has been quiet forSTUCK_JOB_MINUTES(default 15). In production, a worker restart (a deploy) makes ARQ re-run the cancelled job once after the same reset, and the backend's watchdog fails (with a refund) any job whose row has been quiet for 30 minutes, since a live pipeline commits far more often than that. A job only gets one automatic retry; a second interruption fails it with a clear message, refunds the charge, and posts the failure to Slack. Frame extraction used to be silent for its whole run, so a very long 4K video on a slow server could be failed mid-extraction; it now logs and commits a heartbeat everyFRAME_EXTRACTION_HEARTBEAT_SECONDS(10), so only a genuinely dead job goes quiet. - Where to look: the job status and the Logs page (
/logs, admins only; system-wide with team/user filters), orGET /api/jobs/{id}/logs. Finished jobs keep their logs in thejob_log_archivestable, so the detail survives restarts. Map the status to a stage using the table in the user journey. A job that says "Restarting after an interruption" was recovered by the sweep (recovery_attemptson the job row counts the retries). - Likely causes by stage:
analyzingfor a long time: this is the slow stage; large videos take minutes. Check for OpenRouter rate-limit retries in the logs.transcribing: Deepgram is slow or the audio is long.
Is the pipeline worker actually running? (/api/health/pipeline)
Open https://<your-domain>/api/health/pipeline in a browser. No login or server access needed (it routes through the tunnel to the backend). It reports mode: worker (healthy, jobs run in the worker; check worker_alive is true and queued_jobs is not climbing), in-process-fallback (problem: REDIS_URL is set but Redis is unreachable, so jobs run in the web server and the isolation is off), or in-process (REDIS_URL not set, expected only off production). This is the fastest way to confirm a deploy worked.
Sentry says "Queue unavailable, ran ... in-process"
- Meaning: the backend could not reach Redis when a user started a job, so it ran the pipeline inside the web server (the local-development path) rather than fail the click. The job itself is fine. The isolation is off for that one job. Confirm with
/api/health/pipeline(above). - Where to look:
docker compose -f docker-compose.prod.yml ps(isredisup and healthy?),docker compose -f docker-compose.prod.yml logs redis backend. A single event right after a deploy is the containers coming up in a different order than usual; a stream of them means Redis is down or the backend cannot resolve theredishostname (both must be on the same compose project and network).
A backend or worker process crashed at boot with "database is locked" or "table ... already exists"
- Meaning: two or more processes ran the schema check on the same SQLite file at the same moment.
init_dbnow serializes them with a lock file and retries for up to 30 s, so this should not recur; if it does, check that thestoragevolume is a local filesystem (file locks do not work over NFS) and thatstorage/.init_db.lockis writable.
A job ended in error
- Where to look: the job's
error_messageand the pipeline logs (the failing stage is logged as[STAGE FAIL]). - Common causes: FFmpeg missing or a corrupt upload (extract stage), an unreadable image (photo upload), or an AI call exhausting its retries.
A job failed with JSONDecodeError from an AI call
- Likely cause: OpenRouter returned a
200with no completion in the body. It pads a non-streaming connection with whitespace while the model works, and occasionally the padding arrives alone. On a chunked job (audio or long video) this hits one chunk while its siblings succeed, so the same file usually processes fine on a retry, which makes it look file-specific when it is not. - Fixed in AS-964: an unreadable body is now retried on the 5xx backoff instead of failing the job on the first attempt. A single dud reply costs a few seconds, not the upload.
- Where to look: the pipeline log warning
AI service returned an unreadable body, which quotes the body, plus theapi_request_logsrow for the job. Repeated occurrences across many jobs point at OpenRouter or the network path, not the app.
Photos uploaded but nothing is processing
- Likely cause: photo jobs wait in
awaiting_grouping. The user must group photos and press Process; AI does not start automatically. - Fix: finish grouping in the modal and start processing. See the photo pipeline.
"Upload failed" when uploading many or large photos
- Likely cause: historically the frontend sent every selected photo in one request, and the Cloudflare Tunnel caps a single request body at roughly 100MB. A big set of camera JPEGs blew past that and Cloudflare returned an HTML error page the UI could not parse, so it showed a bare "Upload failed".
- Fix (already in place): photos now upload in size-bounded batches (80MB budget), the first to
POST /api/upload-photosand the rest toPOST /api/upload-photos/{id}/append, all on one job. A single photo over the 100MB hard limit is caught in the browser with a clear message. See batched photo upload.
The inventory is "visual only" with no spoken details
- Likely cause: no
DEEPGRAM_API_KEY, so transcription was skipped (the pipeline logs a warning and proceeds visual-only). - Fix: set the Deepgram key.
Progress bar never updates in the UI
- Likely cause: the SSE stream is being buffered. In production this is handled by nginx (
proxy_buffering off); in dev the local SSE route handler owns the stream. - Where to look: confirm
nginx.confis the deployed one and thecloudflared/nginxcontainers are healthy. See deployment.
Output quality
Items look wrong, duplicated, or missing
- Likely causes:
- Quiet or unclear narration: transcription drives naming. Encourage clear narration.
- Photos not grouped: ungrouped same-item shots can become separate items.
- Low-confidence items never reviewed: items below 0.6 arrive flagged and need a human.
- Where to look: the item's confidence and evidence fields, and the job transcript.
Structural items appear (or expected items are excluded)
- By design: the AI is contents-only. It excludes walls, floors, built-in plumbing, and building systems, and includes movable property, cabinets, and appliances. A team can nudge behavior with a custom prompt add-on.
Duplicate items across a chunk boundary in a video
- Mechanism: cross-chunk dedup merges near-boundary duplicates within
VIDEO_DEDUP_WINDOW_SECONDS. Genuine duplicates far apart in time are not merged. Reviewers can merge them manually.
Submission to Adjust Square
"Your team has no Adjust Square API token" (409)
- Likely cause: with per-user auth, the team's
team_tokenis missing (it is set from the SSO login token). - Fix: ensure the user logged in via SSO so the token was captured; confirm the team record has a
team_token. See the Adjust Square integration.
Submission returns 502
- Likely cause: Adjust Square rejected the request or was unreachable.
- Where to look: backend logs for the
AdjustSquareErrorstatus and body. ConfirmADJUST_SQUARE_BASE_URLis set and reachable.
Some items submitted, some did not (re-submission)
- By design: re-submission posts items one at a time; a failed item stays unsubmitted so it can be retried, while the rest proceed. Re-run the submission for the stragglers.
Photos did not appear on the Adjust Square claim
- Likely cause: photo upload happens in a background task after the items are created; a failure there is logged but does not fail the submission.
- Where to look: backend logs for photo upload warnings. Re-submitting re-attempts only unsubmitted items, so badly-failed photos may need manual attention.
Sending to Contents Estimation
No photos arrived on the Contents Estimation claim (the send said it worked)
- Check this first: open
GET /api/auth/debugwhile signed in and readcontents_estimation.public_base_url. If it isnull, that is the whole answer. - Why: the intake takes photos as URLs, not bytes, and those URLs are built by prefixing each stored
/api/photos/...path withPUBLIC_BASE_URL. With it unset there is no URL to send, so the claim goes across with zero photos and still reports success. Nothing is retried, because nothing failed. - Fix: set
PUBLIC_BASE_URLin the server's.envto this app's public origin (https://contentsvision.com) and restart the backend. It must be the externally reachable origin:localhostor an internal hostname is useless to the other side, which fetches these URLs itself. - Where to look: the backend logs a
WARNINGnamingPUBLIC_BASE_URLon any send whose items have photos but which carried none. The send response'sphotos_included(shown on the modal's success screen) is the same signal from the UI. - If
public_base_urlIS set: the URLs went out and the fetch is failing on their side. Confirm the origin ishttpsand reachable from outside the network, then ask Contents Estimation to check their fetch logs. The photo route itself is deliberately unauthenticated (serve_photoinapi/routes/inventory.py), so it is fetchable by design.
Re-sending a claim created a duplicate on the Contents Estimation side
- Likely cause: something changed what goes into
external_reference. It is the dedupe key and must stay the full ContentsVision claim id. See Idempotent, one-shot.
Deployment and ops
A deploy comes up unhealthy
- Where to look:
docker compose -f docker-compose.prod.yml logs -f backend. The frontend waits for the backend's/healthto pass. - Common causes: a missing required env var (AI key, JWT secret) or a bad
.env.
After adding a model column, production errors on that column
- Cause:
create_allonly creates whole new tables; existing tables need anALTER TABLEline ininit_db(). - Fix: add the
ALTER TABLE ... ADD COLUMNstatement. See data model.
Data loss after a rebuild
- Cause: the SQLite database and per-job files live in the
storage(anduploads) Docker volumes. If those volumes were removed, data is gone. - Prevention: back up the
storagevolume regularly.
Where the signals live
| Signal | Where |
|---|---|
| Live job progress and logs | /logs page, GET /api/jobs/{id}/stream, GET /api/jobs/{id}/logs |
| Job timings, tokens, cost | GET /api/jobs/{id}/metrics |
| Durable traces, metrics, and logs | Grafana (Tempo / Prometheus / Loki) when telemetry is enabled |
| Who changed what (and undo) | Settings activity log, GET /api/activity |
| Backend errors | docker compose logs backend / uvicorn output |
| Pipeline errors and per-job progress lines | docker compose logs worker (production); the backend log when running without REDIS_URL |
| Queue depth | docker compose exec redis redis-cli zcard arq:queue |
| Liveness | GET /health |
Reconstructing a past job after a restart
The in-memory /logs buffer (and its Redis mirror, kept 24 hours) is lost on restart, but finished jobs keep their lines in job_log_archives, and when telemetry is enabled the full history persists in Grafana. Open the job.process trace in Tempo (search by job.id or claim.id), read the stage timeline, and follow the trace id into Loki for the log lines. Step-by-step queries are in the observability runbook.