The AI pipeline
The pipeline is the heart of ContentsVision. It takes an uploaded file and produces inventory items. There are three pipelines, one per media type, all orchestrated from backend/app/services/pipeline.py and all sharing the same supporting services. This page traces each one stage by stage.
pipeline.py exposes four entry points: run_pipeline (video), run_grouped_photo_pipeline (photos, the live path), run_audio_pipeline (audio), and the older run_photo_pipeline (legacy batched photos), plus run_read_the_rest_pipeline for the extra paid photo pass. The upload and process routes start them through app/core/queue.enqueue_job: on the ARQ worker in production (REDIS_URL set, see deployment), as a FastAPI background task in local development.
The shared shape
Every pipeline follows the same rhythm: prepare the input, run AI analysis (in parallel where possible), then synthesize and save items. Each stage is wrapped in a stage_timer from observability so timings and token usage are recorded, and each stage updates the job's status, progress_pct, and current_stage so the UI can stream progress.
Tunable knobs
The pipeline's behavior is controlled by settings in backend/app/core/config.py (all environment-overridable). The defaults below are tuned together; the configuration reference lists every variable.
| Setting | Default | Effect |
|---|---|---|
frame_extraction_fps | 1.0 | Frames pulled from video per second. |
frame_long_edge_px | 1600 | Frames are shrunk inside ffmpeg so their long edge is at most this (never upscaled). A 4K frame lands at 1600x900, about 80 KB instead of 1.7 MB. 0 keeps full resolution. |
frame_jpeg_quality | 4 | ffmpeg MJPEG quality scale for the frame files (2 best, 31 worst). |
keyframe_decode_min_ratio | 0.8 | Decode only keyframes (-skip_frame nokey) when the video's keyframe rate is at least this fraction of frame_extraction_fps. Phones write about one keyframe a second, so a phone walkthrough qualifies and decodes about four times faster; an edited export with a long GOP falls back to a full decode. 0 turns the shortcut off. |
frame_extraction_heartbeat_seconds | 10 | While ffmpeg writes frames, the pipeline logs and commits a heartbeat this often (frames written so far), so the progress bar moves and the stuck-job watchdog never mistakes a long extraction for a hung job. |
chunk_duration_seconds | 30.0 | Length of each time chunk analyzed by one AI call. |
frames_per_minute | 60 | Target frame density used when selecting frames per chunk. |
max_frames_per_call | 60 | Hard cap on frames sent to a single vision call. |
video_dedup_window_seconds | 3.0 | How close in time two same-named items must be to merge across a chunk boundary. |
max_photos_per_call | 8 | Photos per call in the legacy batched photo pipeline. |
photo_image_max_pixels | 1150000 | Photos are shrunk to at most this many pixels (Anthropic's 1.15 megapixel ceiling, ~1,530 tokens for a 4:3 photo) before the vision call. Was a 768px long edge (~590 tokens), at which items in a crowded photo were 20-50px wide and got named wrong. |
photo_image_max_long_edge | 1568 | Photos are also never wider than this on the long side, so a panorama cannot blow past the token estimate. |
video_frame_long_edge | 768 | Video frames stay at this long edge. A chunk sends up to 60 frames; at 1.15 MP they would cost about $0.46 a chunk against $0.60 earned. |
DEFAULT_MIN_IMAGE_PX | 80 | Smallest embedded image in a PDF treated as an item photo rather than page furniture (pdf_extractor.py). |
MIN_SIGNATURE_LINES | 2 | Repeating lines a PDF must have before its learned recipe is cached and reused (pdf_import.py). |
MIN_SIGNATURE_OVERLAP | 0.6 | How much two documents' page furniture must agree before one is read as the same vendor format as the other (pdf_import.py). |
MIN_SHARED_SIGNATURE_LINES | 3 | Furniture lines two signatures must share before that overlap is even considered (pdf_import.py). |
MAX_PAGES / MAX_IMAGES | 2000 / 12000 | Ceiling on what one PDF may cost, alongside a 1 GiB extracted-image budget (pdf_extractor.py). |
MAX_PHOTOS_PER_ITEM | 5 | Hard cap on photos in one group. Enforced on the regroup and segment endpoints, and by pdf_import._assign_groups when a document prints more than this under one label. |
photo_budget_credit_rate_usd | 0.10 | The per-credit rate a photo request's output budget is priced at (the lowest rate any customer is on). |
photo_budget_target_margin | 0.8 | Share of that credit one pass may spend. |
photo_budget_min_output_tokens / photo_budget_max_output_tokens | 800 / 3500 | Floor and ceiling on a photo request's output budget. |
photo_budget_truncation_ratio | 0.95 | Share of its budget an answer must use to count as cut short without a finish_reason=length. |
photo_max_passes | 3 | Passes allowed over one photo group: the first call plus two more. |
openrouter_concurrency | 6 | Max simultaneous AI calls per job for video and audio jobs (one job's own semaphore). |
openrouter_photo_concurrency | 12 | Max simultaneous AI calls per job for photo jobs (run_grouped_photo_pipeline, run_read_the_rest_pipeline, and the legacy batched pipeline). Higher than video because a photo call carries one to five images and a short answer, where a video chunk carries up to 60 frames. Watch 429 retries on the admin API Requests page after raising it. |
openrouter_global_concurrency | 24 | Max simultaneous AI calls across all in-flight jobs in the process. A shared limiter (in claude_analyzer) wrapping every OpenRouter call, so several concurrent jobs share provider capacity smoothly instead of each firing its per-job cap and tripping 429s. Keep at or below the httpx pool in claude_analyzer._get_http_client (32). |
openrouter_model | anthropic/claude-opus-4-7 | The vision and analysis model. Also what /health reports. |
openrouter_temperature | 0 | Sampling temperature, sent only on models that accept one (Opus 4.5 and earlier). Pinned to 0 for repeatable results (AS-719). Opus 4.7 and later reject the field, so it is left out for them. |
deepgram_model | nova-2 | The transcription model. |
Video pipeline
run_pipeline(job_id, db) runs five stages. The first two overlap: as soon as the audio track is out, transcription and frame extraction run at the same time. Progress percentages are shown so support can map a stuck job to a stage.
Stage 1: Extract
FFmpeg probes the video for duration, fps, and resolution and extracts the audio track to a file (seconds of work). The duration is stored on the job and later drives billing. Then frame extraction starts, and transcription (Stage 2) starts at the same moment on the audio file: the two have nothing to do with each other, so run_pipeline runs them together (_run_together, an asyncio.gather that cancels the survivor if either fails). Each keeps its own stage_timer; JobObservability clocks stages by name so the overlapping timers do not disturb each other.
Frame extraction (video_processor.extract_frames) writes one frame per second (frame_extraction_fps) and does three things that the old "every frame, full size, q=2" call did not:
- Shrinks inside ffmpeg. A
scalefilter caps the long edge atframe_long_edge_px(1600, never upscaling) and the JPEGs are written atframe_jpeg_quality(4). A 10-minute 4K video used to write about 1 GB of frames; now it is roughly 50 MB. Nothing downstream wants more: the vision model sees a 768 px long edge, item crops are 800x600, context images 1280x720, and the browser shows raw frames (GET /api/jobs/{id}/frames, the Guided Review source strip and the item modal) at under 940 px. The item crop is the tightest consumer, and 1600 keeps a half-frame crop at its full 800 px. - Decodes only keyframes when it can.
probe_keyframe_ratereads the packet flags (no decoding, under 0.1 s) and, if the video has at leastkeyframe_decode_min_ratioxframe_extraction_fpskeyframes per second, ffmpeg runs with-skip_frame nokey, so a 30 fps phone video costs one decode per second instead of thirty. Measured on a real iPhone 14 walkthrough (152 s, 1080p, 157 keyframes): 8.6 s down to 2.0 s, 151 frames instead of 152, and the frame for second N is the keyframe nearest to N (within half a second). Screen recordings and edited exports with long GOPs fail the check and get a full decode, so they never see the same picture repeated. - Heartbeats. Every
frame_extraction_heartbeat_secondswhile ffmpeg runs, the pipeline logs "Still extracting frames: N of ~M" and commitsprogress_pctandcurrent_stageto the job row. Before this, ffmpeg was silent for the whole extraction, so a long 4K video on a slow server could go quiet pastSTUCK_JOB_MINUTESand be failed (and refunded) by the watchdog while it was still working. The log line covers the in-process watchdog, the row commit covers the worker-based (Stage 2 topology) watchdog, and the user sees the bar move.
Frames left in the job's frames/ folder by an earlier attempt (a requeued job) are deleted before extraction starts, so the frame list and the heartbeat count only ever describe the current run.
Stage 2: Transcribe
The audio is sent to Deepgram Nova-2, which returns word-level timestamps. Nova-2 is used instead of Whisper because it is far more accurate on brand and product names, which is exactly what the analysis depends on. This runs while the frames are still being written (see Stage 1). The job's status stays extracting while frames are in progress; if the frames finish before the transcript is back, the status moves to transcribing ("Extracted N frames, waiting for the transcript") so the UI's step list stays honest. Once both are done the job sits at 40%.
The transcript is stored on the job (full text plus per-segment data with word timings). The pipeline also extracts room cues from the narration ("now we're in the master bedroom"), though in the current design each video upload is already scoped to one known room.
The running code transcribes with Deepgram Nova-2 (audio_processor.transcribe_audio), and the DEEPGRAM_API_KEY is what matters in practice. OpenAI and Whisper are not used anywhere in the pipeline; the legacy OPENAI_API_KEY and whisper_model config have been removed.
Stage 3: Chunk
The video timeline is divided into fixed-length chunks of chunk_duration_seconds (30s) by video_processor.compute_chunk_ranges (a pure, unit-tested function). Each extracted frame is assigned to the chunk whose time range contains it. Smaller chunks mean more, smaller AI calls that overlap and finish sooner.
Stage 4: Analyze (the parallel, expensive stage)
This is where items are actually found. All chunks are launched at once, bounded by a semaphore of size openrouter_concurrency (6), and results are collected as each finishes so progress streams live (55% to 85%).
For each chunk, the pipeline:
- Selects frames by density (
select_frames_by_density): rather than sending every frame, it picks the best frames up toframes_per_minute/max_frames_per_call, choosing the sharpest frame in each time bucket. - Slices the transcript to just this chunk's time range, so the model sees the narration that belongs to these frames.
- Calls Claude vision through OpenRouter (
claude_analyzer.analyze_room_segment) with the active system prompt for video.
After all chunks return, cross-chunk dedup merges near-boundary duplicates: if the same item name appears in two adjacent chunks within video_dedup_window_seconds (3s) of the boundary, it is treated as one physical object (a couch straddling the 30s line), not two. The merge is conservative and never collapses items that are far apart in time.
Stage 5: Synthesize
Items are sorted by video time so numbering matches recording order. For each item the pipeline crops the best frame to the item's bounding box (a close-up) plus a wider context image (on the thread pool, see the vision call), writes them to the job's photos/ directory, and saves an InventoryItem with its ItemPhoto rows. Item numbers continue from the claim's current maximum (see numbering). Items with confidence below 0.6 are flagged.
Finally, auto-approval runs if enabled, billing is finalized (video minutes times the per-minute credit rate, recorded for the usage meter), and the job is marked complete. When prepaid billing is enabled (ENABLE_BILLING), the upload did not start the pipeline at all: the job parked in awaiting_confirmation until the user confirmed the quoted cost, and the credits were spent up front at that confirmation (the spend is the gate). So nothing more is charged here; if the job instead fails, that up-front charge is refunded in the error path. All no-ops otherwise. See billing.
Photo pipeline
Photos work differently from video, and the difference is deliberate. Photos are distinct images, not near-identical frames, and the old approach of letting the AI guess which photos showed the same item produced noisy, duplicate-heavy output. The live pipeline removes that guessing.
How it flows (run_grouped_photo_pipeline):
- Upload (
POST /api/upload-photos, thenPOST /api/upload-photos/{id}/appendfor the rest): every photo is normalized to JPEG and saved; oneStagedPhotorow per file, each starting in its own group. The job sits inawaiting_grouping. AI does not start automatically. The frontend splits a large selection into size-bounded batches to stay under the Cloudflare Tunnel's per-request body cap, but all batches land on one job (see batched photo upload). - Group: the user merges photos that show the same item (capped at
MAX_PHOTOS_PER_ITEM, 5). The grouping is persisted via the staged-photo regroup endpoint, which enforces the same cap. A PDF import arrives with its groups already made, and they obey the cap too (see PDF import). The frontend has its own copy of the number inGroupPhotosModal.tsxso the user gets an inline message rather than a 400; the backend is the real guard. - Process (
POST /api/jobs/{id}/process): the user starts analysis. Soft-deleted photos are removed from disk first. - Analyze: one AI call per group runs in parallel,
openrouter_photo_concurrency(12) at a time, under the process-wideopenrouter_global_concurrency(24) cap. A multi-photo group uses the "merged" prompt and yields exactly one item; a single-photo group uses the "unmerged" prompt and may yield several items. - Synthesize: items are saved with every source photo attached as
ItemPhotorows. No cross-batch "is this the same item?" guessing is needed.
Described items (AS-705): the user can name an item during grouping, and a named item skips the AI entirely. The analyze step keys off manual_name: if a group has one, the pipeline builds the item straight from the photos, name, and quantity (confidence=1.0, the whole photo as its image) with no AI call, so it costs no credit. This covers both a whole photo / merged group the user named in the Items panel and a named drawn segment. The user sets the name and quantity via the staged-photos/details endpoint.
Manual segmentation (AS-705): during grouping the user can also draw boxes around items inside one photo (the segment endpoint crops each box into a StagedPhoto with is_segment=True, replacing the original). A segment is always a known item: the pipeline never asks the AI to re-split it, and the user's quantity overrides whatever count the model returns. A named segment skips the AI (as above); a blank-named segment is identified by the AI via the "merged" prompt like any known item, with the user's quantity forced onto the result.
Billing for photos is one request per group that hits the AI, not per photo. Described (named) items make no request and are free. See billing.
How much the AI may write back, and what happens when it runs out of room
A photo request earns exactly one credit however many items come back. That makes the size of the answer a money question, and it used to be answered with a constant: analyze_items_in_one_photo capped output at 1,500 tokens (about 10 items) and analyze_items_in_photo_group at 8,192 (about 54). Both were wrong in opposite directions. A photo showing more than ten items returned ten and the rest were silently lost: the parser salvages the complete items before the cutoff and logs a warning, so nothing crashes and nobody is told. Meanwhile a dense twenty-photo group could write 8,192 tokens and cost more than the credit it earned.
Neither number can be right for both, because the input is not the same size. At $0.10 a credit, the input is about 9% of the credit for one photo and about 65% for a twenty-photo group. So the cap is no longer a constant. Every request works out its own:
input_cost = (system prompt + user text + photos and their labels) x $5/M
output_budget = (credit_rate x target_margin - input_cost) / $25/M
clamped to [photo_budget_min_output_tokens, photo_budget_max_output_tokens]. It is priced at photo_budget_credit_rate_usd, the lowest per-credit rate any customer is on, so the answer is safe for every account. The photo tokens are measured from each image after the resize to 1.15 MP (_build_openai_content returns the count; _vision_size picks the size), not assumed at 1,530 a photo, because a panoramic or square photo is not a 4:3 one.
The result: one photo gets room for about 19 items, a three-photo group about 14, and the biggest group allowed (5 photos) about 10. A group that needs more room gets another pass (and another credit) rather than one pass that loses money.
Why photos are sent at 1.15 megapixels, and why that set the cap at 5. Photos used to be shrunk to a 768px long edge (~590 tokens). A phone photo kept about 4% of its pixels, and in a crowded room (a garage, a shelf) each item ended up 20-50px wide: the model named obvious items wrong. Photos now keep up to 1.15 MP (~1,530 tokens), Anthropic's classic image ceiling, which is about 2.6x the input cost per photo: a single photo's input went from 9% of a credit to about 16%, and it is where the extra detail matters most. At that size a group of 10 would spend about 85% of its credit on input alone, so the cap dropped from 10 to 5 (about 46% on input) in the same change. A 5-photo group is almost always one item shot from several angles, which needs the detail least. Video frames were not raised (video_frame_long_edge), since a chunk sends up to 60 of them.
The 80% target holds everywhere, with no exception. It did not always: the output floor (photo_budget_min_output_tokens, 800) used to take over on the largest group allowed, which was 20 photos, and that pass spent about 85% of its credit instead of 80%. At 1.15 MP the floor only binds above about 7 photos, so capping a group at 5 keeps the exception gone. Past about 11 photos the input alone costs more than the whole credit: a single request holding 12 full-size photos costs about $0.10 against the $0.10 it earns. That is why MAX_PHOTOS_PER_ITEM is 5 and why a PDF label or a homeowner group holding more than 5 photos is split rather than sent as one request. test_no_single_pass_can_cost_more_than_the_credit_it_earns sends full 12 MP phone photos and asserts the 80% bound unconditionally; it will fail if any reachable group size ever lands on the floor again.
| Where the answer stopped | What happens |
|---|---|
finish_reason="length" | The answer ran out of room. Every complete item is kept, and the group is recorded as holding more. |
Anything else, but the answer used at least photo_budget_truncation_ratio of its budget | Treated as cut short anyway. A finish_reason we cannot trust would put us straight back to losing items in silence, which is the bug this exists to fix. |
| Anything else | Finished on its own. Nothing remains, however many items came back. |
The job never stops to ask. A 135-photo import runs unattended in the background, and eight photos hitting the cap would ask eight times. So the groups that were cut short are written to ProcessingJob.photo_truncated_groups (the group index, how many passes it has had, and the item names read out of it so far), the job finishes normally, and the offer to read the rest is made once, afterwards, on the finished upload. photos_with_more_items on the job schemas is the count.
The extra pass is a continuation, the same shape as the audio chunk continuation in _call_word_chunk: the same photos are sent again along with the names already found, asking only for what is not on that list. Those names are part of the input, so they are paid for out of the same budget. photo_max_passes (3) is the ceiling: a group still cut short after three passes is too dense to read, and the user is told to draw boxes with Segment instead.
The user asks for it, and pays for it first. The finished upload's row in the room's Uploads table says how many photos had more items than one pass could read, and offers a Read the rest button. That button quotes the exact price (GET /api/jobs/{id}/read-the-rest/quote), shows the same credit confirmation any other spend does, and only then starts the pass (POST /api/jobs/{id}/read-the-rest, which spends up front and refunds on failure). run_read_the_rest_pipeline re-reads only the recorded groups, saves the new items after the existing ones (its sort_start continues past the highest sort_order already on the job, so a later pass can never overwrite an earlier pass's crops), and adds its calls to photo_request_count. Billing is unchanged in kind: one credit per AI call, so three passes over a group cost three credits.
The photo prompts' OUTPUT FORMAT blocks no longer ask for room or audio_evidence. Neither was ever read off a photo-sourced item (the pipeline sets room from the room name on every path, and audio_evidence is always "" for a photo), and every field the model writes costs output budget. At about 159 output tokens an item a full pass on a single photo left 4% margin at $0.10 a credit; at about 130 after the trim it leaves 20%. The design does not work without it.
PDF import
A PDF with photos in it is dropped in the same upload box as photos and joins the same grouping pile. Everything after extraction is the photo pipeline above, unchanged. The reference case is a vendor contents list (Encircle, Xactimate and the like), but nothing in the reader requires that: it looks for embedded photos and the text printed near them, and falls back to photographing pages when it finds neither.
The design point: the AI is asked once per vendor FORMAT, never per file and never per page. These exports lay their pages out identically on every page of every export from that tool, so the model reads a small structural sample once and writes an extraction recipe; the server executes that recipe over every page in pure Python.
Where it lives. services/pdf_extractor.py is the pure reader (page-range detection, fingerprint, recipe schema, executors, page rendering) and depends only on PyMuPDF. services/pdf_import.py orchestrates: find the inventory pages, resolve the recipe, extract, fall back, stage. services/pdf_claim_import.py is the claim-level path that splits the result across rooms. models/pdf_template.py stores the learned recipes. claude_analyzer.build_pdf_recipe and claude_analyzer.locate_photos_in_page are the only two AI entry points.
Two layout families, chosen by the recipe's strategy:
| Strategy | Shape | Example |
|---|---|---|
card_grid | One photo per item, the item's label drawn ON the photo or immediately under it. Labels and photos come in matching numbers. | Encircle, Xactimate |
header_runs | A heading line owns the RUN of photos below it, until the next heading. No photo carries a label of its own, and one item routinely holds a dozen photos. | Servpro "Content Items List With Images" |
header_runs exists because card_grid cannot read that shape at all: it looks for a label drawn on each photo, finds none, and an unlabelled photo becomes its own item. A real 15-item Servpro page came back as 129 items. The reader walks each page as one stream of events in reading order (room heading, item heading, photo), and the open item carries across pages, because a run of photos routinely spills over the break. A photo printed before any heading is page furniture and is dropped. When a header_runs recipe finds nothing, pdf_import retries with card_grid before falling back to page images.
Room headings (Recipe.room_pattern, group 1 = the name) are read by both strategies and ride along on every extracted photo as ExtractedPhoto.room, then onto StagedPhoto.source_room. An item keeps the room it was OPENED under, so the tail of the last item in the Garage is not filed under the Kitchen because a room heading happened to print above it.
Front and back matter is found and skipped, so nobody has to delete a cover sheet before uploading. find_inventory_range(doc) returns the first and last page holding inventory, and every later step (the fingerprint, the AI sample, the rendered page, the extraction) works inside it. The signal is the item photo itself: a vendor renders every item photo at one pixel size, thousands of times, while a cover carries a logo at some other size and a terms page carries nothing. So it finds the size that dominates the document and takes the first and last page carrying one. Only the ENDS are trimmed; whatever sits between those pages is kept exactly as it is. When the evidence is thin (a scan with no embedded photos, no dominant size) it returns the whole document, because trimming on a guess would silently drop items.
This matters more than it looks: the AI sample comes off the front of the file, so a cover page as page 1 means the reading rules get written from a page with no items on it.
The recipe is data, never code. Recipe.from_dict keeps only known fields, coerces types, and discards any regex that will not compile. Every value has a safe default, so malformed model output degrades the extraction rather than raising.
Reading the layout (Recipe.strategy == "card_grid", validated against a 22-page Encircle export: 135 photos, all 135 carrying a label, a description and a date):
- A photo is an embedded image at least
min_image_pxon both edges and not repeated on most pages (that would be a logo). - Its label is the text block drawn over the image (or just under it), matched by
label_pattern. - Its caption is found by walking outward from the photo: the nearest text block on the
description_positionside must be withindescription_max_gap_ptto count as a caption at all, and each block beyond it is kept only while it is withindescription_line_gap_ptof the one already kept. Wrapped lines of one paragraph sit ~2pt apart and all survive; a section header 18pt above the caption does not. This is the safety net that keeps a page title out of an item name even whenignore_patternsmisses it. - A caption with no photo below it on its own page belongs to the first photo on the next page: the card grid flows across page breaks. Three of the 135 items in the reference export depend on this.
- A leading
15 xon a caption becomesquantity=15, and the printed date is normalized to the sortableYYYY-MM-DDTHH:MM:SSthe grouping step sorts on.
Grouping comes for free. Photos printed under one label are one item, so they are staged sharing a group_index and the grouping screen opens with them already merged. The group key is (room, label), not the label alone: labels restart per room in some formats, and two PDFs dropped in one batch can both number their items from 1.
A label printing more than MAX_PHOTOS_PER_ITEM photos is split into several groups (pdf_import._assign_groups). One group is one AI request earning one credit, and the photos are input that credit has to pay for, so an unsplit group of 31 full-size photos would cost about $0.25 against $0.10. Nothing stops a vendor printing 60 photos under one line number, so the cap cannot live only on the endpoints a person drives. A homeowner grouping photos through their inventory link has no cap on their screen either, so insured_link._materialise_room_photos splits an oversized group the same way at submit.
The blocks are balanced, not filled to the cap: 31 photos with a cap of 5 becomes 5, 5, 5, 4, 4, 4, 4, never six fives and a 1. Both are seven blocks and seven credits, but a one-photo block is not a block at all: it leaves its siblings, goes down the single-photo path, and reads as a mistake on the grouping screen. _balanced_blocks(n, cap) takes the fewest blocks that all fit (ceil(n / cap)) and spreads the photos over them. Photos stay in document order both within a block and across the blocks, and every block keeps the same source_item_key, source_description and source_room, which is what lets the grouping screen show Item 12, part 2 of 7 on each card and tell the user in plain English that one document line became seven groups.
The scanned fallback (strategy == "page_image") runs when there is no text layer or the recipe finds no photos. Each page is rendered and locate_photos_in_page marks where each photograph sits, and the server crops them into separate cards. Its three outcomes are distinct on purpose: boxes found means crop; an empty list means the model looked and found no photographs, so the page is dropped (a terms page must not become an item); None means the call could not be made, so the page is kept whole rather than lost.
Fingerprinting. fingerprint() hashes the lines that repeat on most inventory pages (digits masked, e.g. powered by encircle, page # of #) plus the page size and producer. It returns that hash and how many repeating lines backed it; below MIN_SIGNATURE_LINES the recipe is always re-derived rather than guessed at. A document of fewer than two pages always reports zero: with one page there is no repetition, so every line looks "repeating", including the item text, and caching on that would key the recipe to one document rather than to a vendor's layout.
A near-match on the furniture lines catches the vendors who print the insured's name. The hash is all or nothing, and Servpro repeats Hatfield Tim & Rhonda on all 506 pages of an export, so its "layout" hash is really a hash of the CLAIM and every family would pay for the AI to re-derive a format the app already knows. So the lines the hash was built from are stored alongside the recipe in pdf_templates.signature_json, and on a hash miss pdf_import._match_template compares them. Deliberately conservative, because applying the wrong vendor's recipe reads a document into nonsense: the page size must match exactly, the signatures must share at least MIN_SHARED_SIGNATURE_LINES (3) lines, and their overlap coefficient (shared / smaller set, pdf_extractor.signature_overlap) must reach MIN_SIGNATURE_OVERLAP (0.6). Two Servpro exports for different families share three of five lines, which is 0.6; two different vendors share none.
Learned recipes are global, not team-scoped. They hold layout rules only (coordinates, regexes), never claim content, so the first upload of a new format teaches the whole app and every later import of it is instant and free.
The note printed beside a photo is a HINT for the AI, never the item's name. Each staged photo keeps the document's own wording in source_description. That wording is usually poor ("Charred barber products", or a 60-word run-on list), so it is not used as the item name and there is no user-facing choice about it. Instead:
- It is appended to the user prompt for that item's AI call as
SOURCE_NOTE_BLOCK(claude_analyzer._with_source_note), truncated at 1,200 characters so it cannot crowd out the photo. - The instruction that explains how to use it lives in the user message, not the system prompt, on purpose: a team can replace the photo system prompts entirely (see automation and prompts), and the rule has to stay understood when they do. Keeping the two together means the rule always travels with the data.
- The rule the model is given is explicit about the order of operations, which is the whole point: identify the items from the photo first, exactly as if the note were not there, and only then read the note to confirm what was found. Told the other way around a model will work down the note hunting for matches and report items that are not in the picture, which is precisely what these run-on descriptions invite.
- The rest of the rule: a note may cover one item or several; where it names a detail you can see but could not have named alone (a brand, a model number, a material), use it to sharpen both the item and its description; where it mentions something not visible, leave it out and never invent an item to match the note; where the note and the photo disagree, the photo wins; never copy it in as the item name or quote it back as evidence.
- Photos merged as one item hand over all of their notes, joined and de-duplicated in order (
pipeline._source_note). - The note is also saved onto the finished item's
notesasFrom the source PDF: ..., so the source record survives into the inventory next to what the AI saw.
Because the note assists the AI rather than replacing it, every group is still billed exactly as a normal photo upload. Only a name the user types still skips the AI.
source_description is not only for PDFs. When the insured uploads a photo through their own inventory link, the note they type against that file ("the TV is behind the boxes on the left") is written to the same column at submit and travels the same path into the same prompt block. The ordering rule above is what makes that safe: a homeowner listing five things in one note cannot make the model report five items from a photo showing two.
When an import goes wrong, support gets the file. pdf_import.assess() judges each extraction and returns a reason string when it looks wrong, and both _import_pdfs_background and pdf_claim_import mail that reason plus the PDF itself to SUPPORT_EMAIL via services/support_email.py (Mailgun; a no-op until configured, like the Slack notifier). The user still keeps whatever was found. The checks are deliberately narrow, because each has to mean "we probably misread the layout" rather than "this document is unusual":
| Signal | Why it means we got it wrong |
|---|---|
| No photos found, or the read raised | Nothing to hand over. |
| Every page imported as a whole image while the document has a text layer | A real scan has no text; text plus no embedded photos means the reader could not find photos it should have. A genuine scan is not reported. |
| More than 80% of embedded photos came out with no label and no description | The photos were found but the text could not be matched to them, so the caption rules are wrong. |
A failed import almost always means a layout we have not seen; the fix is to teach the reader that format, and that needs the actual file, which is why it is attached.
Uploading, room-level. A PDF rides the first POST /api/upload-photos request alongside any loose photos; it is refused by the /append endpoint, because a PDF yields an unknown number of photos and cannot fit the client-assigned index ranges parallel appends depend on. The job is created at importing_pdf and the read runs on the worker (run_pdf_import; a background task without REDIS_URL), because a new format costs an AI call and every page has to be walked, which is far longer than a request may stay open. Its staged files use a pdf_#### prefix and take group indices past photo_count_hint, so a PDF's photos can never collide with a loose photo arriving in a later batch.
Uploading, claim-level. POST /api/claims/{claim_id}/import-contents-list exists because a ProcessingJob belongs to a ROOM, so the room-level box can only ever put a file into the room it was opened from. A contents list is not one room: the reference Servpro export is 1,593 items across 15. The endpoint parks the PDFs under <storage>/pdf_imports/<batch_id>/ and returns an ImportBatch immediately; services/pdf_claim_import.py then reads each file through the same pdf_import.extract_photos, splits the photos by ExtractedPhoto.room, finds or creates that room on the claim (matched case-insensitively against its ACTIVE rooms, so a second list adds to "Garage" rather than making another one), and stages each room's photos into that room's own awaiting_grouping photo job. A document that names no room lands in one room called Unsorted.
Progress lives on the same ImportBatch row the Contents Capture import uses, so the UI polls GET /api/integrations/imports/{batch_id} for both. run_pdf_claim_import never raises: every failure lands on the batch as a readable message, because an import that failed must never look like one still running.
Scale ceilings. MAX_PAGES (2,000), MAX_IMAGES (12,000) and MAX_IMAGE_BYTES (1 GiB) bound what one document may cost. They were 400 and 2,000 and would have thrown away two thirds of the 506-page, 5,712-photo reference export. The byte budget is the real backstop: page and image counts say nothing about how big each photo is, and the whole photo list is built in memory before anything is written to disk.
The photo flows use three different system prompts (photo_merged, photo_group_multi and photo_unmerged), each separately customizable by a team. See automation and prompts.
How many items a GROUP may yield is a per-upload choice. photo_merged says "ALL these photos show the SAME item, do NOT split", which is the promise the grouping screen makes when a PERSON merges four angles of a sofa. A vendor contents list does not group that way: it groups by its own line number, and one line is routinely a box of assorted products (Servpro's #3 is thirteen photos of a protein powder, a Swiffer, refills and a bottle of Clorox, all named in its printed note). Asked the merged question, that whole line came back as one item and three products were lost.
So ProcessingJob.groups_may_hold_several_items, set from a checkbox on the grouping screen and passed to POST /api/jobs/{id}/process, picks which question each multi-photo group is asked. It defaults ON in the UI when any staged photo carries a source_item_key (the photos were grouped by a document, not by a person) and OFF otherwise, so hand-merging is untouched. analyze_items_in_photo_group returns a LIST and every item carries the whole group as its photos; the note is a hint about what to look for and never a licence to report an item that is not in the picture.
Billing is unaffected: it counts AI requests, and a group is one request however many items come back. That is what makes a long contents list affordable, since the alternative is one request per item.
Audio pipeline
run_audio_pipeline handles voice memos and call recordings. No frames, no photos: just transcribe and read.
The analyzer records, for each item, the audio_start_seconds and audio_end_seconds of the part of the recording that mentions it, so Guided Review can play back just that slice. If the recording has no speech, the job completes cleanly with zero items and a friendly message rather than failing. Audio is billed per 20 second block, like video, at a lower rate.
The vision call
All analysis goes through claude_analyzer.py, which calls OpenRouter's OpenAI-compatible endpoint (a single routing layer with model fallback and rate-limit headers).
- Model:
openrouter_model, defaultanthropic/claude-opus-4-7. It replaced Opus 4.5 (same $5 / $25 per million tokens) for sharper image recognition; it also accepts images up to 2576px on the long edge, which is what allowed the photo size above to be raised. It counts text as up to ~35% more tokens than 4.5 for the same words, which the output budget absorbs (the prompt is a small share of a request). Nothinkingfield is sent, so the model does not reason before answering and output cost is as before. - Sampling: on models that accept one, every call pins
temperaturetoopenrouter_temperature(default0) so the same video or photo set yields the same items and names run-to-run (ticket AS-719). With no temperature a model defaults to ~1.0 and samples differently each run, which is what made repeated uploads produce different item counts and names. Temperature 0 makes output near-deterministic; it is not a bit-identical guarantee (provider batching, routing, and floating point still vary, and Anthropic models do not honor aseed). Onlytemperatureis sent, nevertop_palongside it, since Claude 4.x rejects both at once. Opus 4.7 and every later Anthropic model reject sampling parameters outright (a 400 on any request carryingtemperature), soclaude_analyzer.model_accepts_samplingleaves the field out for them and run-to-run variation is whatever the model gives. SettingOPENROUTER_MODELback to Opus 4.5 sends it again with no other change. - Concurrency: a shared, connection-pooled
httpxclient; the pipeline limits in-flight calls with a semaphore. - Image preparation runs off the event loop. Every image sent to the model is opened, turned upright (its EXIF orientation baked in), shrunk to a 768px long edge and re-saved as JPEG by
_prepare_image_for_model, which_build_openai_contentfans out across the thread pool (run_in_executor) and gathers. That work is 20 to 300 ms per image (Pillow'sdraftmode decodes a large JPEG at reduced scale, about half the cost of a full decode) and a video chunk carries thirty of them; done inline, as it once was, it froze every other request on the server for seconds at a time while a job started its chunks. The same rule applies to the frame scoring inselect_frames_by_density(OpenCV, run on the thread pool from the pipeline) and to the per-item crop and context images cut in the synthesize stage (process_item_photo, whose body_process_item_photo_syncruns on the thread pool). Nothing in the pipeline that decodes or resizes an image may run on the event loop; it is what keeps uploads and page loads for every other user responsive while a job runs. - Retries: up to 4 attempts per call (
max_retries, not overridden by any caller). On HTTP 429 it honorsRetry-Afterorx-ratelimit-reset; on 5xx, timeouts, and network errors it uses exponential backoff with jitter. - Unreadable replies: a
200whose body will not parse as JSON is retried on the same backoff as a 5xx (ticket AS-964). OpenRouter holds a non-streaming connection open with whitespace padding while the model works, and occasionally that padding arrives with no completion behind it, which reads as a normal200with a blank body. Before AS-964 the decode error escaped everyexceptbranch in the retry loop and failed the whole job on the first attempt, discarding the sibling chunks that had already succeeded. The failure message now quotes the first 400 characters of the body (or its length when it is whitespace only), so a recurrence is diagnosable from the log rather than from a character offset. - Inspection: every call's final outcome is recorded to the
api_request_logstable (request JSON with images redacted, full response, attempts, duration, tokens, real cost) and browsable on the admin Logs page's API Requests tab. - Structured output: every prompt ends with a locked
OUTPUT FORMAT:block defining the JSON contract the app parses. Custom prompt add-ons are spliced in before that block, so the contract is always preserved.
The built-in system prompts encode the domain rules: be thorough, name items like searchable retail product names, put brand in its own field, document damage, group identical items, include movable property plus cabinets and appliances, and exclude structural building parts entirely. The condition scale is fixed (New, Above Average, Average, Below Average), and confidence is scored by how clearly an item is seen and whether it was named in narration.
Age is only ever taken from something a person said. An estimated age is unreliable and drives item value, so the model may set age_years / age_months only when the age is stated aloud, and never from wear, styling, or technology generation. The video and audio prompts keep the age fields but gate them on a spoken statement (video additionally requires the stating words to appear in audio_evidence). The three photo prompts drop the fields entirely, since a photo set has no narration to quote. analyze_room_segment enforces this by mode: mode="photo" hardcodes both fields to 0 regardless of the response, so a team's custom prompt add-on or an off-format reply cannot seed a guessed age; mode="video" and _audio_item_from_json pass a stated age through. Anything unstated stays 0 for the reviewer to fill in.
Item numbering
Numbers are unique and sequential within a claim and continue across every job in that claim (numbering.py). AI items are numbered in video-time order. Manual inserts can append (max + 1) or insert between numbers, shifting later items up by one.
A failed AI call is not an empty result
claude_analyzer swallows its own exceptions and returns ([], {}), which looks identical to the model legitimately finding nothing. That is the difference between a job that should complete and one that should refund, so every pipeline now tells them apart by whether a usage dict came back: a real answer always reports token usage, a failed call reports none.
- Photo (
run_grouped_photo_pipeline) countsn_ai_ok. If any group was sent to the AI and none answered, it raises, which takes the existing failure path: the job is markederrorandrefund_for_jobreturns the up-front charge. It also metersphoto_request_countonn_ai_ok, not on the number of groups attempted, so a partly-failed run charges only for the calls that worked. - Video (
run_pipeline) does the same across its segments. - Audio (
run_audio_pipeline) makes one call, so no usage means it failed. - A job with nothing to send the AI (every item named by the user) is not an outage and still completes, billing nothing.
Found by driving the real app with an expired OpenRouter key: all 135 photo calls returned 401, and the job finished as Complete with 0 items while charging 135 credits.
Error handling
Every pipeline wraps its work in a try/except. On any failure the job is set to status=error with the exception type in current_stage and the message in error_message, and the failure is logged to the job's observability buffer. Partial work already committed (for example, items from chunks that succeeded) is kept. The Adjust Square submission path has its own retry and per-item failure handling, covered in its own page.
Next
- Automation and prompts: auto-approval rules and custom prompt add-ons.
- Observability: how stage timings, tokens, and job logs are captured.
- Billing: how each pipeline's output is priced.