By August 2026 the PawScapes pipeline had mirrored a photo for most venues into R2, at keys like venues/{id}.*. The site served them fine. The admin couldn't see them: they were files in a bucket, not media documents, so they had no alt text, no focal point, no search, and no record of where they came from.

Uploading them again through the admin would have duplicated every file and handed their lifecycle to the CMS, while the pipeline's mirror and re-encode scripts still expected to own venues/{id}.*. highseam 1.1.0's import endpoint was designed against exactly this corpus.

Identity is the venue, not the URL

The obvious key for "have I imported this photo already?" is its source URL. It's also wrong here. The originals came from signed Google URLs that expire and get re-issued differently for the same photo, so the same image would look new on every run. The import keys each photo on something stable instead:

pawscapes-cms/scripts/import-media-provenance.mjs
* Idempotent by server design: re-running returns `unchanged`, never dupes.
* Identity is source.externalKey = "venue:{id}" — originUrl is provenance
* metadata only (gps-cs-s URLs rot and re-mint differently; see
* docs/AI/research/provenance-handoff-2026-08-05.md in the site repo).

The origin URL is still stored, as provenance, alongside when it was mirrored and the first venue it was used on. It just isn't identity.

Adopt in place

Each batch of 50 goes to POST /api/media/import and describes files that already exist in the bucket. highseam creates the documents and marks the files importer-owned. No bytes move, and the pipeline keeps managing the objects.

The run on 6 August went like a long job usually goes: a 50-item trial (which also proved a rename keeps the same identity), then two interruptions when the laptop lid closed and one transient timeout. Each was fixed by running the script again, because a re-run returns unchanged for anything already done. A patch made resumes skip the adopted set entirely. The final run created 2,102 documents with 0 failures, and the total reached 7,954 of 7,954, checked against D1.

Renames, a week later

In mid-August a cleanup renamed and re-slugged a batch of venues. Media display names are derived from slugs ({slug}-hero), so they were now stale. The fix was the same script with --full:

pawscapes/docs/AI/log/DONE.md (2026-08-14)
0 created / 55 updated / 7,899 unchanged / 0 deduped / 0 failed

The 55 are exactly the renamed venues that had photos. Everything else was recognised and left alone.

Dedupe came later

highseam 1.6.0 added content-hash dedupe: identical bytes uploaded twice return the existing document. For imports it only applies when the external key finds nothing, and it never overwrites provenance. On the re-run above, every photo matched on its external key first, so dedupe never came into play (0 deduped).

The media library is now a searchable 7,954-image collection, and searching by a venue's slug finds its photo.