The AI studio I built — and the agents that run it
The in-house AI studio of a creator-marketing agency, now serving external clients. I designed and built its generation platform, run its GPU fleet and agent swarms, and lead the team and the numbers around them.
≈6×
Lower compute cost per video44% → 81%
Output run by agents, in one month89% → 0%
Frames with the wrong face4,000+
Images & videos in four monthsFrom the studio’s own reports and logs, anonymized.
The Challenge
The agency wanted AI photo and video at production volume: consistent characters, delivered every day, at a cost that makes sense. There was no platform, no pipeline and no team. I was the only engineer — which in 2026 means me, a spec, and a lot of AI coding agents.
Hosted video APIs priced every clip the same, quality depended on luck, and nobody could say what a delivered video actually cost.
The Approach
A brief goes in; approved, delivered clips come out. In between: an LLM that writes prompts, a durable job queue, a GPU cluster running several ComfyUI engines with open-weights models and LoRA/DoRA character fine-tunes, automated checks — and one human verdict.
Most of production is run by agent swarms: from pulling client references to fleet rendering, packaging and delivery. By stage count, about three-quarters are agent-only. People keep three gates: the budget, identity sign-off and the final verdict. Autonomy means continuing authorized work — not changing my decisions.
A new external client ordered 100 clips. 19.5 hours after kickoff they had 121 approved videos — the render queue ran unattended overnight. Agents also shut down their own GPUs when the queue drains: idle burn is a bug, not a phase.
stages 1–4 · plan
- 1BriefRefused before any GPU is spent if the models can’t do it
- 2Prompt prepAn LLM fills in values; Python owns the structure
- 3Review gateEvery plan editable before dispatch
- 4Job queueLeases, heartbeats, retries, fair share
The studio pipeline in eight stages: brief, prompt prep, review gate, job queue, GPU cluster, automated QC, human verdict, delivery. Failed QC retries automatically; a retake goes back to prompt prep.
stages 5–8 · render, check, deliver
- 5GPU clusterSeveral ComfyUI engines; idle pods get reaped
- 6Automated QCFace-embedding drift, motion metrics
- 7Human verdictGOOD or RETAKE, from full-resolution frames
- 8DeliveryPackaged with checksum ledgers
QC fail → automatic retry · RETAKE → back to prompt prep
The Result
Four months after starting from zero, the studio had delivered 4,000+ images and videos, went to market in 20 days, and now sells to external clients; AI content from the studio became the agency’s main traffic source.
Beyond the platform, I hired and managed 2–4 operators (hourly output per operator ×3.6 in one month), built pricing and packaging, and reported cost, quality and forecasts to the owners every month — June and July closed 11% and 8% under forecast.
Incident reports
Four problems, what I did about them, and what changed. One number each.
Identity bleed
89% → 0%
Problem
Two agent-run batches came back 16 out of 16 defective: the character’s face slowly turned into the person in the motion reference. The easy verdict was “the model can’t do this.”
What I did
Refused to blame the model without evidence. Built a per-frame face-embedding metric and ran single-factor A/B tests on real material. Found five causes — all on our side — and shipped validator warnings, reference trimming and a new prompting rule for staff and agents.
Result
Frames with the wrong face: 89.3% → 0.0%.
One pod, one clip, zero idle
230 clips overnight
Problem
Batching looked efficient — until 7 of 13 four-clip jobs died of container memory exhaustion and took their in-flight clips with them. Pods sat idle between hand-offs.
What I did
Replaced batches with single-item jobs and a dispatcher with one ledger, so double assignment is impossible by construction. Failed items migrate to another pod; drained pods are killed automatically.
Result
230 clips rendered overnight (5 h 47 min) on up to 12 RTX 5090s; 213 delivered at ~77% first-pass acceptance (83% after retakes).
Rent, buy or API?
≈6× cheaper
Problem
Hosted video APIs priced every clip the same, so every new order made the bill grow in a straight line.
What I did
Priced a dedicated server against rented GPUs (renting won), then moved video to self-hosted open-weights models on our own fleet — and kept the honest comparison on record, including the one lane where the hosted API stayed cheaper.
Result
≈6× lower compute cost per mass-format video; API video spend −89% within a month.
The deploy gate that failed open
1,145-test CI gate
Problem
Deploys were rsync and a manual restart, and a deploy once killed a live render at 20%. The first “never deploy over a render” check failed open: a quoting error plus “|| true” read “no jobs running” while two renders were running.
What I did
A private repo, a hard CI gate, one-command deploys with a journal, a pre-deploy database snapshot, health checks and rollback — and both gates rewritten as tested code.
Result
Every release now passes a CI gate of 1,145 automated tests and refuses to deploy during a render — verified in production. A safety check you haven’t tested is a decoration.
$ Did I type it all myself? No. Most of the code was co-written with AI coding agents. I own what agents can’t: the architecture, the specs, the invariants, the reviews, the acceptance — and the 3 a.m. production calls.
