RZMRN
← cd ..SYSTEMS

The AI studio I built — and the agents that run it

The in-house AI studio of a creator-marketing agency, now serving external clients. I designed and built its generation platform, run its GPU fleet and agent swarms, and lead the team and the numbers around them.

The AI studio I built — and the agents that run it
ComfyUIOpen-weights image & video modelsLoRA / DoRAPythonFastAPISQLiteNext.jsDockerRunPod GPU fleetClaude CodeCodex

≈6×

Lower compute cost per video

44% → 81%

Output run by agents, in one month

89% → 0%

Frames with the wrong face

4,000+

Images & videos in four months

From the studio’s own reports and logs, anonymized.

The Challenge

The agency wanted AI photo and video at production volume: consistent characters, delivered every day, at a cost that makes sense. There was no platform, no pipeline and no team. I was the only engineer — which in 2026 means me, a spec, and a lot of AI coding agents.

Hosted video APIs priced every clip the same, quality depended on luck, and nobody could say what a delivered video actually cost.

The Approach

A brief goes in; approved, delivered clips come out. In between: an LLM that writes prompts, a durable job queue, a GPU cluster running several ComfyUI engines with open-weights models and LoRA/DoRA character fine-tunes, automated checks — and one human verdict.

Most of production is run by agent swarms: from pulling client references to fleet rendering, packaging and delivery. By stage count, about three-quarters are agent-only. People keep three gates: the budget, identity sign-off and the final verdict. Autonomy means continuing authorized work — not changing my decisions.

A new external client ordered 100 clips. 19.5 hours after kickoff they had 121 approved videos — the render queue ran unattended overnight. Agents also shut down their own GPUs when the queue drains: idle burn is a bug, not a phase.

stages 1–4 · plan

  1. 1BriefRefused before any GPU is spent if the models can’t do it
  2. 2Prompt prepAn LLM fills in values; Python owns the structure
  3. 3Review gateEvery plan editable before dispatch
  4. 4Job queueLeases, heartbeats, retries, fair share

The studio pipeline in eight stages: brief, prompt prep, review gate, job queue, GPU cluster, automated QC, human verdict, delivery. Failed QC retries automatically; a retake goes back to prompt prep.

stages 5–8 · render, check, deliver

  1. 5GPU clusterSeveral ComfyUI engines; idle pods get reaped
  2. 6Automated QCFace-embedding drift, motion metrics
  3. 7Human verdictGOOD or RETAKE, from full-resolution frames
  4. 8DeliveryPackaged with checksum ledgers

QC fail → automatic retry · RETAKE → back to prompt prep

The Result

Four months after starting from zero, the studio had delivered 4,000+ images and videos, went to market in 20 days, and now sells to external clients; AI content from the studio became the agency’s main traffic source.

Beyond the platform, I hired and managed 2–4 operators (hourly output per operator ×3.6 in one month), built pricing and packaging, and reported cost, quality and forecasts to the owners every month — June and July closed 11% and 8% under forecast.

Incident reports

Four problems, what I did about them, and what changed. One number each.

Identity bleed

89% → 0%

Problem

Two agent-run batches came back 16 out of 16 defective: the character’s face slowly turned into the person in the motion reference. The easy verdict was “the model can’t do this.”

What I did

Refused to blame the model without evidence. Built a per-frame face-embedding metric and ran single-factor A/B tests on real material. Found five causes — all on our side — and shipped validator warnings, reference trimming and a new prompting rule for staff and agents.

Result

Frames with the wrong face: 89.3% → 0.0%.

One pod, one clip, zero idle

230 clips overnight

Problem

Batching looked efficient — until 7 of 13 four-clip jobs died of container memory exhaustion and took their in-flight clips with them. Pods sat idle between hand-offs.

What I did

Replaced batches with single-item jobs and a dispatcher with one ledger, so double assignment is impossible by construction. Failed items migrate to another pod; drained pods are killed automatically.

Result

230 clips rendered overnight (5 h 47 min) on up to 12 RTX 5090s; 213 delivered at ~77% first-pass acceptance (83% after retakes).

Rent, buy or API?

≈6× cheaper

Problem

Hosted video APIs priced every clip the same, so every new order made the bill grow in a straight line.

What I did

Priced a dedicated server against rented GPUs (renting won), then moved video to self-hosted open-weights models on our own fleet — and kept the honest comparison on record, including the one lane where the hosted API stayed cheaper.

Result

≈6× lower compute cost per mass-format video; API video spend −89% within a month.

The deploy gate that failed open

1,145-test CI gate

Problem

Deploys were rsync and a manual restart, and a deploy once killed a live render at 20%. The first “never deploy over a render” check failed open: a quoting error plus “|| true” read “no jobs running” while two renders were running.

What I did

A private repo, a hard CI gate, one-command deploys with a journal, a pre-deploy database snapshot, health checks and rollback — and both gates rewritten as tested code.

Result

Every release now passes a CI gate of 1,145 automated tests and refuses to deploy during a render — verified in production. A safety check you haven’t tested is a decoration.

$ Did I type it all myself? No. Most of the code was co-written with AI coding agents. I own what agents can’t: the architecture, the specs, the invariants, the reviews, the acceptance — and the 3 a.m. production calls.

Role: AI & Automation Lead · platform architecture, agent systems, GPU fleet, team and unit economics