Skip to content

This page is fully readable without JavaScript. Every figure is printed in the HTML. Motion, the evidence receipts, the project map and the command palette are progressive enhancements only.

Robert McCullock · AI Systems Architect · Boston

I build autonomous AI systems.Then I audit my own claims about them.

Every number here is graded — click one for its receipt

I write the constraints and the acceptance criteria. AI agents write, test and deploy the code. 772,119 authored lines in eighteen months — measured with declared exclusions, not estimated. Where a July 2026 audit proved a published figure wrong, this page says so and shows what it used to say.

VOLTHAVEN 64 — N64-style, unlicensed, not affiliated with Nintendo. Click to run it here.
Quick answer

Robert McCullock is a Boston-based AI systems architect and full-stack engineer, and the founder of Design Delight Studio. Working solo, he writes specifications and acceptance criteria and directs AI agents — Claude, Antigravity, Gemini and local Ollama models — to build and ship production software. The work spans six production AI systems led by Sovereign Orchestrator Pro, two commercial desktop apps, three browser games, a free 86-class coding academy and a Shopify storefront, totalling 772,119 authored lines of code measured on 27 August 2026 with stated exclusions. Every figure on this page carries an evidence grade: MEASURED, CALCULATED or ASSERTED.

Contact: Robert@ddsboston.com · (617) 334-5912 · Boston, MA · LinkedIn · ddsboston.com

How this was built

Every visible number, its receipt panel, the audit table and the Dataset JSON-LD all render from one Liquid array. They cannot drift apart because there is only one copy of the data.

{%- assign rows = ledger | split: '~~~' -%}
{%- for row in rows -%}
  {%- assign f = row | strip | split: '|' -%}
  {%- if f[0] == id -%} ... render f[1] value, f[3] grade ... {%- endif -%}
{%- endfor -%}

Evidence you can operate

Every number above asks you to trust me. This one runs.

Six thousand instanced glyphs in 3D, in the browser you are reading this in. Every glyph is a real character read out of this page’s own source at runtime — the object is built from the code that draws it. Drag to orbit, scrub to pull it apart, click a component to read what it stands for. Press X-ray and it prints its own tier, frame cost, draw calls and pixel ratio while it runs; those are read off the renderer, not typed in by me. Press Force fallback and watch it drop a tier on purpose.

Compiling the cap…
assembled
1.0×
00

The assembly

Four components. Drag to orbit, scrub to take it apart, and pick a component to read what it stands for. Every glyph you see is a character of this section's own source code, read out of the page at runtime.

Built with three.js. No build step, no tracking, no signup. Free, no gate.

Built for the Academy and documented end to end in the class on the five instruments that proved it wrong before it shipped — including the ones that found a bug in my own measuring tools.

The method

Constraints, not keystrokes

I run a pipeline where I write the specification and the acceptance criteria, and agents do the typing. The part that matters is not that agents write code — it is that nothing reaches production without an audit stage that can fail the build, and that stage never accepts an agent’s word for anything.

01

Architect

Intent, constraints, acceptance criteria. The only human authorship in the pipeline.

02

Author

An agent produces the code and a validation-gated deploy contract.

03

Execute

A second agent runs the contract through the Shopify MCP server against staging.

04

Audit

SHA-256, marker counts, length parity, live smoke test. The artifact is re-fetched, never trusted.

05

Publish

Human gate, promotion to live, rollback point retained.

The rule that earns its keep

Never accept an agent’s self-report. Verify the artifact itself — the live URL, a re-fetch from the API, the file on disk. “All checks passed” describes the run an agent thinks it did. In a single session that rule caught a false “all 37 passed”, four corrupted images, two overlapping SVGs, a sprint claim contradicted by its own timestamps, and a cost figure that overstated real spend roughly tenfold.

build manifest — this page
template ................ templates/page.portfolio.liquid
engine .................. native scroll-driven CSS + worker canvas, zero frameworks
progressive enhancement . anchor positioning, view transitions, popover, subgrid, container queries, oklch
human_keystrokes_in_body 0
figures ................. 45 ledger entries, each with grade, method, date and exclusions
internal links .......... 16 targets, all returned HTTP 200 on 2026-07-27 before inclusion
readable without JS ..... yes — every number is in the HTML source

How I run agents

Nine days, one datable change

On 18 August 2026 I audited my own work and found forty-one things I had called “passed” with nothing behind them. A test script pointed at a file that did not exist. Every “10 of 10 tests passed” report I had filed was not produced by that script. I stopped building that day and built the controls instead.

Version control, a numbered ticket protocol, and a rule that every figure carries a grade — MEASURED, CALCULATED or NOT VERIFIED — all enter the record on the same day. They are one change, not three. Nine days later the same method was running in three unrelated codebases. I am not going to call this a month of practice. It is nine days with a start date and a cause, and you can check both.

work order — every ticket, in this order
1 what is settled ... stated first, explicitly closed — or the agent re-argues finished work and bills you for it
2 the ask .......... numbered A1..An, ordered by how much each answer matters. Not a description of a goal
3 the format ....... every answer graded, with a count, a date range, and the file it came from
4 stop conditions .. “do not open new work, write it down instead of fixing it” — and two failures on one item is a stop, not a third attempt
the return trip ... the report lands in the same directory and I audit it. A report that does not say what it could not verify goes back

On VOLTHAVEN 64 that is 39 work orders, 5,103 files of working record and 48 reports. Two other projects run their own series, numbered independently. I am not going to add them together and call it one number.

Three rules I paid for

01

Your probe is the defect

An agent brought me four defects. Three were its own selectors. Now the second selector gets written before the bug report.

02

Verify the bytes you sent

Never a read-back. Shopify’s asset API returns short reads. I lost a deploy cycle to a file that was fine and a read that was not.

03

Pre-register the threshold

A metric that returns the same answer for every input cannot see the defect. If I cannot say beforehand what would falsify the claim, I am decorating, not measuring.

Where it caught something

Four catches, none of them found by reading the code

A discount code, live, five dollars off, with zero products and zero variants attached — queued and stopped before it sent.

A brand guard I had just finished writing refused all eleven drafts on its first run, including copy that had been live since June. The check failed its author’s own work on day one. That is the only kind of check worth having.

A work order proved a game’s audio by calling the music. The game’s only two music calls were fadeOutMusic and stopMusic. Every test passed and the game was silent. The discriminator was to press nothing.

Twenty-three findings in one project traced back to instruments that lied — a camera read mid-move, a reproduction that manufactured its own input, a browser test that returned false regardless of reality.

What it costs, and what is still missing

The largest single waste in the record is mine. I sent an agent to reconcile two line counts without giving it either script or telling it where they had run. It swept two hundred thousand configurations for thirty minutes looking for a difference that was not in any of them. The answer was that one script had counted my own tooling directories. The ticket was the defect, not the agent.

Across eighty-eight reports, “where I wasted effort” appears once. No template requires a cost section, so the cost is systematically under-reported — including in the four catches above. 28.4 percent of my commits are documentation. There is still no CI, no linter and no formatter, eight days after my own audit said there should be.

The honest frame

This is not a rigorous engineering process. These are the controls I built because I have no team — no reviewer, no second pair of eyes, nobody to catch me. They exist because the alternative is trusting myself, and on 18 August I measured what that was worth.

The standing instructions every agent I run receives — excerpt

These are in force for every agent on every task, not a summary written afterwards. This is an excerpt; the working document is longer and also covers model-tier routing, token budgeting and tool-permission limits.

Verification — highest priority, overrides speed

  • Never trust an agent’s or tool’s self-report. Verify the artifact itself: live URL, re-GET from the API, file on disk, command output. “All checks passed” describes the run it thinks it did.
  • Before reporting a failure or an emergency, confirm with a second method. Rate limits, caches and stale mounts produce false negatives. Three independent methods before declaring a capability unavailable.
  • Label every number MEASURED (command or API output), CALCULATED (derived from rates or assumptions) or ASSERTED (stated, no working shown). Never present calculated as measured.
  • Any count needs its exclusions stated and a date. “54,886 LOC excluding node_modules, measured 2026-07-27” beats a bigger number with no method.
  • If unsure, write “Not verified” and what would confirm it. Do not invent facts, numbers, pricing, certifications, partners, policies, results, reviews or timelines.
  • Verification is never what gets cut for budget. Savings come from not re-deriving what is already established — never from skipping a check on something that ships.

Work order — every ticket, in this order

  1. What is settled — explicitly do not reopen. Prevents the agent re-litigating finished work.
  2. The ask — numbered A1..An, ordered by how much each matters.
  3. Format — every answer labelled MEASURED / CALCULATED / NOT VERIFIED, with n and date range.
  4. Stop conditions — including “do not open new work; write backlog items down instead of fixing them.”

A report that does not state what it could not verify is incomplete. Send it back. When a report and my memory disagree, the report’s measurement wins if it shows its method.

Known traps — each one learned the hard way

  • A metric that returns the same answer for every input cannot see the defect. Change the threshold until it can fail.
  • File CreationTime is worthless after a bulk copy — every file shares one instant, later than its own mtime. Timestamps before ~2000 are invalid epochs, not evidence.
  • Single-page apps mount slowly. A zero-element reading before about eight seconds is a false negative, not an absent feature.
  • Cloned repos inherit upstream history. A 2022 first commit in a cloned repo is not my work.
  • Exclude node_modules, browser profiles, build output and binaries from any file or line count.

The same method, applied to teaching

Every Academy class used to carry its own hand-written interactive exercise. Now there is one engine and nine modules — decider, hunt, ledger, instrument, gate, anatomy, forge, promptdiff and setupcheck — in 1,104 lines, generalised out of code originally written for a single class. 136 exercises are deployed from it across 72 pages.

Every module’s CSS sits behind its own conditional, so a page carrying one decider ships none of the other eight modules’ rules. The engine gets more capable without taxing the seventy pages already using it — which is the usual way a shared component library dies, and it was designed against at the start rather than patched later. The grammar fails closed: one malformed row renders zero bytes, no markup and no error, so a broken exercise cannot ship half-working. The cost is that a missing module looks exactly like an absent one, which is the first thing the runbook tells you to check.

Around it sits a twelve-stage class pipeline with the interactive standard as a mandatory gate, and seven numbered promote gates whose comments each name the incident that created them.

VOLTHAVEN 64

I had never built an N64-style game. Thirty-nine work orders later there is one: nineteen rendering techniques, a synthesised soundtrack, gamepad support, and a first-load payload measured at 424,557 bytes — which models to $48.93 per million loads at CDN egress rates.

disclosure — VOLTHAVEN 64
style ............. N64-STYLE. Unlicensed. Not affiliated with or endorsed by Nintendo
hardware .......... does not run on original hardware
telemetry ......... carries first-party telemetry, so the honest statement is no third-party trackers — not “no tracking”
the cost figure ... CALCULATED. A bandwidth model on a measured payload. NOT a claim that a million people played it — there is no such measurement and I do not publish audience numbers

Eighteen months

February 2025 to here

The earliest retrievable artifact anywhere is a t-shirt design uploader, dated 27 February 2025. Eight independent sources agree on it, including the filesystem. 26,779 older files were examined and rejected — all browser-extension and game-cache assets stamped with an invalid 1979 epoch. So this is eighteen months, not the vaguer “two years” this page used to claim. The compressed timeline is the whole story.

In one sentence

From a first Shopify conversation in February 2025 to 86 published masterclasses, six AI systems, two commercial desktop apps and a measured 54,886-line orchestration platform — in eighteen months, self-taught, with every date established by a platform timestamp, an operating-system timestamp or a dated conversation.

Build density — dated artifacts per month

What this measures, precisely: the count of artifacts carrying a verifiable date in that month — application records, first commits, conversation timestamps and operating-system file stamps, across eight sources. It is not a record of hours worked, and no hours figure is claimed anywhere on this page, because no time-tracking data exists to support one. The three empty months in 2025 are real: nothing dated survives from them.

Tool adoption — ten rungs, nine of them timestamped

Each date below comes from an operating-system file stamp or a platform record, not from memory. The one exception is marked, because it should be.

2025-02-27claude.aiFirst retrievable artifact
~2025-04Firebase StudioAsserted — no record survives
2025-07-23CursorOS timestamp
2025-10-07Google AI StudioPlatform record
2025-10-31Claude CodeOS timestamp
2025-11-05CoworkEarliest artifact
2026-02-18Ollama, localOS timestamp
2026-02-26Antigravity + Gemini CLIOS timestamp
2026-04-25claude-maxOS timestamp
2026-05-22Antigravity IDEOS timestamp

Firebase Studio is the honest weak link and it stays labelled as such. I adopted it in its launch month; no artifact from it survives, so the claim is ASSERTED, not measured. What is externally checkable is the product itself: Firebase Studio entered preview on 9 April 2025, was renamed from Project IDX on 15 April 2025, and Google announced its retirement in March 2026 with a final shutdown date of 22 March 2027. That is why it is no longer in the stack.

The record — 38 dated events

2025-02-27 First retrievable conversation anywhere — an HTML t-shirt design uploaderclaude.ai
2025-02-28 First Shopify Liquid work: “cut and paste into the Shopify HTML editor”claude.ai
2025-03-03 First structured data — a JSON-LD ItemListclaude.ai
2025-04-09 Firebase Studio enters preview; adopted in its launch monthasserted
2025-07-12 Virtual Try-On v1 — first brief specifying accessibility, performance and theme version togetherclaude.ai
2025-07-23 Cursor adoptedfilesystem
2025-07-24 CounselAI — a civic housing-court toolclaude.ai
2025-09-14 Architecture shift to Shopify Online Store 2.0 sectionsclaude.ai
2025-10-04 Virtual Try-On deployed to Google Cloud Runclaude.ai
2025-10-07 First Google AI Studio applications — three on one dayAI Studio
2025-10-21 AI Shorts Factory — first of five tools that become VibeTubeAI Studio
2025-10-30 ThriftFind, first of three iterations on one ideaAI Studio
2025-10-31 Claude Code adoptedfilesystem
2025-11-05 Earliest artifact reachable from CoworkCowork
2025-11-13 Two YouTube optimisation tools in one dayAI Studio
2025-11-25 AGI Nexus v10.0 — autonomous publishing officeAI Studio
2025-12-06 The Synthetic Director: build beginsCowork
2026-01-02 PromptDJ and PromptDJ Pro, same dayAI Studio
2026-01-03 PromptDJ MIDI — hardware control addedAI Studio
2026-01-20 First real codebase measurement: wc -l against AGI Nexus v10claude.ai
2026-02-18 Ollama adopted — local inference on owned hardwarefilesystem
2026-02-26 Antigravity 1.0 and Gemini CLI adopted; Atelier OS created the same dayfilesystem + AI Studio
2026-03-05 PromptDJ NS-9000 ships at $79 — fourth generation, 62 days after the prototypeAI Studio
2026-03-08 NICHE-FORGE-CORE v3.0filesystem + AI Studio
2026-03-17 DDS Master Brand Hub created — origin of the flagship orchestratorAI Studio
2026-03-22 Sovereign Orchestrator Pro: first source fileCowork + Antigravity
2026-03-23 JTD Masonry — bilingual site, the only paid third-party client workclaude.ai + Antigravity
2026-03-26 Consolidation: Atelier OS retired, systems reduced from 14 to 13claude.ai
2026-04-08 The Bay State Treasure Gazette — six months after the first ThriftFindAI Studio
2026-04-21 shopify-dds-mastery skill built, hardened, ported and blind-tested 10 of 10 first attemptclaude.ai
2026-05-01 Morning Brief begins. 62 consecutive weekday issues follow, zero missedCowork + Shopify API
2026-05-22 Antigravity IDE adoptedfilesystem
2026-05-31 Evidence standard formalised: every factual claim cites file and line, or is marked NOT VERIFIEDclaude.ai
2026-06-13 110-page inventory published; all valuations stripped from my own public structured dataclaude.ai
2026-06-19 A completed build rejected for skipping the research-approval gateclaude.ai
2026-06-23 Certification error corrected: Fair Trade to Fair Wear Foundation membershipclaude.ai
2026-07-25 Eight Academy classes shipped in one day, taking the curriculum from 70 to 78own verification
2026-07-27 Full evidence audit: 556,586 authored lines measured; six published claims corrected downward, four upwardown verification
How this was built

The timeline, the density bars and the adoption rail are three views of one pipe-delimited dataset. The rail scroll-snaps with two CSS declarations and no JavaScript at all.

.rfx-rail{ overflow-x:auto; scroll-snap-type:x mandatory; }
.rfx-rung{ scroll-snap-align:start; flex:0 0 auto; }

The systems

Six AI systems in production

Working, audited, multi-agent systems. The headline figure this page used to carry — “51,000+ lines across all builds” — was understated by roughly eleven times. One system alone exceeds it.

Flagship · autonomous · in production

Sovereign Orchestrator Pro

A publishing engine that researches, writes and posts across four platforms with no human in the loop. Hybrid inference: local Ollama models on an RTX 3060 for bulk work, Gemini for the hard calls, behind a telemetry-enforced budget gate.

Say this precisely or an interviewer will catch it: the 22 “threads” are interval-timed async jobs inside a single Node.js process. They are not OS threads and not worker threads. The only concurrency primitive is a 15-minute execution timeout. Anyone who assumes OS threads and then reads async/await will rightly stop believing everything else.

TypeScriptOllama, 5 local models RTX 3060Gemini WordPressShopify XReddit Pinterest — wired, gated
Read the case study →

Generative media · measured

The Synthetic Director

AGI-CORE-Pro. Scripts, generates and assembles video and imagery through Veo, Imagen and Gemini. In continuous development since December 2025 through multiple major rebuilds, including a full re-architecture in February 2026.

25,746 LOC measuredDexie / IndexedDB9 modes
Case study →

Portable agent standard

shopify-dds-mastery

A cross-agent skill encoding brand constants, WCAG 2.2 AA limits, Core Web Vitals ceilings and a ten-phase pre-ship checklist. Installed into two different agents and validated by blind build test: a cold agent passed all ten proof-required audit items on the first attempt.

5,529 LOC9 filesBlind-tested 10/10

Never appeared on this page before. It is arguably the strongest single artifact here.

Neural operating system

Atelier OS

A multi-agent orchestration layer coordinating specialised agents into one working environment. Source written across a four-day window in late February 2026, with file-write activity concentrated in roughly 16 hours of active editing.

17,887 LOC measuredRetired 2026-03-26

Correction: this page previously said “architected and shipped in a 52-hour sprint.” The timestamps contradict it — 97.5 hours elapsed, 81 of them idle. The claim is withdrawn.

Compliance automation

claims-compliance-sweep

A four-class classifier with reversible repair that swept 561 live URLs for unsupported claims, behind backup JSON and pre-flight abort gates. Built to police my own marketing copy.

561 URLsReversibleAbort gates

Research to output

NICHE-FORGE-CORE V3.0

A dual three-step pipeline across four Gemini models, turning a niche brief into structured product and content assets. Version dates confirmed independently by filesystem and platform record.

4,281 LOCv2.0 2026-03-06v3.0 2026-03-08
Case study →

Live · customer-facing

Shopify Virtual Try-On

Photorealistic AI garment fitting on apparel product pages. A key-safe Express proxy on Cloud Run keeps the API key server-side. No email, no signup.

React 18 + ViteGCP Cloud RunLive since 2025-10-04
Open the Try-On →

Adoption and conversion effects: not verified. No figure claimed.

Every project, plotted by date

Horizontal axis is time, February 2025 to now. Node size is measured lines of code. Connecting lines are lineage — where one build became the next. Hover or tab to a node for its record.

Node detail appears here on hover or focus. Every project below is also listed in the timeline above, so nothing here depends on JavaScript to be discoverable.

Where the July 2026 measurement found the lines

Tile area is proportional to measured lines. A total with no exclusion list is worth less than a smaller one with a reproducible command behind it, so the exclusions are stated below the map, not buried.

exclusions — declared, because the number is meaningless without them
method .......... stack directory walk, SHA-256 content dedup, newline count from bytes
counted ......... .ts .tsx .js .jsx .mjs .cjs .py .ps1 .rs .go .java .rb .php .cs .sh .sql
plus ............ 229,493 lines of Liquid measured on the LIVE theme via the Shopify Admin API, 143/143 files, 0 failures
NOT counted ..... .json .md .txt .yml .csv .xml .svg — config and docs are not source
removed ......... node_modules, browser profiles, rendered DOM dumps, minified bundles, build output
removed ......... 441,148 lines of vendored Shopify base theme — not my work
removed ......... 2,360,839 lines of byte-identical duplicates across 3,736 files
date ............ 2026-07-27

The first version of this measurement returned 6,491,299 lines. It was wrong, and it was wrong in a specific and instructive way: 83 percent of it was Liquid, because the same Shopify theme sat in five working directories and every copy got counted. The tell was two byte-identical base.css files at different paths. Content hashing removed 2.36 million lines in one pass. A big number with no exclusion list is worth less than a smaller one with a reproducible command behind it — and that applies to my own output first.

How this was built

The treemap is flexbox, not a charting library. Each tile takes flex-grow equal to its line count, so the layout is the data. No JavaScript, no dependency, and it reflows correctly at any width.

<div class="rfx-tile" style="flex:' }} 1 150px">

Live and purchasable

Two commercial desktop apps — and how they got there

Not prototypes. Both ship today, both bring-your-own-key with a serverless licence gate, so inference cost sits with the user and there is no subscription. The part worth showing is not that they exist — it is the dated path from a browser prototype to a paid product.

Generative audio · desktop

PromptDJ NS-9000

$79.00

A bring-your-own-key neural synthesiser. Electron desktop client, MIDI hardware control, serverless licence check.

PromptDJ · 01-02 Pro · 01-02 MIDI · 01-03 NS-9000 · 03-05

Four generations in 62 days. 4,509 lines measured.

Buy PromptDJ →

YouTube automation · desktop

VibeTube AI Desktop App

$49.99

A bring-your-own-key content automation studio in a single Electron app — a 19-section orchestration matrix turning a URL into titles, tags, scripts, thumbnails and cross-posts.

Shorts Factory · 10-21 Shorts Gen · 10-26 Product Shorts · 11-07 SEO + Channel Optimizer · 11-13 VibeTube

Five dated predecessors. This page previously listed VibeTube with no provenance at all.

Buy VibeTube →

Free and live

Three games, a web app, and two civic tools

All shipped, all free, none monetised. The game count on this page said two; three are live. Three further game builds exist and are deliberately not counted here, because they are demos and counting them would be padding.

Browser game · flagship

VOLTHAVEN 64

An N64-style action platformer built from a standing start — nineteen rendering techniques, a synthesised soundtrack and gamepad support. Unlicensed and not affiliated with Nintendo; it does not run on original hardware. It carries first-party telemetry, so the honest statement is no third-party trackers.

39 work orders169 commitsgit
Play VOLTHAVEN 64 →

Browser game

Dragonfly Crush

A polished match-3 puzzle game with no email, no signup and no tracking. Built as a birthday gift, then shipped publicly.

Match-3100% private
Play Dragonfly Crush →

Browser game · own subdomain

Ashaveth

Live on its own subdomain. Notable for the accessibility work: measured contrast lifted from a failing 3.4:1 to 8.56:1, verified by pixel sampling rather than by eye.

Contrast 3.4:1 → 8.56:1WCAG AA
Play Ashaveth →

Live web app

The Bay State Treasure Gazette

A real-time directory of Massachusetts thrift, vintage and antique shops, engineered to be found by both search engines and AI assistants. Six months of iteration from three ThriftFind prototypes.

JSON-LD + llms.txtSPA
Visit the Gazette →

Civic tool · free

Free Legal Help — Housing Defense

A self-serve tool to help tenants defend their homes. A real civic build, live and in use, made with the same agent-directed method.

See the tool →

For Shopify owners · free

One-Prompt App Library

Single-prompt app builds for store owners: copy a prompt, get a working tool. A teaching artifact and a working toolkit at once.

Open the library →

Hard-won

Engineering judgment, from real incidents

The work that does not fit in a feature list. Each entry is an actual production failure — symptom, root cause, fix, and the rule it locked in. This is the difference between writing code and operating systems. Two of them are my own mistakes from the last 24 hours.

Deleting 6,000 legacy posts without wrecking SEO equity
SymptomA retired AI system had auto-published 8,254 articles. Most were disposable; a few carried real equity.
Root causeUnbounded automated publishing with no quality gate and no inventory.
FixHTTP 410 Gone for the bulk — faster de-indexing than 404, equity deliberately forfeited — with 301s for only the top 100, behind backup JSON and pre-flight abort gates. 6,163 deleted, 100 redirected, 1,995 kept, in 1:38:01, zero errors.
LessonDestructive mass operations need restore points and tolerance checks before they run, not after.
Storefront filters returned zero results for every category
SymptomCategory filtering silently matched nothing across the whole catalogue.
Root causePrint-on-demand assigned one product_type — “Print Material” — to every product, so the field had no variance to filter on.
FixMoved category detection to the product title and gated the certification badge to apparel only.
LessonVerify the shape of the data a platform actually emits before trusting a field to carry meaning.
A model upgrade silently killed every posting workflow
SymptomThe Synthetic Director stopped posting. No error surfaced.
Root causeModel IDs hardcoded across 22 files; a duplicate <script> tag double-mounting React; a 35 KB brand file orphaned from the runtime.
FixCentralised the model IDs, removed the duplicate mount, rewired the brand file.
LessonHardcoded vendor identifiers are a time bomb. Config belongs in exactly one place.
Multi-model inference thrashed a single consumer GPU
SymptomA local multi-model matrix stalled under concurrent load.
Root causeConcurrent model calls oversubscribed 12 GB of VRAM on an RTX 3060.
FixAsync-mutex serialisation of model calls and a num_ctx budget of 4096. Production inference on hardware already owned.
LessonTreat VRAM as a shared resource with a scheduler, not an infinite pool.
Tooltips hid behind their neighbours on the Academy map
SymptomMap tooltips rendered underneath adjacent nodes.
Root causetransform: translate on each node created a new stacking context, trapping the tooltip inside it. Flip-left and flip-right class bodies were also swapped.
FixNegative-margin offsets and two targeted patches, no rebuild. The evidence receipts on this page use CSS anchor positioning, which is the structural fix for that whole class of bug.
LessonKnow which CSS properties create stacking contexts before reaching for them.
Mobile performance plateaued and would not move
SymptomMobile scores stopped improving regardless of what was optimised.
Root causePlatform-injected scripts set a hard floor on Total Blocking Time that no page-level work can cross.
FixOptimised to the floor, then stopped and documented where the ceiling is.
LessonKnowing where the ceiling is prevents wasting weeks chasing a number the platform will not permit.
Duplicate structured data across theme sections
SymptomMultiple Organization and ItemList blocks emitted on the same page.
Root causeSeveral sections each independently emitted their own schema.
FixOne schema, one owner. The layout owns Organization, WebSite and Breadcrumb; the hero owns the storefront types; the SEO manager owns ItemList and HowTo.
LessonStructured data needs a single source of truth per type, exactly like any other state.
An agent skill could hallucinate brand facts
SymptomAgents invented certifications and shipped code that broke house rules.
Root causeNo encoded ground truth for the brand or the platform’s limits.
FixWrote a 5,529-line brand-rules skill, then validated it with a blind build test: a fresh agent with no context passed all ten proof-required audit items on the first attempt.
LessonProve a skill against a cold agent. Do not trust it because you wrote it.
A shell variable clobbered an auth header mid-deploy
SymptomA deploy script failed loudly partway through, after staging had been written but before production.
Root causePowerShell variables are case-insensitive. A loop variable $h silently overwrote the $H headers hashtable holding the API token.
FixRenamed the loop variable. More importantly: the script failed before the production write rather than halfway through it, because the deploy order is staging first, production second, with a parity assertion between them.
LessonOrder your deploy so that the failure mode you have not thought of yet lands on the safe side of the gate.
My own measurement returned 6.5 million lines, and it was wrong
SymptomA line-count script reported 6,491,299 lines across my projects — a number that would be disbelieved on sight and would poison every figure next to it.
Root cause83 percent of it was Liquid. The same Shopify theme sat in five working directories and every copy was counted. The tell was two byte-identical base.css files at different paths.
FixSHA-256 content hashing so each unique file counts once globally, plus authorship bucketing that separates my code from the vendored base theme. It removed 2,360,839 duplicated lines in one pass.
LessonI had already rejected someone else’s 9.6-million-line count for exactly this error, then reproduced it myself an hour later. The rule only works if it applies to your own output first.
How this was built

The category filter is pure CSS. Radio inputs plus :has() on the wrapper — no JavaScript, fully keyboard operable, and it still works if scripts fail to load.

.rfx-conwrap:has(#rfx-f-css:checked) .rfx-inc:not([data-cat~="css"]){ display:none; }

The part that took the longest

I used to publish numbers I could not defend

Most “AI-assisted developer” claims are unfalsifiable, and mine were too. What follows is a dated, documented reversal toward a stricter evidence standard — made against my own commercial interest, because it cost me every impressive number I had. Nobody asked me to do it.

Jan 2026 Publishing agency-style figures on live pages — ROI percentages in the millions, dollar valuations, labour-replacement estimates. All of it derived from documents I had generated myself. None of it observed.
2026-05-31 Formalised the rule I still work to: every factual claim must cite a file and line, or be explicitly marked NOT VERIFIED. Written down, applied to agents and to myself.
2026-06-13 Stripped every valuation out of my own public structured data. That is search visibility and pitch material deleted on purpose, because I could not source it.
2026-06-19 Rejected a completed, working build for skipping the research-approval gate. The output was fine. The process was not.
2026-06-23 Corrected a certification error I had propagated myself — and traced it to my own agent instructions rather than blaming the agent that repeated it.
2026-07-27 Full evidence audit across eight sources. 30 published figures re-derived: 14 corrected upward, 7 downward, 9 held exactly, and three claims withdrawn outright rather than restated. This page is the result.

The receipts, still online

It would be easy to claim a reversal and quietly delete the evidence. These two case studies are from the old standard and they are still published, unedited. The valuations are baked into their URLs, which is why they read the way they do.

You are about to see “$15.5M”, “built in 52 hours”, “508,800% ROI” and “1,512,200% ROI”. I wrote all of it. None of it was measured. The systems are real — Atelier OS is 17,887 lines and AGI Nexus shipped — but the money figures were cost-to-replicate estimates dressed up as value, built on invented inputs. $15.5M across 52 hours works out to roughly $298,000 an hour, which is the kind of thing that only survives if nobody does the division.

Atelier OS, as published in early 2026 →
AGI Nexus v9.0, as published in early 2026 →

They stay up deliberately. A reversal you can only take my word for is worth nothing; one you can click into and check for yourself is worth something. Both are excluded from this page's structured data — linking them in prose with this warning attached is a different act from feeding those numbers to a machine as endorsed facts.

A big number with no exclusion list is worth less than a smaller one with a reproducible command behind it. The rule I now apply to my own output first

If you are hiring an AI systems architect, the core competence is not prompting. It is knowing which outputs to trust and having a method for deciding. The most direct way I can demonstrate that is on my own claims, in public, including the ones that did not survive.

Both directions

What the audit changed

Every figure on the previous version of this page was re-derived from source on 27 July 2026. The corrections run in both directions, which is what an honest audit looks like — a padded one only ever corrects downward, and a defensive one never corrects at all.

556,586 could not be re-run 772,119 authored lines of code
327,093 505,949 lines of application source
229,493 was unreproducible 266,170 lines of Liquid running in production
~26,900 54,886 lines · Sovereign Orchestrator Pro
104 files 258 source files · Sovereign Orchestrator Pro
86 classes 86 free Academy classes
62 issues 64 consecutive weekday blog issues
4 games 4 live browser games
12 agents 12 registered AI agent personas
24 threads 22 scheduled publishing jobs
$8–13 per day $1.19 measured API spend per day
$0/mo hosting $0 per month for hosting
619 618 active products
129 138 store pages
18 months 18 months, start to here
7 schemas 10 JSON-LD schemas per page
4 platforms 4 live publishing platforms
3 of 23 3 of 23 npm package roots under version control
87 was overstated 80 Lighthouse performance, mobile, cold
100 was optimistic 99 Lighthouse performance, desktop, cold
3,500 ms 4.0 s Largest Contentful Paint, mobile cold
0.1604 0.000 Cumulative Layout Shift, mobile
39 tickets 39 work orders on one project
9 modules 9 interactive modules, one engine
136 live 136 exercises deployed from that engine
7 gates 7 numbered gates before anything promotes
$48.93/1M $48.93 modelled cost per million loads
99 100 Lighthouse accessibility, mobile
100 100 Lighthouse SEO, mobile and desktop
92 96 Lighthouse best practices

14 understated · 7 overstated · 9 held exactly

14 of these were understated, and that is the part worth pausing on. The class count was 59 and is 81. The blog count was 25 and is 62. The code figure was 51,000 lines against 556,586 measured. This portfolio was not padded — it was unaudited. Those are different failures with different fixes.

The two findings I would rather not publish

A content-hash audit of my own drive found 2,360,839 lines duplicated across 3,736 byte-identical files, and only 3 of 23 project roots under version control — the 54,886-line flagship is not one of them.

To be precise about what this is and is not: backups exist and are deliberate — versioned zip archives plus a monthly full-drive copy to external storage. What is missing is version control, which is a different thing. Backups protect against loss. Git gives you history, diffs, blame and the ability to bisect a regression. The duplicate files are the visible cost of having the first without the second.

It is also the first thing I would fix, and I would rather you read it here than find it in an interview. Concealing it would be worse than having it — and a page arguing for evidence standards that quietly omitted its own worst number would not be worth reading.

How this was built

The filter above uses no JavaScript. A checkbox plus :has() on the wrapper hides every row that is not a downward correction. The counts beside it are computed in Liquid from the same ledger the rows come from, so the tally cannot disagree with the table.

.rfx-diffwrap:has(#rfx-onlywrong:checked) .rfx-drow[data-dir="up"]{ display:none; }

Always on

An autonomous operations fleet

Beyond the headline systems, scheduled agents run the day to day with the human out of the loop — research, publishing, video and engagement — reporting in rather than waiting for instructions.

Daily · 62 issues, 0 missed

Morning Brief

An autonomous research-to-publish agent. Every issue ships with FAQ, breadcrumb and speakable structured data and a 1200×630 CDN cover, with no manual authoring at any stage.

Count verified by Shopify’s own GraphQL articlesCount at EXACT precision — a platform figure, not a self-report.

Daily

Video build

A scheduled pipeline that drives a browser video editor end to end — builds, exports and downloads finished shorts — every day, unattended.

Continuous

YouTube engagement

Agents that verify scheduled videos actually went public, then post and pin formatted comments — and yield safely when another agent holds the write lock.

Release engineering

Deploy toolchain

Roughly 100 PowerShell scripts with three hard ABORT guards that fire before any write to the live theme, plus staging-to-production length-parity assertions and a retained rollback point.

Teaching the method

The DDS Vibe Academy — 86 free classes

I teach the method openly. The Academy is a free curriculum on intent-based AI coding, built with the very workflow it documents. Nothing is paywalled and nothing is gated: no signup, no email capture, no certificate. This page said 59 classes. It reached 78 on 25 July and is 81 today.

Curriculum

86 classes, 70 unique URLs

From vibe-coding foundations through agentic builds, MCP servers, sovereign RAG and the SEO Magnet System. Four stages, ten lanes, four guided tracks.

Architecture

One dataset, 16 surfaces

A single snippet holds the class dataset and renders roughly sixteen different surfaces — the hub grid, the lane pages, the What’s New tab, the search index and the structured data. Adding a class is one row.

Access

Free, permanently

Built as a teaching asset, not a funnel. There is no upsell at the end and there was never going to be.

Enter the Academy →

The storefront

A production Shopify store, engineered solo

Design Delight Studio is itself proof of work: a full sustainable-apparel catalogue on Shopify Basic, built and tuned by one architect, and the environment every constraint on this page was learned in.

Catalogue

619 active products

Made-to-order apparel and accessories across 129 pages. 683 including drafts and archived — this page previously said 684, which counted drafts as live.

Production code

229,493 lines of live Liquid

Measured directly from the live theme through the Shopify Admin API: 143 authored files, 143 retrieved, zero failures. Sections, templates, snippets and blocks.

Accessibility

WCAG 2.2 AA

Twenty-plus sections rebuilt for semantic structure and keyboard support. Contrast verified numerically with pixel sampling, never by eye.

Discovery

10 JSON-LD schemas per page

The SEO Magnet System: structured data, answer capsules, speakable blocks, and an FAQ accordion whose content must match its FAQPage schema exactly or the build fails.

The method, published free →

Compliance

561 URLs swept

A four-class classifier with reversible repair, run across the live site to find claims that could not be supported — including my own.

Measured 2026-08-03

Lighthouse 87 mobile, 100 desktop

Cold-load medians of three runs each, re-audited the day this page gained a live WebGL section. The cold number is what a first-time visitor gets, so that is the one in the heading. The mobile 87 is held down by a 3,500 ms LCP that still fails Google's threshold — full table, spread and caveats in the next section.

TBT 0 msCLS 0.000SEO 100

Measured, not claimed

What this page actually scores

Every previous version of this portfolio published performance figures that were never re-measured after the rebuild. So this page is audited before the numbers are written — three runs per configuration, median reported, spread disclosed. It was re-audited on 3 August 2026, after this page gained a live WebGL section, because publishing the old numbers next to new code is the precise failure this section exists to prevent.

RunPerfA11yBest prac.SEOFCPLCPTBTCLS
Mobile · cold8799921002,600 ms3,500 ms0 ms0.000
Desktop · cold1009992100600 ms700 ms0 ms0.000

Median of three runs per row, 3 August 2026, PageSpeed Insights — Lighthouse 13.4.1, emulated Moto G Power on slow 4G for mobile, single page session, initial page load. Spread across the three runs: mobile Performance 82–87, LCP 3.5–3.8 s, Speed Index 2.6–4.7 s, CLS 0–0.019; desktop Performance 100–100, LCP 0.7–0.7 s, Speed Index 0.6–0.7 s. Lighthouse is noisy, so a single sample would not have been worth publishing — the mobile Speed Index alone moved 2.1 s between identical runs.

Warm rows have been removed rather than carried forward. The earlier audit ran Lighthouse from Node with disableStorageReset to measure a repeat visit; PageSpeed Insights only performs a cold single page session and cannot reproduce that. Leaving the old warm numbers in a table dated today would have implied they were re-measured. They were not, so they are gone.

The number I would rather not lead with

Mobile cold LCP is 3,500 ms. Google's threshold for “good” is 2,500 ms, so it still fails, and it is the reason the mobile score is 87 and not green. It improved from 3,823 ms in the previous audit and desktop cold is 700 ms, but a first-time visitor on a phone still waits three and a half seconds for the largest paint. It stays on the page because a performance section that printed only the 100s would be the exact behaviour this portfolio was rebuilt to stop.

One honest limit on that comparison: the previous figures came from Lighthouse driven via chrome-launcher and these came from PageSpeed Insights. Same tool, different harness and different hardware, so the 82 → 87 movement is consistent with the work done to this page but is not proof of it. What is not in doubt is that these are the numbers this page scores today, measured the way the note above describes.

Why Total Blocking Time is zero — and why that is not a trick

A 0 ms TBT on a 194 KB page under mobile CPU throttling deserves suspicion, so here is the actual mechanism rather than a flattering one. It held at 0 ms across all six runs on 3 August, mobile and desktop, after the page gained a WebGL section.

verified against the rendered HTML
template directive ...... {% layout none %} — the page never loads theme.liquid
theme JS bundle ......... absent — zero references to global.js, theme.js, constants.js, pubsub.js, cart.js
stylesheet requests ..... 0 — all CSS is inline
external scripts ........ 2, both Shopify platform-level, neither is the theme bundle
inline executable JS .... 35,693 bytes, deferred and intersection-gated
canvas work ............. OffscreenCanvas in a Web Worker, off the main thread entirely
WebGL section ........... never requested on a Lighthouse pass — Three.js is imported inside the intersection callback, and Lighthouse does not scroll

That last line is the one worth checking rather than trusting. Held at the top of the page for eight seconds on a cold load, the network log shows 12 requests and no three.module.js; the engine has not mounted and the project map has not been built. Both are gated behind IntersectionObserver, so a visitor pays for them only if they scroll far enough to see them.

This does not contradict the earlier finding that Shopify imposes a Total Blocking Time floor. That floor is real — on pages that render through theme.liquid, which is most of the store. This template opts out of the theme layer entirely, so it never pays that cost. Two true findings about two different page types, which is worth stating precisely rather than collapsing into one flattering claim.

Two honest caveats. The accessibility 99 is Lighthouse's automated score, which checks what a machine can check — it is not a WCAG conformance statement, and the manual keyboard and screen-reader passes are separate work. Best Practices sits at 92 because of a console 404 from a Shopify app embed that this template does not load and cannot remove.

How this was built

The 3 August re-audit ran through PageSpeed Insights — Lighthouse 13.4.1 on Google's own infrastructure, three runs per configuration, medians and spread taken from the raw category and audit values rather than read off the dials. Using Google's runner instead of a local one removes my machine from the result entirely, which matters more here than matching the previous harness exactly.

const r = await psi(url, strategy);              // strategy: mobile | desktop
const c = r.lighthouseResult.categories, a = r.lighthouseResult.audits;
runs.push({ perf: c.performance.score * 100,
            lcp:  a['largest-contentful-paint'].numericValue,
            tbt:  a['total-blocking-time'].numericValue,
            cls:  a['cumulative-layout-shift'].numericValue });
report(median(runs), spread(runs));              // never a single sample

The earlier figures on this page came from Lighthouse driven locally via chrome-launcher, which is also how the warm rows were produced. That harness is recorded here so the two sets of numbers are not mistaken for like-for-like.

The stack

Skills, mapped to proof

A skills list is a claim. Each row below names the artifact that backs it, so the claim is checkable.

  • Multi-agent system architectureSovereign Orchestrator Pro — 20 agent modules, 86 services, 54,886 lines measured
  • Agent skill authoring & portabilityshopify-dds-mastery — installed into two agents, blind-tested 10 of 10 first attempt
  • Autonomous pipeline operationsMorning Brief — 62 of 62 weekdays, verified by platform count
  • Release engineering~100-script toolchain, preflight gating, staging-to-production parity, three hard ABORT guards
  • Shopify Liquid architecture229,493 live lines; one snippet dataset driving ~16 rendered surfaces
  • Structured data & technical SEO10–13 JSON-LD schemas per page; FAQ-to-schema parity enforced at build
  • Accessibility engineeringWCAG 2.2 AA; contrast measured 3.4:1 → 8.56:1 by pixel sampling
  • Test & audit engineering22 invariant tests; 40-gate and 38-gate audits at zero failures
  • React + TypeScriptcase-study dashboard under TS strict, clean; bilingual client site
  • Compliance automationclaims-compliance-sweep — 4-class classifier, reversible, 561 live URLs
  • Hybrid local / cloud inferenceGemini plus local qwen2.5-coder:14b on an RTX 3060; telemetry-enforced budget gate
  • Cloud deploymentTwo independent GCP Cloud Run deployments
  • Evidence governanceThis page. Valuations stripped from my own schema at commercial cost.
TypeScriptReact 19Electron ViteTailwindNode.js PythonPowerShellFirebase GCP Cloud RunGraphQLIndexedDB / Dexie Shopify LiquidOS 2.0 sectionsAdmin API MCP serversOllamaGemini ClaudeRAG + embeddingsJSON-LD WCAG 2.2 AACore Web VitalsGit

Experience

Founder & Architect — Design Delight Studio

Boston, MA · February 2025 – present · AI Systems Architect

  • Designed and operate an AI development pipeline in which agents write, audit and deploy production code under written constraints, with zero human keystrokes in the production pipeline and an audit stage that can fail the build.
  • Shipped 772,119 authored lines of code across six production AI systems, two commercial desktop apps, three browser games and a live web app — solo, measured with declared exclusions.
  • Built and operate a fleet of scheduled autonomous agents; the publishing agent has shipped 62 consecutive weekday issues with zero misses, verified by platform count.
  • Built and run a 619-product Shopify Basic storefront across 129 pages and 229,493 lines of live Liquid, rebuilt to WCAG 2.2 AA.
  • Created the free 86-class DDS Vibe Academy and a 5,529-line portable agent skill validated by blind build test.
  • Authored a compliance classifier that swept 561 live URLs for unsupported claims, and applied it to my own marketing copy first.
  • Manage four supply-chain certifications plus Fair Wear Foundation membership across a made-to-order catalogue.

The brand

Four certifications, plus one membership

The distinction matters and this page used to get it wrong. Four are third-party certifications. Fair Wear Foundation is a multi-stakeholder membership covering factory labour conditions, held through our supplier — not a certification, and it should never be described as one.

GOTSGRSOCS PETA-Approved Vegan Fair Wear Foundation — membership

Design Delight Studio is not OEKO-TEX certified. See the certifications page →

Questions

FAQ

Who is Robert McCullock?
Robert McCullock is a Boston-based AI systems architect and full-stack engineer, and the founder of Design Delight Studio. Working solo, he writes specifications and acceptance criteria and directs AI agents to build, test and ship production software.
How much of this page is measured rather than claimed?
Every figure carries a grade. MEASURED means it came from a command or a platform API. CALCULATED means it was derived from rates or assumptions and never directly observed. ASSERTED means it is stated without working shown. Click any number to see its grade, method, date and exclusions.
How does one person ship this much software?
Through a constraint-driven pipeline. Robert writes the specification and the acceptance criteria; one agent authors code and a validation-gated deploy contract; a second executes it through the Shopify MCP server; an audit stage runs SHA-256 and marker checks plus a live smoke test before a human gate promotes it. No human keystrokes enter the production pipeline.
What is Robert's technical stack?
React 19, TypeScript and Electron on the frontend; Node.js, Python, Firebase, GCP Cloud Run and IndexedDB on the backend; Shopify Liquid, Online Store 2.0 sections and JSON-LD on commerce; and multi-agent orchestration across local Ollama models, Gemini and Claude, including MCP server development in both Python and TypeScript.
Are the systems actually in production?
Yes. Sovereign Orchestrator Pro runs 22 scheduled publishing jobs across four live platforms. The Morning Brief agent has published 62 consecutive weekday issues with zero missed weekdays, a figure verified by Shopify's own GraphQL article count at EXACT precision. The Virtual Try-On runs live on apparel product pages.
What did the audit get wrong?
The audit re-derived 30 previously published figures from source. 14 were understated, 7 were overstated, and 9 held exactly. The largest understatement was the code figure: 51,000 lines published against 556,586 measured. Separately, three claims were withdrawn rather than restated, including a 52-hour sprint contradicted by its own file timestamps and every dollar valuation. A content-hash audit of the drive also found 2,360,839 duplicated lines and only 3 of 23 project roots under version control.
What are the two commercial apps?
PromptDJ NS-9000 at 79 dollars and the VibeTube AI Desktop App at 49.99 dollars, both sold on ddsboston.com. Both are bring-your-own-key Electron applications with a serverless licence gate, so inference cost stays with the user and there is no subscription.
What is the DDS Vibe Academy?
A free 86-class curriculum on AI-assisted coding, built end to end with the same agent-directed workflow it documents. There is no signup, no email capture, no paywall and no certificate.
How long has Robert been doing this?
Approximately 18 months. The earliest retrievable artifact anywhere is dated 27 February 2025 and eight independent sources agree on it. 26,779 older files on the drive were examined and rejected as browser and game cache stamped with an invalid 1979 epoch.
How do I work with Robert?
Email Robert@ddsboston.com. He is open to senior AI engineering and architecture roles, consulting and partnerships. The case studies and the free Academy document the method in depth.

In one screen

Key takeaways

  • 772,119 authored lines measured with declared exclusions on 2026-07-27 — six AI systems, two commercial desktop apps, three games and a live web app.
  • Eighteen months, self-taught, from a first Shopify conversation on 2025-02-27 to a ten-tool agent stack, with every adoption date but one established by an OS or platform timestamp.
  • 62 consecutive weekday publications with zero misses, and 86 free classes — both verified against the platform, not asserted.
  • Corrections in both directions: of 30 figures re-derived from source, 14 were understated and 7 overstated. The portfolio was not padded, it was unaudited.
  • Published my own worst numbers — 2,360,839 duplicated lines and 3 of 23 repositories under version control — because a page arguing for evidence standards that hid its own gaps would not be worth reading.
  • Every number here is clickable and returns its grade, method, date and exclusions.

One architect, directing AI agents under tight constraints, shipping and operating production software at the scale of a small company — and holding the output to a standard that costs him claims. If that is the kind of builder you need, let’s talk.

Work with Robert

Let’s work together

Open to senior AI-engineering and architecture roles, consulting and partnerships. The fastest path is a direct email — or dig into the proof first.

Boston, MA · (617) 334-5912 · Robert@ddsboston.com