“Don't tell me what the user wants. Tell me what they fear. That's where the margin is.”
Sets strategy from raw signal. Search-grounded. Everything downstream inherits its directive.
Design Delight Studio · Boston · Investor & partner brief
Design Delight Studio is a sustainable streetwear brand. It is also an autonomous production company — a twelve-agent system that finds a market signal, forms a strategy, drafts the work, holds an internal argument about it, blocks any claim it cannot source, and publishes across four platforms. It has done that for 67 consecutive days without a human touching it. The storefront is not the business. It is the proof the machine runs.
Every number here is graded and clickable — open one for its receipt
Design Delight Studio is a Boston sustainable-streetwear brand and AI engineering studio founded and operated by one person, Robert McCullock. Its storefront, publishing and product pipelines run on Sovereign Orchestrator Pro — a twelve-agent system with eight psychometric staff personas that research, draft, negotiate with each other, and gate their own output before it reaches a customer. The system has published for 67 consecutive days unattended, stopped 73% of its own blog publish slots before anything shipped, and runs at a cost of about a $5 per-day ceiling — $10 after a sale — that has never been hit. Every figure on this page carries an evidence grade and a receipt.
MEASURED means counted. CALCULATED means derived from something counted. ASSERTED means stated without proof. As of 3 August 2026 every figure on this page is MEASURED — nothing here is calculated or asserted. That became true today, the hard way: the last CALCULATED figure — our own $1.19 estimate of daily cost — was replaced by an invoiced one, and then the invoiced one was withdrawn too, because it measured the whole billing account rather than this system. A figure that is real but attributed to the wrong thing is not MEASURED, so it is not on this page. If a calculated or asserted figure ever appears here it will be labelled as one, in colour, and it will say why.
01 · The staff
Most AI systems are one model wearing different hats. Ours is a staff. Each of these carries a personality profile, a core drive, a defense mechanism, memory of what the others have done, and a documented relationship with every one of them. They are not prompts with names on. They disagree, on the record, before anything ships.
“Don't tell me what the user wants. Tell me what they fear. That's where the margin is.”
Sets strategy from raw signal. Search-grounded. Everything downstream inherits its directive.
“This claim in section 3 is unverifiable. Cut it. Precision is persuasion.”
The truth gate. Every factual claim must resolve to a verified source or it is flagged for removal — never silently rewritten. It has hard-halted 28 drafts on legal-class findings and is one of four layers that stopped 491 publish slots.
“The rhythm is off in paragraph two. It stumbles. Let's make it sing.”
Writes the long-form draft. Then defends it in the War Room, and loses often.
“The AOV is too low. We are shipping air. Denied. Bundle it or kill it.”
Vetoes products on unit economics before anything is written about them. A margin gate that runs upstream of the copy, not after it.
“Target acquired. JSON-LD extracted. The source is a Collection, not a Product.”
Runs before strategy exists. Surfaces raw external signal so ATLAS is reasoning about the world rather than about itself.
“Beautiful prose, Scribe, but the H1 tag is missing the primary keyword. Google won't read poetry.”
Grades the finished draft for SEO and schema health as a pre-publish condition.
“'Sustainable' is not a visual. 'Unbleached cotton under harsh noon sun' is a visual.”
Art direction, plus every accessibility alt-text and social image description the system emits.
“I hear what you're saying about the price point. Let me share exactly how we source this.”
The only agent optimised for being liked rather than being right. Deliberately.
Before a draft ships, three agents run a tri-agent negotiation. ATLAS argues
the strategy, SCRIBE defends the writing, VERITAS attacks anything unsupported. The function
returns a negotiation_log, a disagreement_flag, and a
revised_draft.
They also carry episodic memory of one another, and their sampling temperature is modulated by live stress and confidence state — a system under pressure literally becomes more conservative.
Why it matters commercially: the expensive failure mode of generative AI is confident, fluent, unsupported output. An adversarial step with a dedicated sceptic is the cheapest known defence, and it runs on every asset.
02 · The machine
This is the loop that runs every day. The interesting parts are not the generation steps — everyone has those. The interesting parts are the three gates, drawn in red, that exist purely to prevent output.
This animation is a simulation of the pipeline’s shape, not a feed of production telemetry. The counters above tick because the simulation is running in your browser right now. The real figures are the 491 refusals below, and they have receipts.
Raw external signal is surfaced before any strategy exists, so the system reasons about the world rather than about itself.
ARGUSSignal becomes a campaign directive, search-grounded. Everything downstream inherits it.
ATLASProducts failing unit economics are killed here — before a word is written about them, not after.
LEGIONLong-form copy is written against the directive and the brand memory.
SCRIBETri-agent negotiation. Returns a disagreement flag and a revised draft, or escalates.
ATLAS · SCRIBE · VERITASAny claim that does not resolve to a verified source is flagged for removal. Then cadence, novelty and quality gates run cheapest-first, fail-fast.
VERITASSEO and schema health graded, accessibility text written, asset published to the scheduled slot.
HERMES · LUMIEREIf any gate blocks, the system skips the slot. It does not lower the bar and publish something weaker. The comment in the source reads: “An empty day beats a thin/duplicate farm post.”
That is the whole philosophy in one line. A content system that cannot say no is a liability generator, and most of them cannot.
03 · The gate
Every founder claims AI leverage. Almost none can show the machinery that catches their own system lying, with a number on it. This is ours, and here is the breakdown by category rather than a single flattering total.
$ grep -c "slot SKIPPED|slot APPROVED|HALT|B3 BLOCK" logs/app-*.log window 2026-05-30 .. 2026-08-04 (67 daily logs) ───────────────────────────────────────────────── slot APPROVED .......... 184 reached a customer slot SKIPPED ........... 491 stopped by the gate stack pre-flight auto-fixed .. 296 corrected before any model saw it B3 substance block ..... 66 too thin to publish VERITAS HALT ........... 28 legal-class finding, hard stop ───────────────────────────────────────────────── REFUSAL RATE ........... 73% of blog publish slots $
“Every factual claim — certification, statistic, review, result — must resolve to a verified source in the product context or brand memory. If a claim is unverified, uncited or fabricated, you must flag it for removal, not silently rewrite it.”
Flag rather than rewrite is the deliberate part. A model that quietly fixes its own bad claim teaches you nothing and hides the failure rate. This one leaves a record, which is why there is a number to publish at all.
The blacklist is hard-coded and dated. Fair Trade claims were banned on 3 July 2026 after we checked our supply chain and found we had never held the certification. Worldwide-shipping claims were banned on 22 July 2026. Each ruling stays in the source with its date, so a rule can be audited back to the day it was made.
Applied to ourselves
A four-class classifier with reversible repair, run across our own live site to find claims that could not be supported — including ours. Reversible because an automated fix that cannot be undone is a second failure waiting to happen.
Portable
The Shopify build discipline extracted into a portable skill and blind-tested 10 out of 10 by a cold agent with no prior context. Evidence the method transfers rather than living in one person's head.
Verified
Full suite, zero failures at time of measurement. Stated with its date because a test result without one is decoration.
04 · Unit economics
This is the second question every investor asks. The answer here is narrower than the one this page carried four hours ago, because we grouped the invoice by project and found we had been attributing a whole portfolio’s spend to a single system.
First correction, 19:40. This page had published $1.19 a day,
graded CALCULATED, from our own token counting, carrying a standing admission that it had never
been reconciled against a bill and that costService.ts had four documented
under-count vectors. We opened the invoice. The vectors were real: billed spend across 1 May to
31 July 2026 was $285.59 over 92 days. We replaced $1.19 with $3.10 a day.
Second correction, 20:35 — and this one was ours, not the software’s. $3.10 a day is the whole billing account divided by 92. It is not what it costs to run the publishing system, because that account carries every Google Cloud project DDS owns. Grouped by project, $250.52 of the $285.59 — 87.7% — is Veo Weaver, a video-generation project that has nothing to do with the agent loop described on this page. Presenting it under this heading was over-attribution, and it was ours.
So the figure is gone rather than adjusted. What is measured is on the cards above: the whole estate, the share that is video, and what everything else costs. The isolated cost of the publishing system is not stated on this page, because it is not yet attributable. The six projects on the billing account do not map cleanly onto it, and the operator’s own estimate of roughly $1.20 a day is exactly the kind of number this page refuses to print without a receipt — it is also known to be bimodal, since newsletter runs cost materially more than ordinary days.
Two corrections, four hours apart, both published in place with times on them. The first was our telemetry under-reporting. The second was us over-claiming a real number by putting it under the wrong heading. The second is the more instructive one, because nothing was fabricated — an accurate figure was simply asked to mean more than it did. That is the failure mode this page exists to demonstrate catching.
$ gcloud billing accounts describe 017A1A-509220-DFB891 --costs charge period 2026-05-01 .. 2026-07-31 (92 days, 3 full months) ───────────────────────────────────────────────── veo-weaver ............ $250.52 video generation — NOT the publishing loop try-on-app ............ $18.30 virtual try-on thrift-finder ......... $16.77 separate product dragonfly-crush ....... $0.00 free tier thrift-finder-476801 .. $0.00 free tier getpaidai ............. $0.00 free tier ───────────────────────────────────────────────── TOTAL BILLED .......... $285.59 tax $0.00, 92 days VIDEO SHARE ........... 87.7% one project dominates the account EX-VIDEO / DAY ........ $0.38 still not the publishing system alone sop_isolated .......... UNATTRIBUTED no project maps to it — not published $
What is not in doubt: inference runs on local hardware — a single consumer RTX 3060 running Ollama for bulk work, with Gemini reserved for the calls that need it. That routing decision is why serverless compute across those three months came to one cent — the one part of the bill that unambiguously belongs to this system — and why the cost ceiling is a design parameter rather than a hope.
05 · The proof instance
The agents, the gates, the scheduler and the publishing adapters are general. Only the brand memory, the compliance blacklist and the product context are apparel-specific, and those are configuration rather than architecture. The storefront is what a live instance of the machine looks like after seventeen months.
Storefront
Built and maintained by one person directing agents, on Shopify Basic. Product drafts, mockups and SEO copy are generated by the pipeline and synced through the Shopify Admin API and Printful, rate-limit aware.
Taught openly
The method is published, not hoarded. No signup, no email capture, no paywall, no certificate. A company confident in its execution can afford to give away its method, and the Academy is the argument that we are.
Commercial software
PromptDJ NS-9000, $79 — a real-time neural music synthesiser driven by
MIDI hardware. VibeTube AI, $49.99 — a native Windows YouTube strategy
engine on a bring-your-own-key architecture, so it carries no API cost for us. Both are sold as
licensed desktop applications and both were confirmed ACTIVE against the live
Shopify Admin API on 3 August 2026. An earlier version of this page claimed only one, because
VibeTube’s listing could not be verified at the time.
Free, and staying free
Ashaveth, an AI-driven dark-fantasy RPG with generative NPCs, grew out of Legend of the Red Dragon 2026 — a playable beta never launched commercially for copyright reasons. Dragonfly Crush was built as a birthday gift and shipped anyway. Three further builds exist and are deliberately not counted, because they are demos and counting them would be padding.
Four certifications: GOTS, GRS, OCS and PETA-Approved Vegan, held product by product rather than brand-wide. Plus one membership: a Fair Wear Foundation supplier membership through Stanley/Stella — a membership, not a certification, and never described as one.
Design Delight Studio does not hold Fair Trade and has never held it. It is not OEKO-TEX certified at brand level. An earlier version of this page said “five certifications” in a headline sitting directly above a paragraph that correctly said four plus a membership. That headline was wrong and has been corrected.
06 · Diligence
A page that lists only strengths is not evidence of anything. These are the things a technical buyer would find in the first afternoon, so they are here in the first reading.
| Weakness | Measure | Status |
|---|---|---|
| Version control | 2 of 23 project roots under git | Real and unresolved. Twenty-one roots have no history, no branching and no rollback beyond file copies. The first thing to fix under an acquisition |
| Cost telemetry under-counted | 2.6× low | Reconciled against the Google Cloud invoice on 3 August 2026. The four documented under-count vectors in costService.ts were real. That module is still wrong and will keep under-reporting inside the app until it is fixed |
| A safety control that enforces nothing | 0 of 652 skipped | Found by an independent audit on 4 August 2026 and not yet fixed. The agent bench is written by one module as a bare array and read by the enforcing module as an object, so the enforcer always sees an empty list. The dashboard reports an agent as benched while the six threads it should have skipped executed 652 times. The same break means a standing proposal to bench the content generator, on a measured 48% publish-success rate, can never take effect either way |
| The revenue-truth check has never run | 0 executions | 311 lines written specifically to stop agents citing analytics revenue the storefront never received. Its verified wrapper has zero callers — the one production consumer calls the unverified version. Mitigating: order classification is applied at the storefront layer, so the published store metrics are honest at $0.00. The reconciliation layer the agents read from is the unpatched one |
| Image spend outside the meter | 59 generations | One call site bypasses the budget gate and the spend recorder entirely: 59 image generations between 5 July and 3 August that no counter, log line or dashboard has ever seen. At the project’s own modelled unit price that is about $7.91 — roughly 14% of everything the system believes it has spent. The daily budget gate can be walked straight past |
| The learning loop steers on a fossil | 69.5% of the counter | The violation counter that shapes every writer prompt is dominated by a category that was, until a bug fix on 30 July, the default bucket for anything unmatched. It is ranked first and labelled “recent” while being cumulative since inception, and it out-ranks the two categories that actually block. Separately, the strategy analyst that consumes this has effectively run once per restart since 15 June because its change-detection compares against a list that has been pinned at its retention cap |
| Per-project cost attribution | not established | Open, and the reason there is no per-day cost figure on this page. The billing account holds six projects covering every DDS product; 87.7% of spend is a video-generation project unrelated to the publishing loop. No project maps cleanly to Sovereign Orchestrator Pro, so its isolated running cost cannot be stated from the invoice. Fixing this means labelling resources per system and re-reading a full month. Until then the number is absent rather than estimated |
| Duplication on disk | 2,360,839 duplicated lines | Found by our own audit. Inflates any naive line count, which is why the headline figure declares exclusions |
| Key-person risk | 1 human | Structural. The system runs unattended for 67 days, but the architecture, the constraints and the acceptance criteria come from one person. The portable agent skill exists partly to test whether that transfers — it scored 10/10 with a cold agent, which is evidence, not a solution |
| Revenue not published | — | Deliberate. This page argues about the machine and its operating cost, both measurable. Sales figures are available under NDA, not on a public page |
| No published revenue on either app | 2 listed, units undisclosed | Both apps are listed and buyable — verified ACTIVE against the Admin API on 3 August 2026 — but unit sales are not published on this page. That is deliberate, not an omission: revenue is available under NDA. Listed is not the same as selling, and you should ask |
Earlier Design Delight Studio material carried figures — valuations, ROI multiples, replacement-cost estimates — that were never measured. Those pages are left up as a record, and the correction is documented in full on the portfolio page.
Nothing on this page descends from them. Every figure here is rebuilt from a counted source with a date attached, and there is no valuation, no multiplier and no replacement-cost claim anywhere on it — including in the software's own documentation, which was corrected on 3 August 2026.
07 · Questions
08 · The short version
09 · Contact
Sales figures, the full system inventory and a live walkthrough of the pipeline are available under NDA. Everything on this page is already verifiable without one.