Skip to content
AUTONOMOUS 67 DAY STREAK · 0 MISSED CEILING $5/day · $10 ON A SALE OLLAMA: LOCAL 12 AGENTS TRUTH GATE: ARMED

Design Delight Studio · Boston · Investor & partner brief

We built a company that runs itself.Then we taught it to argue with itself.

Design Delight Studio is a sustainable streetwear brand. It is also an autonomous production company — a twelve-agent system that finds a market signal, forms a strategy, drafts the work, holds an internal argument about it, blocks any claim it cannot source, and publishes across four platforms. It has done that for 67 consecutive days without a human touching it. The storefront is not the business. It is the proof the machine runs.

Every number here is graded and clickable — open one for its receipt

Quick answer

Design Delight Studio is a Boston sustainable-streetwear brand and AI engineering studio founded and operated by one person, Robert McCullock. Its storefront, publishing and product pipelines run on Sovereign Orchestrator Pro — a twelve-agent system with eight psychometric staff personas that research, draft, negotiate with each other, and gate their own output before it reaches a customer. The system has published for 67 consecutive days unattended, stopped 73% of its own blog publish slots before anything shipped, and runs at a cost of about a $5 per-day ceiling — $10 after a sale — that has never been hit. Every figure on this page carries an evidence grade and a receipt.

MEASURED means counted. CALCULATED means derived from something counted. ASSERTED means stated without proof. As of 3 August 2026 every figure on this page is MEASURED — nothing here is calculated or asserted. That became true today, the hard way: the last CALCULATED figure — our own $1.19 estimate of daily cost — was replaced by an invoiced one, and then the invoiced one was withdrawn too, because it measured the whole billing account rather than this system. A figure that is real but attributed to the wrong thing is not MEASURED, so it is not on this page. If a calculated or asserted figure ever appears here it will be labelled as one, in colour, and it will say why.

01 · The staff

Eight personalities. One of them exists to tell the others they are wrong.

Most AI systems are one model wearing different hats. Ours is a staff. Each of these carries a personality profile, a core drive, a defense mechanism, memory of what the others have done, and a documented relationship with every one of them. They are not prompts with names on. They disagree, on the record, before anything ships.

ATLAS
Founder & Head of Product · drive: significance

“Don't tell me what the user wants. Tell me what they fear. That's where the margin is.”

Sets strategy from raw signal. Search-grounded. Everything downstream inherits its directive.

VERITAS
Quality Assurance & Ethics Auditor · drive: certainty

“This claim in section 3 is unverifiable. Cut it. Precision is persuasion.”

The truth gate. Every factual claim must resolve to a verified source or it is flagged for removal — never silently rewritten. It has hard-halted 28 drafts on legal-class findings and is one of four layers that stopped 491 publish slots.

SCRIBE
Brand Storyteller · drive: mastery

“The rhythm is off in paragraph two. It stumbles. Let's make it sing.”

Writes the long-form draft. Then defends it in the War Room, and loses often.

LEGION
Logistics & Commercial Efficiency · drive: efficiency

“The AOV is too low. We are shipping air. Denied. Bundle it or kill it.”

Vetoes products on unit economics before anything is written about them. A margin gate that runs upstream of the copy, not after it.

ARGUS
Lead Research Scout · drive: accuracy

“Target acquired. JSON-LD extracted. The source is a Collection, not a Product.”

Runs before strategy exists. Surfaces raw external signal so ATLAS is reasoning about the world rather than about itself.

HERMES
Technical SEO & Deployment · drive: optimisation

“Beautiful prose, Scribe, but the H1 tag is missing the primary keyword. Google won't read poetry.”

Grades the finished draft for SEO and schema health as a pre-publish condition.

LUMIERE
Avant-Garde Art Director · drive: aesthetic truth

“'Sustainable' is not a visual. 'Unbleached cotton under harsh noon sun' is a visual.”

Art direction, plus every accessibility alt-text and social image description the system emits.

AMITY
Community Manager · drive: connection

“I hear what you're saying about the price point. Let me share exactly how we source this.”

The only agent optimised for being liked rather than being right. Deliberately.

The War Room — this part is real code, not a metaphor

Before a draft ships, three agents run a tri-agent negotiation. ATLAS argues the strategy, SCRIBE defends the writing, VERITAS attacks anything unsupported. The function returns a negotiation_log, a disagreement_flag, and a revised_draft.

They also carry episodic memory of one another, and their sampling temperature is modulated by live stress and confidence state — a system under pressure literally becomes more conservative.

Why it matters commercially: the expensive failure mode of generative AI is confident, fluent, unsupported output. An adversarial step with a dedicated sceptic is the cheapest known defence, and it runs on every asset.

02 · The machine

Signal in, published asset out, and four places it can be stopped.

This is the loop that runs every day. The interesting parts are not the generation steps — everyone has those. The interesting parts are the three gates, drawn in red, that exist purely to prevent output.

Pipeline — live simulation seven stages · three gates · ~28% of work is killed before it reaches a customer 0in flight 0shipped 0blocked

This animation is a simulation of the pipeline’s shape, not a feed of production telemetry. The counters above tick because the simulation is running in your browser right now. The real figures are the 491 refusals below, and they have receipts.

  1. 01

    Scout

    Raw external signal is surfaced before any strategy exists, so the system reasons about the world rather than about itself.

    ARGUS
  2. 02

    Strategise

    Signal becomes a campaign directive, search-grounded. Everything downstream inherits it.

    ATLAS
  3. GATE 01

    Margin veto

    Products failing unit economics are killed here — before a word is written about them, not after.

    LEGION
  4. 03

    Draft

    Long-form copy is written against the directive and the brand memory.

    SCRIBE
  5. GATE 02

    The argument

    Tri-agent negotiation. Returns a disagreement flag and a revised draft, or escalates.

    ATLAS · SCRIBE · VERITAS
  6. GATE 03

    Truth gate

    Any claim that does not resolve to a verified source is flagged for removal. Then cadence, novelty and quality gates run cheapest-first, fail-fast.

    VERITAS
  7. 04

    Grade & publish

    SEO and schema health graded, accessibility text written, asset published to the scheduled slot.

    HERMES · LUMIERE
The design decision that matters most

If any gate blocks, the system skips the slot. It does not lower the bar and publish something weaker. The comment in the source reads: “An empty day beats a thin/duplicate farm post.”

That is the whole philosophy in one line. A content system that cannot say no is a liability generator, and most of them cannot.

03 · The gate

It refused to publish 491 times.

Every founder claims AI leverage. Almost none can show the machinery that catches their own system lying, with a number on it. This is ours, and here is the breakdown by category rather than a single flattering total.

brand-inquisitor — truth gate● LIVE
The rule the gate enforces

“Every factual claim — certification, statistic, review, result — must resolve to a verified source in the product context or brand memory. If a claim is unverified, uncited or fabricated, you must flag it for removal, not silently rewrite it.”

Flag rather than rewrite is the deliberate part. A model that quietly fixes its own bad claim teaches you nothing and hides the failure rate. This one leaves a record, which is why there is a number to publish at all.

The blacklist is hard-coded and dated. Fair Trade claims were banned on 3 July 2026 after we checked our supply chain and found we had never held the certification. Worldwide-shipping claims were banned on 22 July 2026. Each ruling stays in the source with its date, so a rule can be audited back to the day it was made.

Applied to ourselves

561 live URLs swept

A four-class classifier with reversible repair, run across our own live site to find claims that could not be supported — including ours. Reversible because an automated fix that cannot be undone is a second failure waiting to happen.

Portable

A 5,529-line agent skill

The Shopify build discipline extracted into a portable skill and blind-tested 10 out of 10 by a cold agent with no prior context. Evidence the method transfers rather than living in one person's head.

Verified

71 of 71 tests passing

Full suite, zero failures at time of measurement. Stated with its date because a test result without one is decoration.

04 · Unit economics

What it costs, and the number we took back.

This is the second question every investor asks. The answer here is narrower than the one this page carried four hours ago, because we grouped the invoice by project and found we had been attributing a whole portfolio’s spend to a single system.

Corrected twice in one day — 3 August 2026

First correction, 19:40. This page had published $1.19 a day, graded CALCULATED, from our own token counting, carrying a standing admission that it had never been reconciled against a bill and that costService.ts had four documented under-count vectors. We opened the invoice. The vectors were real: billed spend across 1 May to 31 July 2026 was $285.59 over 92 days. We replaced $1.19 with $3.10 a day.

Second correction, 20:35 — and this one was ours, not the software’s. $3.10 a day is the whole billing account divided by 92. It is not what it costs to run the publishing system, because that account carries every Google Cloud project DDS owns. Grouped by project, $250.52 of the $285.59 — 87.7% — is Veo Weaver, a video-generation project that has nothing to do with the agent loop described on this page. Presenting it under this heading was over-attribution, and it was ours.

So the figure is gone rather than adjusted. What is measured is on the cards above: the whole estate, the share that is video, and what everything else costs. The isolated cost of the publishing system is not stated on this page, because it is not yet attributable. The six projects on the billing account do not map cleanly onto it, and the operator’s own estimate of roughly $1.20 a day is exactly the kind of number this page refuses to print without a receipt — it is also known to be bimodal, since newsletter runs cost materially more than ordinary days.

Two corrections, four hours apart, both published in place with times on them. The first was our telemetry under-reporting. The second was us over-claiming a real number by putting it under the wrong heading. The second is the more instructive one, because nothing was fabricated — an accurate figure was simply asked to mean more than it did. That is the failure mode this page exists to demonstrate catching.

gcloud billing — by project, 2026-05-01 → 2026-07-31● INVOICED

What is not in doubt: inference runs on local hardware — a single consumer RTX 3060 running Ollama for bulk work, with Gemini reserved for the calls that need it. That routing decision is why serverless compute across those three months came to one cent — the one part of the bill that unambiguously belongs to this system — and why the cost ceiling is a design parameter rather than a hope.

05 · The proof instance

Pointed at sustainable streetwear. It did not have to be.

The agents, the gates, the scheduler and the publishing adapters are general. Only the brand memory, the compliance blacklist and the product context are apparel-specific, and those are configuration rather than architecture. The storefront is what a live instance of the machine looks like after seventeen months.

Storefront

683 products, engineered solo

Built and maintained by one person directing agents, on Shopify Basic. Product drafts, mockups and SEO copy are generated by the pipeline and synced through the Shopify Admin API and Printful, rate-limit aware.

Taught openly

81 free Academy classes

The method is published, not hoarded. No signup, no email capture, no paywall, no certificate. A company confident in its execution can afford to give away its method, and the Academy is the argument that we are.

Commercial software

Two desktop apps, $128.99 of list price

PromptDJ NS-9000, $79 — a real-time neural music synthesiser driven by MIDI hardware. VibeTube AI, $49.99 — a native Windows YouTube strategy engine on a bring-your-own-key architecture, so it carries no API cost for us. Both are sold as licensed desktop applications and both were confirmed ACTIVE against the live Shopify Admin API on 3 August 2026. An earlier version of this page claimed only one, because VibeTube’s listing could not be verified at the time.

Free, and staying free

Three playable games

Ashaveth, an AI-driven dark-fantasy RPG with generative NPCs, grew out of Legend of the Red Dragon 2026 — a playable beta never launched commercially for copyright reasons. Dragonfly Crush was built as a birthday gift and shipped anyway. Three further builds exist and are deliberately not counted, because they are demos and counting them would be padding.

Certifications — stated precisely, because this is where brands overreach

Four certifications: GOTS, GRS, OCS and PETA-Approved Vegan, held product by product rather than brand-wide. Plus one membership: a Fair Wear Foundation supplier membership through Stanley/Stella — a membership, not a certification, and never described as one.

Design Delight Studio does not hold Fair Trade and has never held it. It is not OEKO-TEX certified at brand level. An earlier version of this page said “five certifications” in a headline sitting directly above a paragraph that correctly said four plus a membership. That headline was wrong and has been corrected.

06 · Diligence

What is wrong with it.

A page that lists only strengths is not evidence of anything. These are the things a technical buyer would find in the first afternoon, so they are here in the first reading.

WeaknessMeasureStatus
Version control2 of 23 project roots under gitReal and unresolved. Twenty-one roots have no history, no branching and no rollback beyond file copies. The first thing to fix under an acquisition
Cost telemetry under-counted2.6× lowReconciled against the Google Cloud invoice on 3 August 2026. The four documented under-count vectors in costService.ts were real. That module is still wrong and will keep under-reporting inside the app until it is fixed
A safety control that enforces nothing0 of 652 skippedFound by an independent audit on 4 August 2026 and not yet fixed. The agent bench is written by one module as a bare array and read by the enforcing module as an object, so the enforcer always sees an empty list. The dashboard reports an agent as benched while the six threads it should have skipped executed 652 times. The same break means a standing proposal to bench the content generator, on a measured 48% publish-success rate, can never take effect either way
The revenue-truth check has never run0 executions311 lines written specifically to stop agents citing analytics revenue the storefront never received. Its verified wrapper has zero callers — the one production consumer calls the unverified version. Mitigating: order classification is applied at the storefront layer, so the published store metrics are honest at $0.00. The reconciliation layer the agents read from is the unpatched one
Image spend outside the meter59 generationsOne call site bypasses the budget gate and the spend recorder entirely: 59 image generations between 5 July and 3 August that no counter, log line or dashboard has ever seen. At the project’s own modelled unit price that is about $7.91 — roughly 14% of everything the system believes it has spent. The daily budget gate can be walked straight past
The learning loop steers on a fossil69.5% of the counterThe violation counter that shapes every writer prompt is dominated by a category that was, until a bug fix on 30 July, the default bucket for anything unmatched. It is ranked first and labelled “recent” while being cumulative since inception, and it out-ranks the two categories that actually block. Separately, the strategy analyst that consumes this has effectively run once per restart since 15 June because its change-detection compares against a list that has been pinned at its retention cap
Per-project cost attributionnot establishedOpen, and the reason there is no per-day cost figure on this page. The billing account holds six projects covering every DDS product; 87.7% of spend is a video-generation project unrelated to the publishing loop. No project maps cleanly to Sovereign Orchestrator Pro, so its isolated running cost cannot be stated from the invoice. Fixing this means labelling resources per system and re-reading a full month. Until then the number is absent rather than estimated
Duplication on disk2,360,839 duplicated linesFound by our own audit. Inflates any naive line count, which is why the headline figure declares exclusions
Key-person risk1 humanStructural. The system runs unattended for 67 days, but the architecture, the constraints and the acceptance criteria come from one person. The portable agent skill exists partly to test whether that transfers — it scored 10/10 with a cold agent, which is evidence, not a solution
Revenue not publishedDeliberate. This page argues about the machine and its operating cost, both measurable. Sales figures are available under NDA, not on a public page
No published revenue on either app2 listed, units undisclosedBoth apps are listed and buyable — verified ACTIVE against the Admin API on 3 August 2026 — but unit sales are not published on this page. That is deliberate, not an omission: revenue is available under NDA. Listed is not the same as selling, and you should ask
On numbers we have published before

Earlier Design Delight Studio material carried figures — valuations, ROI multiples, replacement-cost estimates — that were never measured. Those pages are left up as a record, and the correction is documented in full on the portfolio page.

Nothing on this page descends from them. Every figure here is rebuilt from a counted source with a date attached, and there is no valuation, no multiplier and no replacement-cost claim anywhere on it — including in the software's own documentation, which was corrected on 3 August 2026.

07 · Questions

FAQ

What is Design Delight Studio actually selling?
Sustainable streetwear, through a 683-product Shopify storefront. But the storefront is also the proof: it is a live, revenue-generating instance of an autonomous production system that one person builds and operates. The apparel is the current target of the machine, not the ceiling of it.
How many people work at Design Delight Studio?
One human. Robert McCullock writes the constraints, the architecture and the acceptance criteria. Twelve registered agent personas do the research, drafting, grading and publishing, and eight of them carry full psychometric profiles with a relationship matrix so they argue with each other before anything ships.
What stops the AI from publishing something false?
A truth gate called the Brand Inquisitor, and it is four layers rather than one: a deterministic regular-expression sanitiser that rewrites banned phrasing before any model sees the text, an LLM audit, a classifier, and a pre-publish gate stack. Measured across 67 days of logs to 4 August 2026: 491 blog publish slots were skipped against 184 approved, so 73% of publication opportunities were stopped; 296 drafts were auto-corrected deterministically; 66 were blocked for insufficient substance; and 28 were hard-halted on legal-class findings. The page previously reported 3,108 “claims blocked” from the Inquisitor’s cumulative counter. That figure was withdrawn on 4 August after an independent code audit found 69.5% of it was an adjective-bloat bucket the auditor is explicitly instructed never to raise as a violation, so it never blocked anything. Worth stating plainly: every LLM-dependent layer of this gate fails open, meaning a model outage degrades to passing rather than blocking. The deterministic layer does not.
What does it cost to run?
The honest answer is that we do not publish a single per-day figure, because we cannot yet attribute one. What is measured: Google Cloud billed $285.59 across all six DDS projects in the 92 days from 1 May to 31 July 2026, tax $0. Grouped by project, $250.52 of that - 87.7% - is Veo Weaver, a video-generation project unrelated to the publishing loop. Excluding it, everything else runs $0.38 a day, and that still bundles the try-on app and Thrift Finder in with the publishing stack. A $5 per-day telemetry-enforced ceiling exists, rising to $10 once the day’s first sale lands, and has never been hit. Serverless compute across all three months came to one cent, because inference runs locally on a single consumer GPU. This page corrected this figure twice on 3 August 2026 - first replacing a calculated estimate with the invoice, then withdrawing the per-day number entirely once the invoice was grouped by project and turned out to be dominated by video.
How long has it run without a human?
67 consecutive days with zero missed publications, verified by enumerating every daily log file against the calendar. 61 is a floor rather than a ceiling because log retention is 61 days, so the true streak may be longer and cannot be proven from logs alone.
What sustainability certifications does Design Delight Studio hold?
Four: GOTS, GRS, OCS and PETA-Approved Vegan, held product by product rather than brand-wide. Separately there is a Fair Wear Foundation supplier membership through Stanley/Stella, which is a membership and not a certification. Design Delight Studio does not hold Fair Trade and is not OEKO-TEX certified at brand level.
What is the weakest part of the operation?
Version control. Only 2 of 23 project roots are under git. That is published here deliberately because it is the first thing a technical buyer should ask about, and because a page that lists only strengths is not evidence of anything.
Is any of this actually generating revenue?
The storefront has been live and selling throughout, and PromptDJ NS-9000 is a paid desktop application at $79 verified against the Shopify Admin API. Detailed sales figures are not published. The claim this page makes is about the machine and its operating cost, both of which are measurable, not about revenue.
Could this system be pointed at something other than apparel?
Yes, and that is the argument. The agents, the gates, the scheduler and the publishing adapters are general. Only the brand memory, the compliance blacklist and the product context are apparel-specific, and those are configuration rather than architecture.
How do I discuss investment or partnership?
Email Robert@ddsboston.com. Every figure on this page is clickable and opens its own receipt showing the method, the date and the caveats, so diligence can start before the first call.

08 · The short version

Key takeaways

If you read nothing else
  • A twelve-agent production system with eight psychometric staff personas publishes across four platforms unattended. 67 consecutive days, zero missed.
  • Three agents run a documented argument before anything ships, and a dedicated sceptic stopped 491 unsupported claims — 359 of them fabricated statistics.
  • A $5/day enforced ceiling, rising to $10 after the day’s first sale, that has never been hit, and one cent of serverless compute across three months, because inference runs on one consumer GPU. We do not publish a per-day cost for the publishing system: the billing account covers six products and 87.7% of it is video generation, so the figure is absent rather than estimated.
  • The 683-product storefront is the proof instance, not the business. The architecture is general; only the brand config is apparel-specific.
  • The method is published in 81 free classes and extracted into a portable agent skill that scored 10/10 blind-tested by a cold agent.
  • Every figure here carries a grade, a method and a date — and the weaknesses are printed next to the strengths, including 2-of-23 version control and an unreconciled cost figure.

09 · Contact

Design Delight Studio

Investor & partner inquiries

Sales figures, the full system inventory and a live walkthrough of the pipeline are available under NDA. Everything on this page is already verifiable without one.