Four years owning 0-to-1 AI products. Ten years leading high-stakes commercial work.
I own two products end to end. Stem is a local-first AI agent with real users, an eval ledger, and a shipped roadmap, and I write the Playbook from its logs. The bid-review desk surfaced $221,915.70 in missed margin and estimating errors across $3,442,234.02 of live commercial bids, every finding linked to source evidence and held for human approval before it reached a customer. Before either, ten years leading high-stakes commercial developments where a missed milestone carried liquidated damages, which is where the habit of grading work before it ships comes from.
A local-first AI orchestration harness I designed and run daily: a chat interface, an eval ledger that grades every output, a retrieval vault, a voice loop, and a constitution that lets it improve itself. The requirements, the evals, and the iteration loop are mine.
Mostly AUv3 audio instruments, built with AI assistance and mined into a seven-phase build pipeline I now teach from. Perihelion is in App Store review right now.
Each one links to the artifact that proves it, or tells you it's on the way.
01 / Context and system design
Give the model the right context
I write the context a model needs before I ask it for anything, the same discipline that turns a vague RFI into a decision someone can actually act on.
Every output Stem produces gets graded against a ledger before it ships, the same instinct that made me itemize $221,915.70 in the review desk before it ever reached a client.
I know when to fan a task out across cheap agents and when to keep it on one thread, because the wrong call burns tokens and still hands back the wrong answer.
Ten years leading delivery on commercial developments where a missed milestone carried liquidated damages taught me where the money leaks on a job and who has to be in the room when it does. That judgment is the whole reason the review desk works.