Selected work + evidence search

31 systems with something you can inspect

The strongest governance cases lead. A searchable evidence index follows, covering all 31 systems, their controls, dates, demos, and 11 live products.

Start with the selected cases when time is short. Use the evidence search when you need a project, year, control, model, record, or working artifact.

The systems

  1. Stillwell — a governed voice line for older adults

    Live in production

    A family answering line an older adult can call, built so it cannot pretend to be a person, cannot be talked past by a caller who keeps trying, and cannot quietly keep a recording of the conversation. Consent is checked at the moment of the call rather than at signup, and revoking it lands on the next call, not the next release.

    What I did: Founded it, wrote its governing standard, and then audited the code against that standard rather than assuming they agreed. They did not: the audit found the adopted policy sitting on an unmerged branch while production ran the superseded version, and the spoken lines the system actually used still implied a live human was on the phone. Both are written up as open findings rather than quietly fixed and forgotten.

    • 3,831 tests / 253 files
    • 84-attack red team, 0 escapes
    • RUAIH self-assessment
    • consent checked per call
  2. Calm Couch — a practice platform for couples counsellors

    Live in production

    Clients practise between sessions and their counsellor sees what was worked on, without the platform ever holding a free-text clinical note or a conversation transcript. What a therapist sees is a structured record of activity, by design, because the alternative is a therapy transcript in a database. Isolation between practices is proved by 413 assertions whose only job is to cross a boundary and fail.

    What I did: Founded it and set the constraint before the first feature, then built the programme that proves it: a HIPAA control matrix, a risk register carrying an owner, a treatment and a named approver for every risk, a vendor register, an access matrix and a data inventory — with machine-readable sources so the evidence can be queried rather than read. The readiness decision is made by a script rather than by me, and it stays closed until signed agreements, the risk analysis, retention proof and independent review are all in.

    • HIPAA control matrix
    • risk register with named approvers
    • vendor and access registers
    • 413 isolation assertions
  3. Plate Check — where the medical-device line actually is

    Live in production

    It will not tell you what dose to take, and it cannot be talked into it. The model's own words are screened before they reach the screen, a dose-shaped sentence is removed surgically rather than the whole answer being dumped, and the data field a dose would live in does not exist.

    What I did: Read the statute first and drew the line, then deleted four features that crossed it — including the one users ask for most — because shipping them would have moved the product from not-a-device to regulated software overnight. Then scored the vision model against a labelled dataset, missed the accuracy target, and left the feature switched off. The live product is not linked here, by a standing privacy decision.

    • device-boundary analysis
    • HIPAA scoping memo
    • dose-language screening
    • labelled-dataset eval
  4. agent-gate — safety fence

    Open source

    Stops a coding agent from leaking secrets, deleting data, or claiming success without proof — and lets the safe actions through. Stdlib-only Python, no dependencies, 79 tests.

    What I did: Wrote it as a standalone reusable layer, then audited it against my own résumé and found I had overstated it: the evidence path did refuse an unproven claim, but the action classifier let an unmatched command through. Writing that down and closing it took the suite from 54 tests to 79.

    • Python
    • deterministic rules
    • fail-closed
    • 79 tests
  5. Multi-Model QA Cascade

    Internal tooling

    No model grades its own homework: three providers propose answers in parallel, a scoring layer containing no model ranks them, and a person approves before anything is written to the database. Runs at zero marginal cost locally, which is the reason it actually gets used.

    What I did: Designed the routing and the scoring layer. The cost work is the governance work: controls that are expensive get quietly dropped in month nine.

    • Ollama
    • Groq
    • Codex
    • deterministic scoring
    • human-in-the-loop
  6. TradeTEST.TRAINING

    Live in production

    Paying customers study for a state licence exam in English or Spanish, and no translated question reaches them until five checks agree it kept its meaning. The pipeline is resumable and cost-capped, so a bad run is cheap to abandon.

    What I did: Built and shipped end to end — billing, scheduling, content pipeline, and the QA gate that decides whether a translation batch is allowed through. Its study coach runs on a written intent list with refusals for everything outside it.

    • Groq
    • Next.js
    • Supabase
    • TypeScript
    • Stripe
    • SM-2
  7. Find Your Vote

    Live in production

    Every voter sees exactly why a candidate ranked where they did — and the same address always gets the same answer, audited run after run, across 25+ public-record sources.

    What I did: Built the scoring engine and the determinism audit. For anything touching an election, a ranking nobody can explain is not a feature.

    • Next.js
    • TypeScript
    • determinism audit
    • 25+ record sources
  8. Vision Factory — dual-model consensus

    Internal tooling

    Stops an inconsistent asset run before it reaches production: one model generates, a second verifies placement, and the run halts when disagreement crosses a fixed threshold. A single model generating a large run drifts confidently and quietly — nobody notices until the whole set is wrong.

    What I did: Built the consensus gate and the status board that makes a long run inspectable while it is still running.

    • GPT-4o
    • Gemini vision
    • consensus gate
    • spec-driven pipeline
  9. Sea Star — publish automation

    Live in production

    Non-technical staff publish to four channels from one post, nightly, with no engineer in the loop. It runs as an admin tool in production, which means it fails in front of someone who will tell me.

    What I did: Designed the fan-out, wrote the publish integrations, and put it in the hands of non-technical staff who use it daily.

    • Groq
    • Next.js
    • publish API
    • nightly cron

Ask the work

Find the evidence, not another archive

Search every system by plain-language terms, dates, technologies, controls, and outcomes. Results always lead to an inspectable demo, live product, or case record where one exists.

Try:
31 recordsLocal deterministic search · no query leaves this browser
  1. agent-gate — safety fence

    Governance & safety tooling

    Stops a coding agent from leaking secrets, deleting data, or claiming success without proof — 79 tests, zero dependencies

    A portable safety fence for AI coding agents. It blocks the dangerous actions — leaking secrets, deleting things, claiming done without proof, rewriting history — and lets the safe ones through. Off by default and on purpose.

    Python · deterministic rules · fail-closed · 79 tests

    Version-control record: 2026-05-29 → 2026-08-20 · 5 commits

  2. AI observability control room

    Governance & safety tooling

    A weak model or agent run leaves evidence and can stop the release instead of disappearing inside an average

    A runnable model canary with structured performance telemetry, automated output checks, evaluator calibration, agent tool and loop controls, and one exportable evidence chain.

    Next.js · structured telemetry · evaluator calibration · agent traces · SHA-256 evidence chain

  3. Asset-Ops pipeline dashboard

    AI pipelines

    You can see a long generation run going wrong while it is still cheap to stop

    A status board that makes a long generation run auditable while it is still running, instead of after it has finished going wrong.

    spec-driven pipeline · status board

  4. audit-kit — review workbench

    Governance & safety tooling

    A reviewer's decision becomes data the pipeline obeys — not a comment somebody may read later

    Human-in-the-loop review as a first-class component: the reviewer's decisions are structured output the pipeline consumes, not a comment thread someone has to interpret later.

    typed JSON · review gates · human-in-the-loop

  5. Bilingual exam localisation gate

    Evals & measurement

    3,400 exam questions crossed a language without changing meaning — and any batch that cannot prove it stays behind

    Translation at volume where the interesting problem is not the translation but knowing which batches are safe to ship. Every stage is a gate with a written pass condition.

    multi-model · staged QA gates · resumable runs

  6. Calm Couch — couples practice platform

    Live products

    A couple and their therapist see the same week, and the couple decides what the therapist sees

    A shared practice workspace for couples counselling. The demo runs both lenses: switch between the couple view and the therapist view and watch a sharing scope turned off disappear from the therapist's session brief on the next render. A couple who has not finished consent shows a consent boundary instead of a brief. Grew out of two precursor builds, The Palms and gotta-guy.

    Next.js · consent scopes · role-scoped views · synthetic data

  7. Coach Vale — explain-back tutor

    AI pipelines

    The learner has to explain their answer before moving on — which is where the actual learning happens

    A Socratic tutor for the CSLB contractor licence exam: the learner picks an answer and then has to explain why, which is where the actual teaching happens.

    Groq · Socratic prompting · exam prep

  8. coolcook.ing — cost engine

    Live products

    A food vendor sees which menu item is losing money while there is still time to change it

    Food-vendor cost and margin calculator: edit menu items inline and watch per-SKU margin move. Unit economics a vendor can actually operate, not a spreadsheet they abandon.

    Next.js · unit economics · margin modelling

  9. Crown Ridge — field service portal

    Live products

    The owner, the manager and the tech each see the one screen they need — and a slipping job surfaces before the customer calls

    Role-aware field-service operations: portfolio KPIs and threshold alerts for the owner, a one-click assignment queue for the manager, today's jobs with photo-evidence capture for the tech.

    Next.js · field ops · role-based UI · dashboards

  10. Datum & Plane — admin toolset

    Live products

    One login gives the owner the briefing, the estimate and the quote — instead of four tools and a lost afternoon

    Gated owner toolset: AI business briefing, walk-site estimator, quote builder, and admin controls in one place.

    Next.js · Groq · estimator · quote builder

  11. Delta King — book-writing dashboard

    Dashboards, calculators & ops

    A book gets written in one workspace instead of three tabs — progress visible, drafts assisted, reading alongside

    A themed chapter reader alongside manuscript tracking and an LLM drafting assistant, in one workspace rather than three tabs.

    Next.js · manuscript tracking · LLM assist

  12. FEST — facility emergency safety & tracking

    Dashboards, calculators & ops

    After an incident, you can replay exactly where everything was and when — instead of reconstructing it from memory

    Real-time asset tracking on an SVG floor plan, with breach alerts and a scrubber to replay what happened and when. Built for the moment after the incident, when someone asks where things actually were.

    SVG mapping · geofencing · realtime · history replay

  13. Find Your Vote

    Live products

    Every voter sees exactly why a candidate ranked where they did — and the same address always gets the same answer

    Address in, ranked candidate matches out, drawn from 25+ public-record sources. Same inputs produce the same ranking every run — for anything touching an election, an unexplainable ranking is not a feature.

    Next.js · TypeScript · determinism audit · 25+ record sources

    Version-control record: 2023-07-04 → 2026-08-20 · 1,391 commits

  14. Game Generator

    AI pipelines

    Flashcards become a playable game only when the material actually suits one — it scores the fit instead of forcing all three modes

    Turns flashcards into playable Match, Order or MCQ mini-games, and scores which mode actually suits the material rather than generating all three and hoping.

    deterministic generation · suitability scoring · HITL playtest

  15. Gotta-Guy — realtime exercises

    Live products

    Both partners see the same exercise at the same moment — and the guidance comes from published sources, not the model's imagination

    Guided communication exercises with both screens synced live. The content is source-grounded rather than model-improvised, which is the whole point in this domain.

    Supabase Realtime · Next.js · TypeScript

  16. Image-gen workflow tool

    AI pipelines

    Any image it made can be made again — the settings are recorded, so a result is reproducible instead of lucky

    Image generation as a repeatable pipeline with recorded settings, so a result can be reproduced rather than re-improvised.

    image generation · workflow capture

  17. Intake / fact-finding tool

    Dashboards, calculators & ops

    Change the questions without rebuilding the form — and every submission arrives structured, ready to use

    A JS object defines the sections, field types and conditional logic; the form and its typed output follow from it. Change the schema, not the form.

    schema-driven · conditional fields · typed JSON export

  18. Money reconciliation dashboard

    Dashboards, calculators & ops

    Mismatched money surfaces where someone will actually see it — exceptions raised, not buried in a spreadsheet

    Transactions matched and exceptions raised where someone will see them, which is the only place a reconciliation tool earns its keep.

    TypeScript · reconciliation · exception surfacing

  19. Multi-Model QA Cascade

    AI pipelines

    No model grades its own homework — three propose, plain rules rank them, and a person makes the final call

    One item routed to Ollama, Groq and Codex at once. A deterministic scoring layer picks the winner, then a person approves before anything is written to SQL. The cost work is the governance work — controls that are expensive get quietly dropped in month nine.

    Ollama · Groq · Codex · deterministic scoring · human-in-the-loop

  20. OSTUP — scaffold & model router

    Governance & safety tooling

    A new project deploys with its guardrails already installed — five minutes in, the checks are running

    One command scaffolds a repo, a Vercel deploy and an agent-ready kit, with a router that sends each job to the cheapest model tier that succeeds. Governance that is expensive to run stops getting run.

    Next.js · scaffolding · agent kit · model-tier routing

    Version-control record: 2026-05-21 → 2026-08-20 · 152 commits

  21. Plate Check — vision estimate accuracy

    Evals & measurement

    It refuses to give medical advice, and the refusal is enforced in code rather than promised in a policy

    A health dashboard with a vision model on top, and the project where I had to decide where a display tool stops and a regulated medical device starts. Under the Cures Act, software that only stores and displays device data is statutorily not a device; the moment it recommends a dose it is Clinical Decision Support. So four features were specified and then permanently killed, and a screen now runs over every word the model writes before anyone reads it. Separately, a model call per dish returns a carb and calorie range scored against a 5,800-image labelled dataset in real time and reported as mean absolute percentage error (MAPE) — it missed its own accuracy target, so the feature is switched off. The live app is not linked, by a standing privacy decision.

    FDA device boundary · HIPAA scoping · dose-language screening · Groq · dataset-backed eval

    Version-control record: 2026-05-21 → 2026-08-10 · 104 commits

  22. Project portfolio audit dashboard

    Evals & measurement

    Caught this portfolio's own demos overstating themselves — the audit tool turned on its owner

    The tool I used to audit my own work — which is how several of the defects on this site were found, including demos that claimed to be live while making no model call.

    Next.js · scoring rubric · audit harness

  23. rbl.land — underwriting tools

    Live products

    Screens a housing deal in under a minute and says go or no — not a spreadsheet to interpret

    NorCal affordable-housing underwriting suite: HCV rent-cap screen, RCFE calculator, and deal screens that return a verdict rather than a number to interpret.

    Next.js · TypeScript · real-estate underwriting

  24. run-run — local-AI studio

    Live products

    Packaging copy and product imagery for pennies, on your own machine — nothing uploaded, no per-use bill

    Local-AI packaging studio: brand positioning copy and Stable Diffusion image prompts, with an SSRF-safe image API. Runs on local Ollama, so the marginal cost is zero.

    Ollama · Stable Diffusion · Next.js · cost engine

  25. Sea Star — publish automation

    Live products

    One post becomes four channels overnight, and non-technical staff run it daily

    One blog topic fans out to a Facebook caption, an Instagram caption with hashtags, and an email subject and body from a single model call. Runs as an admin tool used daily by non-technical staff.

    Groq · Next.js · publish API · nightly cron

    Version-control record: 2026-03-05 → 2026-08-10 · 187 commits

  26. Stillwell

    Live products

    Shipped, live, and serving users on its own domain

    Shipped and serving users. The demo runs the governance end to end on fictional families: consent that fails closed with teach-back and revocation, honest identity disclosure on a direct ask, and the deterministic family-alert ladder. The live product is not demonstrated directly, so this is the mechanism rather than described from guesswork.

    Next.js

    Version-control record: 2026-07-18 → 2026-08-20 · 433 commits

  27. Tax + benefits calculator

    Dashboards, calculators & ops

    One set of questions answers the two things people actually need together — the tax bill and the benefits they qualify for

    A wizard that collects filing status, income, household and deductions, then returns an estimate alongside benefits eligibility (including CalFresh) — the two questions people actually need answered together.

    Next.js · multi-step wizard · eligibility rules

  28. The Palms — couples dashboard

    Dashboards, calculators & ops

    A couple sees the same picture of the week — load, plans, check-ins — instead of arguing from two different ones

    The shared-home foundation for Calm Couch’s Our World: weekly focus and load tracking, family moments, structured check-ins, requests, agreements and repair flows. The demo runs on fictional data and ships no images; the live app is not linked, by a standing privacy decision.

    Next.js · TypeScript · Supabase

  29. TradeTEST.TRAINING

    Live products

    Paying customers study for a licence exam in two languages — and no translated question ships until five checks agree it kept its meaning

    Tutored exam-prep product with spaced-repetition scheduling and bilingual EN/ES content. The localisation pipeline is resumable and cost-capped, so a bad run is cheap to abandon.

    Groq · Next.js · Supabase · TypeScript · Stripe · SM-2

    Version-control record: 2026-02-18 → 2026-08-10 · 1,132 commits

  30. Unit-economics financial model

    Evals & measurement

    You can argue with the inputs instead of having to trust the output — every projection shows the assumptions it stands on

    A model whose inputs you can argue with. Numbers you cannot interrogate are numbers nobody should act on.

    TypeScript · scenario modelling

  31. Vision Factory — dual-model consensus

    AI pipelines

    Stops an inconsistent asset run before it reaches production — a second model checks the first and halts the line on disagreement

    One model generates assets, a second verifies them, and the run halts if drift crosses a fixed threshold. It exists because a single model generating a large asset run drifts confidently and quietly, and nobody notices until the whole set is wrong.

    GPT-4o · Gemini vision · consensus gate · spec-driven pipeline

    Version-control record: 2026-04-09 → 2026-08-20 · 169 commits

Receipts

A compact public record of private work

The source stays private. The aggregate timeline does not: it shows when the work began, how much of it exists, and the dated records behind the named systems.

Private repositories
62
Commits
7,010
First AI commit
2023-07-04
Open the named system ledger (10 records)
First commit, last commit, and commit count for each named system
SystemFirst commitLast commitCommits
Find Your Vote2023-07-042026-08-201,391
TradeTEST.TRAINING2026-02-182026-08-101,132
TradeTEST content pipeline2026-03-222026-08-10229
Stillwell2026-07-182026-08-20433
Sea Star2026-03-052026-08-10187
Vision Factory2026-04-092026-08-20169
ostup2026-05-212026-08-20152
Plate Check — vision estimate accuracy2026-05-212026-08-10104
This site2026-06-162026-08-2076
agent-gate2026-05-292026-08-205

Generated from local git history on 2026-08-20. Only aggregate dates and counts are published. No code, file names, or commit messages are published. Counts are self-reported; the 4 July 2023 start date is independently visible on GitHub’s contribution record (opens in a new tab).

Stack

Capabilities

The models, infrastructure, evaluation patterns, and delivery systems used across the evidence index.

Models orchestrated

  • Claude
  • OpenAI / Codex / GPT-4o
  • Groq
  • Gemini vision
  • local Ollama
  • Stable Diffusion

Infrastructure

  • Next.js
  • TypeScript
  • Supabase
  • Playwright
  • Tailwind
  • Vercel

AI & pipeline

  • Deterministic pipelines
  • dataset-backed evals
  • non-LLM scoring
  • multi-model routing
  • fail-closed safety
  • MAPE / eval loops

Systems shipped

  • SaaS with billing
  • publish automation
  • human-in-loop workbench
  • agentic safety fence
  • schema-driven intake
  • unit-economics engine

Legacy visual view

Prefer browsing to searching?

The older visual gallery remains available for scanning thumbnails and capability filters. The evidence search above is the maintained index and the better route when you are looking for a particular term or date.

Next step

Worth a conversation?

The governance portfolio carries the selected cases. The résumé has the compressed career record and opens without an email wall.