Models orchestrated
- Claude
- OpenAI / Codex / GPT-4o
- Groq
- Gemini vision
- local Ollama
- Stable Diffusion
Selected work + evidence search
The strongest governance cases lead. A searchable evidence index follows, covering all 31 systems, their controls, dates, demos, and 11 live products.
Start with the selected cases when time is short. Use the evidence search when you need a project, year, control, model, record, or working artifact.
A family answering line an older adult can call, built so it cannot pretend to be a person, cannot be talked past by a caller who keeps trying, and cannot quietly keep a recording of the conversation. Consent is checked at the moment of the call rather than at signup, and revoking it lands on the next call, not the next release.
What I did: Founded it, wrote its governing standard, and then audited the code against that standard rather than assuming they agreed. They did not: the audit found the adopted policy sitting on an unmerged branch while production ran the superseded version, and the spoken lines the system actually used still implied a live human was on the phone. Both are written up as open findings rather than quietly fixed and forgotten.
Clients practise between sessions and their counsellor sees what was worked on, without the platform ever holding a free-text clinical note or a conversation transcript. What a therapist sees is a structured record of activity, by design, because the alternative is a therapy transcript in a database. Isolation between practices is proved by 413 assertions whose only job is to cross a boundary and fail.
What I did: Founded it and set the constraint before the first feature, then built the programme that proves it: a HIPAA control matrix, a risk register carrying an owner, a treatment and a named approver for every risk, a vendor register, an access matrix and a data inventory — with machine-readable sources so the evidence can be queried rather than read. The readiness decision is made by a script rather than by me, and it stays closed until signed agreements, the risk analysis, retention proof and independent review are all in.
It will not tell you what dose to take, and it cannot be talked into it. The model's own words are screened before they reach the screen, a dose-shaped sentence is removed surgically rather than the whole answer being dumped, and the data field a dose would live in does not exist.
What I did: Read the statute first and drew the line, then deleted four features that crossed it — including the one users ask for most — because shipping them would have moved the product from not-a-device to regulated software overnight. Then scored the vision model against a labelled dataset, missed the accuracy target, and left the feature switched off. The live product is not linked here, by a standing privacy decision.
Stops a coding agent from leaking secrets, deleting data, or claiming success without proof — and lets the safe actions through. Stdlib-only Python, no dependencies, 79 tests.
What I did: Wrote it as a standalone reusable layer, then audited it against my own résumé and found I had overstated it: the evidence path did refuse an unproven claim, but the action classifier let an unmatched command through. Writing that down and closing it took the suite from 54 tests to 79.
No model grades its own homework: three providers propose answers in parallel, a scoring layer containing no model ranks them, and a person approves before anything is written to the database. Runs at zero marginal cost locally, which is the reason it actually gets used.
What I did: Designed the routing and the scoring layer. The cost work is the governance work: controls that are expensive get quietly dropped in month nine.
Paying customers study for a state licence exam in English or Spanish, and no translated question reaches them until five checks agree it kept its meaning. The pipeline is resumable and cost-capped, so a bad run is cheap to abandon.
What I did: Built and shipped end to end — billing, scheduling, content pipeline, and the QA gate that decides whether a translation batch is allowed through. Its study coach runs on a written intent list with refusals for everything outside it.
Every voter sees exactly why a candidate ranked where they did — and the same address always gets the same answer, audited run after run, across 25+ public-record sources.
What I did: Built the scoring engine and the determinism audit. For anything touching an election, a ranking nobody can explain is not a feature.
Stops an inconsistent asset run before it reaches production: one model generates, a second verifies placement, and the run halts when disagreement crosses a fixed threshold. A single model generating a large run drifts confidently and quietly — nobody notices until the whole set is wrong.
What I did: Built the consensus gate and the status board that makes a long run inspectable while it is still running.
Non-technical staff publish to four channels from one post, nightly, with no engineer in the loop. It runs as an admin tool in production, which means it fails in front of someone who will tell me.
What I did: Designed the fan-out, wrote the publish integrations, and put it in the hands of non-technical staff who use it daily.
Ask the work
Search every system by plain-language terms, dates, technologies, controls, and outcomes. Results always lead to an inspectable demo, live product, or case record where one exists.
Stops a coding agent from leaking secrets, deleting data, or claiming success without proof — 79 tests, zero dependencies
A portable safety fence for AI coding agents. It blocks the dangerous actions — leaking secrets, deleting things, claiming done without proof, rewriting history — and lets the safe ones through. Off by default and on purpose.
Python · deterministic rules · fail-closed · 79 tests
Version-control record: 2026-05-29 → 2026-08-20 · 5 commits
A weak model or agent run leaves evidence and can stop the release instead of disappearing inside an average
A runnable model canary with structured performance telemetry, automated output checks, evaluator calibration, agent tool and loop controls, and one exportable evidence chain.
Next.js · structured telemetry · evaluator calibration · agent traces · SHA-256 evidence chain
You can see a long generation run going wrong while it is still cheap to stop
A status board that makes a long generation run auditable while it is still running, instead of after it has finished going wrong.
spec-driven pipeline · status board
A reviewer's decision becomes data the pipeline obeys — not a comment somebody may read later
Human-in-the-loop review as a first-class component: the reviewer's decisions are structured output the pipeline consumes, not a comment thread someone has to interpret later.
typed JSON · review gates · human-in-the-loop
3,400 exam questions crossed a language without changing meaning — and any batch that cannot prove it stays behind
Translation at volume where the interesting problem is not the translation but knowing which batches are safe to ship. Every stage is a gate with a written pass condition.
multi-model · staged QA gates · resumable runs
A couple and their therapist see the same week, and the couple decides what the therapist sees
A shared practice workspace for couples counselling. The demo runs both lenses: switch between the couple view and the therapist view and watch a sharing scope turned off disappear from the therapist's session brief on the next render. A couple who has not finished consent shows a consent boundary instead of a brief. Grew out of two precursor builds, The Palms and gotta-guy.
Next.js · consent scopes · role-scoped views · synthetic data
The learner has to explain their answer before moving on — which is where the actual learning happens
A Socratic tutor for the CSLB contractor licence exam: the learner picks an answer and then has to explain why, which is where the actual teaching happens.
Groq · Socratic prompting · exam prep
A food vendor sees which menu item is losing money while there is still time to change it
Food-vendor cost and margin calculator: edit menu items inline and watch per-SKU margin move. Unit economics a vendor can actually operate, not a spreadsheet they abandon.
Next.js · unit economics · margin modelling
The owner, the manager and the tech each see the one screen they need — and a slipping job surfaces before the customer calls
Role-aware field-service operations: portfolio KPIs and threshold alerts for the owner, a one-click assignment queue for the manager, today's jobs with photo-evidence capture for the tech.
Next.js · field ops · role-based UI · dashboards
One login gives the owner the briefing, the estimate and the quote — instead of four tools and a lost afternoon
Gated owner toolset: AI business briefing, walk-site estimator, quote builder, and admin controls in one place.
Next.js · Groq · estimator · quote builder
A book gets written in one workspace instead of three tabs — progress visible, drafts assisted, reading alongside
A themed chapter reader alongside manuscript tracking and an LLM drafting assistant, in one workspace rather than three tabs.
Next.js · manuscript tracking · LLM assist
After an incident, you can replay exactly where everything was and when — instead of reconstructing it from memory
Real-time asset tracking on an SVG floor plan, with breach alerts and a scrubber to replay what happened and when. Built for the moment after the incident, when someone asks where things actually were.
SVG mapping · geofencing · realtime · history replay
Every voter sees exactly why a candidate ranked where they did — and the same address always gets the same answer
Address in, ranked candidate matches out, drawn from 25+ public-record sources. Same inputs produce the same ranking every run — for anything touching an election, an unexplainable ranking is not a feature.
Next.js · TypeScript · determinism audit · 25+ record sources
Version-control record: 2023-07-04 → 2026-08-20 · 1,391 commits
Flashcards become a playable game only when the material actually suits one — it scores the fit instead of forcing all three modes
Turns flashcards into playable Match, Order or MCQ mini-games, and scores which mode actually suits the material rather than generating all three and hoping.
deterministic generation · suitability scoring · HITL playtest
Both partners see the same exercise at the same moment — and the guidance comes from published sources, not the model's imagination
Guided communication exercises with both screens synced live. The content is source-grounded rather than model-improvised, which is the whole point in this domain.
Supabase Realtime · Next.js · TypeScript
Any image it made can be made again — the settings are recorded, so a result is reproducible instead of lucky
Image generation as a repeatable pipeline with recorded settings, so a result can be reproduced rather than re-improvised.
image generation · workflow capture
Change the questions without rebuilding the form — and every submission arrives structured, ready to use
A JS object defines the sections, field types and conditional logic; the form and its typed output follow from it. Change the schema, not the form.
schema-driven · conditional fields · typed JSON export
Mismatched money surfaces where someone will actually see it — exceptions raised, not buried in a spreadsheet
Transactions matched and exceptions raised where someone will see them, which is the only place a reconciliation tool earns its keep.
TypeScript · reconciliation · exception surfacing
No model grades its own homework — three propose, plain rules rank them, and a person makes the final call
One item routed to Ollama, Groq and Codex at once. A deterministic scoring layer picks the winner, then a person approves before anything is written to SQL. The cost work is the governance work — controls that are expensive get quietly dropped in month nine.
Ollama · Groq · Codex · deterministic scoring · human-in-the-loop
A new project deploys with its guardrails already installed — five minutes in, the checks are running
One command scaffolds a repo, a Vercel deploy and an agent-ready kit, with a router that sends each job to the cheapest model tier that succeeds. Governance that is expensive to run stops getting run.
Next.js · scaffolding · agent kit · model-tier routing
Version-control record: 2026-05-21 → 2026-08-20 · 152 commits
It refuses to give medical advice, and the refusal is enforced in code rather than promised in a policy
A health dashboard with a vision model on top, and the project where I had to decide where a display tool stops and a regulated medical device starts. Under the Cures Act, software that only stores and displays device data is statutorily not a device; the moment it recommends a dose it is Clinical Decision Support. So four features were specified and then permanently killed, and a screen now runs over every word the model writes before anyone reads it. Separately, a model call per dish returns a carb and calorie range scored against a 5,800-image labelled dataset in real time and reported as mean absolute percentage error (MAPE) — it missed its own accuracy target, so the feature is switched off. The live app is not linked, by a standing privacy decision.
FDA device boundary · HIPAA scoping · dose-language screening · Groq · dataset-backed eval
Version-control record: 2026-05-21 → 2026-08-10 · 104 commits
Caught this portfolio's own demos overstating themselves — the audit tool turned on its owner
The tool I used to audit my own work — which is how several of the defects on this site were found, including demos that claimed to be live while making no model call.
Next.js · scoring rubric · audit harness
Screens a housing deal in under a minute and says go or no — not a spreadsheet to interpret
NorCal affordable-housing underwriting suite: HCV rent-cap screen, RCFE calculator, and deal screens that return a verdict rather than a number to interpret.
Next.js · TypeScript · real-estate underwriting
Packaging copy and product imagery for pennies, on your own machine — nothing uploaded, no per-use bill
Local-AI packaging studio: brand positioning copy and Stable Diffusion image prompts, with an SSRF-safe image API. Runs on local Ollama, so the marginal cost is zero.
Ollama · Stable Diffusion · Next.js · cost engine
One post becomes four channels overnight, and non-technical staff run it daily
One blog topic fans out to a Facebook caption, an Instagram caption with hashtags, and an email subject and body from a single model call. Runs as an admin tool used daily by non-technical staff.
Groq · Next.js · publish API · nightly cron
Version-control record: 2026-03-05 → 2026-08-10 · 187 commits
Shipped, live, and serving users on its own domain
Shipped and serving users. The demo runs the governance end to end on fictional families: consent that fails closed with teach-back and revocation, honest identity disclosure on a direct ask, and the deterministic family-alert ladder. The live product is not demonstrated directly, so this is the mechanism rather than described from guesswork.
Next.js
Version-control record: 2026-07-18 → 2026-08-20 · 433 commits
One set of questions answers the two things people actually need together — the tax bill and the benefits they qualify for
A wizard that collects filing status, income, household and deductions, then returns an estimate alongside benefits eligibility (including CalFresh) — the two questions people actually need answered together.
Next.js · multi-step wizard · eligibility rules
A couple sees the same picture of the week — load, plans, check-ins — instead of arguing from two different ones
The shared-home foundation for Calm Couch’s Our World: weekly focus and load tracking, family moments, structured check-ins, requests, agreements and repair flows. The demo runs on fictional data and ships no images; the live app is not linked, by a standing privacy decision.
Next.js · TypeScript · Supabase
Paying customers study for a licence exam in two languages — and no translated question ships until five checks agree it kept its meaning
Tutored exam-prep product with spaced-repetition scheduling and bilingual EN/ES content. The localisation pipeline is resumable and cost-capped, so a bad run is cheap to abandon.
Groq · Next.js · Supabase · TypeScript · Stripe · SM-2
Version-control record: 2026-02-18 → 2026-08-10 · 1,132 commits
You can argue with the inputs instead of having to trust the output — every projection shows the assumptions it stands on
A model whose inputs you can argue with. Numbers you cannot interrogate are numbers nobody should act on.
TypeScript · scenario modelling
Stops an inconsistent asset run before it reaches production — a second model checks the first and halts the line on disagreement
One model generates assets, a second verifies them, and the run halts if drift crosses a fixed threshold. It exists because a single model generating a large asset run drifts confidently and quietly, and nobody notices until the whole set is wrong.
GPT-4o · Gemini vision · consensus gate · spec-driven pipeline
Version-control record: 2026-04-09 → 2026-08-20 · 169 commits
Receipts
The source stays private. The aggregate timeline does not: it shows when the work began, how much of it exists, and the dated records behind the named systems.
| System | First commit | Last commit | Commits |
|---|---|---|---|
| Find Your Vote | 2023-07-04 | 2026-08-20 | 1,391 |
| TradeTEST.TRAINING | 2026-02-18 | 2026-08-10 | 1,132 |
| TradeTEST content pipeline | 2026-03-22 | 2026-08-10 | 229 |
| Stillwell | 2026-07-18 | 2026-08-20 | 433 |
| Sea Star | 2026-03-05 | 2026-08-10 | 187 |
| Vision Factory | 2026-04-09 | 2026-08-20 | 169 |
| ostup | 2026-05-21 | 2026-08-20 | 152 |
| Plate Check — vision estimate accuracy | 2026-05-21 | 2026-08-10 | 104 |
| This site | 2026-06-16 | 2026-08-20 | 76 |
| agent-gate | 2026-05-29 | 2026-08-20 | 5 |
Generated from local git history on 2026-08-20. Only aggregate dates and counts are published. No code, file names, or commit messages are published. Counts are self-reported; the 4 July 2023 start date is independently visible on GitHub’s contribution record (opens in a new tab).
Stack
The models, infrastructure, evaluation patterns, and delivery systems used across the evidence index.
Legacy visual view
The older visual gallery remains available for scanning thumbnails and capability filters. The evidence search above is the maintained index and the better route when you are looking for a particular term or date.
Next step
The governance portfolio carries the selected cases. The résumé has the compressed career record and opens without an email wall.