Skip to content
Hardpack.

Case study · The build system

Your project is not the first one through this

Most engagements start by setting up the same things over again. A Hardpack build starts on a system that has already been through 155 reviewed pull requests. It was debugged on my time rather than on yours.

That system is a work board, a set of guardrails that stop the dangerous git commands before they run, and a package that carries them from one project into the next. What it buys a client is a first week spent on the work rather than on deciding how the work will be run.

Below is what runs on it today, what a new project gets on day one, and what it cost to build.

5
Repositories on the same standard
155
Reviewed, merged PRs behind it
5–9
Est. weeks of engineering time avoided
5.8×
Human review vs. the local model, over 7 runs

The situation

An AI agent hands you volume immediately. Checking that volume is where the time comes back.

If I have to read 200 generated commits the way I would read a stranger's pull request, I have saved nothing. And "looks right" stops being a usable standard somewhere around the fifth one.

The thing worth building was never a fast week. It was speed that survives into the next project, and into the one after that. So the checking went into the system, and I instrumented what it cost.

What runs on it now

The same system runs every project Hardpack takes on. It comes in three parts, and they do not all travel the same distance.

One board, 11 projects, 1,165 tickets. Every project's work is tracked in one place, whether or not that project uses any of the rest.

5 repositories run the versioned standard. The guardrails are pinned by tag out of ticket-workflow, a public package with its own CI gate and a protected main. Pinning by tag is what makes it one standard rather than five copies that have quietly drifted apart.

The machine-level wiring does not travel. The board server, the telemetry writer and the guard that stops a subagent crossing a human approval all live in machine-local configuration. A fresh clone on another machine gets none of them. That part of the setup is real, and it is not transferable.

The system was built under its own rules and tracked on its own board. 362 commits across 38 active days, 330 of them AI-co-authored, 1,772 tests.

The three reach figures were measured 2026-08-20 and are self-reported: the board is not a public repository, and most of those repositories are private, so unlike the figures below there is no command you can run against them. The repo figures were verified 2026-08-20and you can check them yourself — git rev-list --count HEAD · npx vitest run.

Two client builds run on it today. The domains have nothing in common. The starting setup was identical in both.

Equipment schedules and submittal review for a mechanical engineer. He specifies the equipment, the vendor returns a submittal PDF, and the app checks it back against the schedule for values outside tolerance and units nobody quoted. Doing that by hand means re-keying the vendor's numbers into the schedule row by row.

An auction filtering tool for an independent mechanic. It pulls a capped result set off a salvage auction site and makes it workable on a tablet out in the yard.

What a new project inherits on day one

Four decisions, each one a guarantee bought by giving up a little of the agent's speed. A new project starts with all four already made and already debugged. The later ones only hold because the earlier ones do.

1. Typed tools instead of a shell

The board is exposed through a from-scratch MCP server, so the agent calls typed functions rather than hand-editing markdown. Failures move from runtime to the boundary. update_ticket validates status against a shared enum, and that check sits in the service layer every caller goes through, so the agent cannot invent status: "in progres" through the board's tools.

2. Rules enforced by code rather than by honor system

One ticket, one branch. A guard hook inspects every shell command and blocks the dangerous shapes (git add -A, commits to main, force-push, reset --hard). The merge itself is held by branch protection instead: a pull request, a green gate, and no bypass for anyone. That runs on GitHub's side and never reaches the hook. Squashing to one commit is my own convention on top of it, and nothing in the platform enforces that part. An agent that can be talked out of a rule does not have a rule. In a prompt it holds until the context gets long. In a hook it holds because the command does not run.

3. "Done" has an enforced half and a recorded half

The definition of done is a checklist: typecheck, lint, tests, and a summary carrying a test line, either Tests: N added or Tests: none — <reason>. In the board's own repository those three run in a required check on every pull request, and a red result blocks the merge for everyone, including me. The fourth is a different kind of item. Ticket bodies live outside the repository, so CI never reads that line. It is what the workflow asks for and what I check at review, and calling it enforced would be the same overstatement this checklist exists to prevent. "None" is allowed. Silence is not. Making the skip explicit turns an omission into a decision somebody has to defend.

4. Cost measured rather than estimated

The intake agent runs local-first against an OpenAI-compatible endpoint, with no cloud key and no per-call charge. That creates a problem. Every LLM cost dashboard multiplies tokens by a published price, and locally there is no price. So the cost model reports measured, assumed and externalities separately and refuses to blur them. The numbers further down show why that matters.

Two more things a new project inherits, and they matter most when the data is yours. The intake agent talks to a local model, so there is no cloud key at runtime and unplugging the network does not stop it answering. And the board's tickets are working notes that stay on the machine they were written on.

What it cost to build (with Claude)

This is the Claude side of the ledger, separate from the local intake agent the app runs. I reconstructed the build economics from Claude Code's own session logs and committed the script plus a frozen snapshot to the repo, so the method is auditable rather than asserted.

The part anyone can check against the git history: 155 reviewed, merged PRs and 26,693 lines of code, about 5.3 merged PRs a day.

Verified 2026-07-21, reproducible by running git log --oneline --until=2026-07-21 | grep -Ec '\(#[0-9]+\)' in each of the 2 repositories the snapshot covers and summing the result.

Set against a by-hand counterfactual at a deliberately low 2–3 hours per PR, that comes to roughly 5–9 weeks of engineering time avoided, against a measured Claude bill of ~$2,485. Do not take the range on faith. Move the assumptions yourself and watch the hours, dollars and ROI recompute.

The asterisks, because a number without them is marketing: the dollar cost is self-reported from local session logs. Only the PRs, lines of code and velocity are independently reproducible. It is a floor, because the CI code-review agent's usage runs on GitHub's servers rather than in those logs, so it is not counted. And the by-hand hours are an estimate. I anchored on merged PRs rather than ticket counts (git-verifiable, and I had already watched a ticket-based count swing threefold), deduplicated the raw token counts (each streamed response is logged several times), and scoped them to these repositories. A case study about honest measurement does not get to skip its own footnotes.

Why two models at all? Different jobs. Building the app is a one-time, capability-hungry task where a frontier model earns its keep. The intake agent runs on every report, over operational data you may not want leaving the building, so recurring cost and privacy both point local. Frontier model for the build, local model for the runtime.

The proof

It searched before it wrote

The intake agent's first move on any note is retrieval, not generation. One real run, nothing tidied: someone reported that list_tickets silently truncates at ~440 tickets, and the board already had a ticket for exactly that.

query list_tickets overflow · top 5, by cosine similarity

  • 74%list_tickets overflows the tool-output cap at scalebacklog
  • 52%Drag-and-drop tickets between columnsbacklog
  • 45%Local-first intake agent with retrievaldone
  • 42%MCP server exposes the board as typed toolsdone
  • 37%Accessibility pass on the board UIbacklog

tool call update_ticket on tkt-list-tickets-output-cap

human gate approved

The agent found the ticket that already existed and called update_ticket on it instead of create_ticket, which is the anti-duplicate rule working, and a human approved it before anything touched the board. To step through a whole pipeline, including the note, every hit and its score, the reasoning, and the gate, walk through three real runs in the replay viewer.

The number that changed my mind

This is the other model, the app's own local intake agent, not Claude. A different run from the one above, costed, passed back through the same renderSummary() the CLI prints. I am not going to pretend one recording shows both. The ragged 0.2848888888888889 is why I do not round any of it. Real instruments do not round for you.

runId    087f1884-f6af-4689-918d-d05c4c9121d1
at       2026-07-14T16:31:36.484Z
model    openai/gpt-oss-20b
calls    2  activeMs 4787  reviewMs 20512
outcome  {"created":1,"updated":0,"declined":0,"noProposal":false,"errored":false}
— Measured —
  active compute                           4787 ms
  prompt tokens                            2246 tokens
  completion tokens                         138 tokens
  total tokens                             2384 tokens
  cacheable prefix                         89.2 %   (estimate (~4 chars/token))
  accepted tickets                            1 count

— Assumed ($) —
  marginal energy                      notional kWh   (notional — set active watts, idle watts)
  marginal energy cost                 notional USD   (notional — set active watts, idle watts)
  keep-warm energy cost                notional USD   (notional — set idle watts)
  hardware amortization                notional USD   (notional — set hardware cost)
  review time cost             0.2848888888888889 USD   (21s x 50/hr)
  total run cost                       notional USD   (notional - some cost inputs unset)
  manual value (avoided)       4.166666666666666 USD   (5min x 50/hr)
  cloud-equivalent (openai/gpt-oss-20b)         notional USD   (notional — no price configured for this model)

— Externalities (report-only) —
  water footprint                      notional L   (notional)
  carbon footprint                     notional gCO2e   (notional)

— Headline —
  cost per accepted ticket             notional USD   (notional - total cost incomplete)
  net savings                          notional USD   (notional - value or cost incomplete)
  local vs cloud (saved)               notional USD   (notional - cloud or local incomplete)

The line that reframed the whole project is the timing pair:

4,787ms model computevs20,512ms human review

On this run the human took 4.3× longer than the local model, which put the model at 19% of the total. A faster model has 5 seconds to play with. Cutting the review burden with better drafts and a tighter approval screen goes after 21. The industry conversation is about tokens and latency. In a human-in-the-loop workflow the human is the cost center.

The caveat, stated plainly: reviewMs is instrumented on 7 of the logged runs, a real sample now but a small one. Across them the review-to-compute ratio runs from 3.7× to 12.1×, median 5.8×. The run costed in full above, at 4.3×, actually sits near the low end. That is enough to trust the direction, that the human dominates the loop, without pretending the exact multiple is a law. A case study about honest measurement does not get to round that up.

The win it refused to book

On another run the agent read the note, searched the board, and proposed nothing, because an existing ticket already covered it. Watch what the cost model does with that. On the accepted run it books manual value (avoided): 4.166666666666666 USD, the typing the agent saved. Here it books 0: no accepted ticket - nothing avoided. Across 66 runs, 53 earn that credit and 13 do not, so $220.83 is honest. Booking it unconditionally would read $275.00, a 25% overstatement of the headline from a one-line change nobody would ever see.

If you have ever watched a demo fall apart under a customer's questions, that is the discipline that stops it happening to you. A dashboard that only reports wins is marketing wearing the clothes of measurement, and the first person it fools is the one holding it.

Who this is for

If you are weighing whether AI-assisted delivery is reliable enough to put real work through, this is what the measured version looks like. The system above is what your project would start on, and none of it would be assembled for the first time on your budget.

The one number I would put in front of anyone betting on this space is 5.8×. It was measured on the local agent rather than on Claude, and the lesson carries either way. The model is cheap. The human is the cost.

For the same method on a client’s problem rather than my own, read an auction filtering tool for an independent mechanic.

For the same method on somebody else’s problem rather than my own, read a submittal check built for a practicing mechanical engineer.