This series followed one feature from a one-line ticket to an open pull request: paying with a saved card at checkout, in a Spring Boot and React pizza app. This last post puts the whole run in one place, with the real numbers, where the assistant helped most, and every point where reading the work changed the outcome.
The whole run
Every step was a real Claude Code session, and every session's time and usage cost was recorded. Implementation steps resumed one conversation so each built on the last; the debugging and review sessions started fresh on purpose.
Step Mode Assistant time
Analysis of the existing checkout plan (r/o) 1m 41s
Requirements: 9 questions plan (r/o) 1m 02s
Seven-step plan plan (r/o) 2m 46s
Step 1 data access + expiry accept edits 2m 01s
Step 2 service + Stripe update accept edits 3m 08s
Step 3 endpoint + security accept edits 1m 46s
Step 4 frontend helpers accept edits 56s
Step 5 card chooser (+ a11y fix) accept edits 7m 42s
Step 6 Playwright tests accept edits 10m 43s
Step 7 docs accept edits 1m 14s
Debugging the report tests fresh, r/o 2m 15s
Pull request description resumed 1m 52s
Independent review fresh, plan 3m 10s
Review fixes accept edits 3m 08s
~43 min, about $14.50 in usage
Forty-three minutes of assistant time produced ten commits: 25 files, about 1,700 lines, 34 new backend tests and 7 new browser tests. My own time, reading, verifying, deciding and committing, was not measured precisely, but it was several times longer than the assistant's. That ratio is the most honest number in the series. The typing got cheap. The judgement did not.
Where the assistant helped most
- Reading widely, fast. An eleven-hop trace of an unfamiliar payment flow, with a line reference for every claim, in under two minutes. Eight of eight claims checked out.
- Finding questions. Nine pointed questions for the product owner, including the one nobody had asked: customers can only save cards on their profile page, so for many of them the ticket would change nothing.
- Disciplined implementation. Each step stayed in scope, ran its own checks, and listed what it did that the plan had not said.
- Proving tests. Breaking the code on purpose to show each important test fails, then restoring it.
- Honesty about gaps. "Couldn't verify", "needed approval", "nobody has looked at it in a real browser". Every one of those lines pointed at real remaining work.
Every point where reading mattered
None of these were caught by a passing test suite. All were caught by someone reading:
- Step 1: an injected dependency nothing used, wired "for later". Removed.
- Step 1: four failing tests reported as "probably not mine" without proof. Proven pre-existing by stashing the change and re-running them.
- Step 3: a test expectation changed from the plan, 401 to 403. Checked against an existing test before accepting; justified.
- Step 5: a radio group with no accessible name. Sent back, fixed, and verified by querying the page for the group by name.
- Step 5: my own false finding about button colours, caught by checking it before sending it.
- Walkthrough: a real, pre-existing checkout bug, found because the screenshot script loaded the page the way a user with a bookmark would. It became a support ticket.
- Step 7: a documented root cause and fix that were both wrong. Corrected after a fresh debugging session found the real cause and SQL confirmed it to the digit.
- Debugging: my own wrong assumption about which database the tests used, caught by the session checking its premises.
- Pull request: a description with numbers that had gone stale. Re-checked against the branch and corrected.
- Independent review: an already-paid order telling the customer to "choose another card". Missed by the author session, its tests and me. Fixed, with tests that fail without the fix.
Ten catches across one feature. Most were small. Two would have reached customers: the payment message, and the checkout bug that was already live. One would have sent the next engineer chasing a fix that could not work. Any of them could have been missed by clicking Accept on a well-written summary.
What is still open
A finished feature is not a finished list. These were all written down rather than quietly dropped, and each has an owner and a reason it waited:
- Saving a card at checkout, the product owner's separate ticket.
- The same feature in the Angular, React Native and SwiftUI apps.
- Four review findings declined for this pull request: two tabs choosing different cards, a late card list swapping the form, a failed reload hiding a message, and a Stripe call inside a database transaction.
- The report tests that depend on how old the database is, with the root cause attached.
- The checkout name-and-email bug from the support ticket.
A follow-up list like this is what makes a scoped pull request trustworthy: the reviewer can see exactly what was left out, and exactly why.
What the run would look like without the habits
It is worth being concrete about the alternative. Paste the ticket, accept the code, accept the tests, accept the description, push. The feature would mostly have worked. It would also have shipped an unused dependency, a misleading note about why the build is red, a pull request with wrong numbers, and a checkout that can invite a paying customer to pay twice. The team would have found out later, from someone else, at a worse time.
With the habits, the same assistant did the same fast work, and the problems were found before they mattered. The difference was never the tool. It was whether a person read what it produced.
What to take from the series
If you keep one thing, keep the order of work: understand, ask, plan, then build one small step at a time, reading each step against the plan. Use fresh sessions for review and debugging, use read-only modes and allowlists wherever you can, verify the claims you depend on, and make every decision that belongs to a person yourself. The assistant makes each of those steps faster. It does not make any of them optional.
Before you accept
- Can you explain every change in the pull request without the chat open?
- Did every step get read against the plan before it was committed?
- Was the work reviewed by a reader who did not write it?
- Is every known gap, deferred finding and leftover side effect written down?
- Were the product, publishing and destructive decisions made by a person?
- Would you be comfortable being paged for this at 2am?