AI-Assisted Engineering – Planning: The Plan Is the Contract

September 14, 20265 min readUpdated 10/4/2026

A plan is the cheapest place to be wrong. Changing a sentence in a plan costs seconds; changing the same decision after three commits costs an afternoon and a rebase. With an AI assistant the gap is wider, because code is now so cheap to produce that it is tempting to skip the plan and "just see what it writes". This post argues for the opposite: the plan is the contract every later diff is reviewed against, so it deserves more of your attention than any single line of code.

Plan mode, and what to ask for

Planning ran in Claude Code's plan mode, which is read-only: the session can read the codebase and think, but cannot edit anything. It resumed the same conversation as the analysis and requirements steps, so it already had the code, the product owner's answers, and the edge cases. The prompt asked for a specific shape:

Write the implementation plan. Requirements for the plan:
- Split into numbered steps, each small enough to be ONE reviewable commit,
  backend before frontend.
- For each step: the files it touches, what changes, and how we verify it
  before moving on (which test, which command).
- Follow the rules in CLAUDE.md (DAO rule, service interface + impl, 404 not
  403, spotless, never edit an applied changeset).
- List explicitly what is OUT of scope, so I can hold every later diff against
  this plan.
Don't write any code yet.

Each requirement is there for the review that comes later. "One reviewable commit" per step keeps every diff small enough to read. "Files it touches" gives you a list to compare git status against. "How we verify it" means each step ends with evidence, not a feeling. And the out-of-scope list is what makes drift detectable at all.

What came back

The plan opened with the design in two paragraphs, then flagged the decisions a reviewer should check rather than burying them:

Two decisions to check in review:
- The update call has no idempotency key, on purpose. A key per order and card
  would make a later switch from card A to B and back to A return Stripe's
  saved result from the first A request without changing anything. Setting the
  same values twice is already safe.
- The card is chosen on the payment step, as the PO asked. That's why this
  updates the PaymentIntent rather than adding a field to the order request.

The idempotency point is subtle and correct, and it is exactly the kind of thing that gets "fixed" later by someone who sees a payment call without a key and adds one. Writing the reason into the plan, and later into a code comment, protects it.

Then seven steps, one commit each: data access and the expiry rule; the service method and Stripe update; the endpoint and security rule; frontend helpers with no visible change; the card chooser; Playwright tests; and docs. Each listed its files and its checks. Step 3's checks, for example, named every status code the new endpoint should return and the test that would assert each one. It closed with ten out-of-scope items and two named risks to test first.

Review the plan like code

Read the plan with the same suspicion you would bring to a diff, because every mistake in it will be faithfully implemented. Questions worth asking of any plan:

  • Is each step independently verifiable? Step 4, "helpers, no visible change", existed so the risky UI step 5 would start from tested building blocks.
  • Is the order right? Backend first meant the frontend was built against an endpoint that already had passing tests.
  • Do the checks prove anything? "Run the tests" is not a check. "This test asserts a deleted card returns 404" is.
  • Are the risks real and testable early? The plan admitted it had not confirmed that confirmCardPayment would work with how the PaymentIntent was created, and named a fallback. Step 5 confirmed it worked; no fallback was needed.

The plan also contained one expectation that turned out to be wrong: it said a request with no token should get 401. The app returns 403 for every protected endpoint, and step 3 had to deviate. That is fine. A plan is a contract, not a prophecy. What matters is that the deviation was visible because the plan had been specific. A vague plan ("secure the endpoint") would have hidden it.

When to re-plan

A deviation in one step can invalidate assumptions in a later one. The 403 surprise did: the step 3 summary pointed out that "the frontend can't tell 'signed out' from 'forbidden' by status code on this endpoint", which mattered for the error handling planned in step 5. When that happens, update the plan before continuing, even by one line. Otherwise later steps are reviewed against a contract nobody believes any more, and review quietly turns back into scrolling. Ten seconds spent editing the plan prevents that drift entirely.

Keep the plan where the work can see it

A plan that only lives in a chat scrollback is lost when the session ends. Plan mode saved this one to a file automatically, and the important decisions went into the project's progress report as the work went along. When implementation started, each step's prompt was short because the plan carried the detail:

The plan is approved. Implement STEP 1 ONLY (backend: data access and expiry
rule), exactly as planned. Run the verification for step 1 (spotless:apply,
then the tests). Do not start step 2 and do not commit; I'll review the diff
and commit it myself.

"Step 1 only" and "do not commit" are doing real work there. One step at a time keeps each diff reviewable. Committing yourself means nothing lands until you have read it.

The planning session also noticed something outside the code entirely: the user's global instruction file required a co-author trailer on commits and the project file forbade it. It asked which to follow instead of guessing. A good plan surfaces conflicts like that before seven commits inherit the wrong answer.

Before you accept

  • Is every step small enough to be one commit you will actually read?
  • Does every step name its files, so you can compare them to the diff?
  • Does every step end with a check that would fail if the step were wrong?
  • Is there an explicit out-of-scope list?
  • Are the non-obvious decisions written down with their reasons, so nobody "fixes" them later?
  • Are the risks named, with a step that tests them early?
  • Is the plan saved somewhere the next session will read it?