← Blog Posts
September 6, 2026 9 min read AI Claude Code

Four Projects, Four Stacks, One Process

I shipped four things this year, and no two of them share a line of code.

  • Current: Recovery Score. Swift, iOS and Watch. Paid, shipped, now on 2.2, with 161 unit tests behind it.
  • Reply All. A browser game on a stack that costs nothing. The interesting decision was a rules-based content pool with an eligibility engine rather than a branching story graph.
  • A book club app. A day’s work for an internal, employee-run club.
  • Salesforce demo environments. Complex demos seeded and customized in an hour or two.

Four unrelated technologies. Nothing carried over. Not a language, not a framework, not a deployment target, not a single reusable file.

What carried over was how I worked.

All four were built in Claude Code. I’ve written about the mechanics of using it well before, so this isn’t a list of prompting tips. The tips change every few months anyway. Each of these four modes is an answer to the same question, asked at a different point in the work: where does the human belong in this piece of it.

The four modes

Each one is a different relationship, and the output is a different kind of thing in each. That second part took me longest to see.

1. Planning: Claude as a senior peer

I bring an idea. Claude pressure-tests it before I commit to anything. Is it feasible, what shape should it take, what’s going to bite me in week three.

Reply All started as a feeling more than a design. I wanted the old dial-up BBS door games, the ones you logged into once a day to find out what had happened to you overnight. The obvious way to build that is a branching story graph, because that’s what a text game looks like from the outside.

The useful question wasn’t how to build one. It was narrower and less flattering: what breaks first if I write it that way. Two answers came back. Every new scene multiplies the paths I’d be keeping coherent by hand, forever. And a branch tree was never how those games worked anyway. They were stat machines. What happened to you tonight depended on what you were carrying and who you’d annoyed, not on a path someone drew in advance.

Reply All shipped as a rules-based content pool for that reason. Every scenario declares its own eligibility conditions and an engine picks what fits your state. Adding a scene means writing one file and not thinking about any other file.

What to withhold here is the answer I already want. Lead with my preferred approach and I get agreement instead of assessment. Agreement is worth nothing at this stage.

2. Execution: architect to developer

The design is settled. Build it.

Salesforce demo environments are almost pure execution. I know what the demo has to prove before I open anything, because the shape of a demo is decided by the conversation it’s supporting. There’s no design question left. There’s a lot of careful, specific work: objects, data that hangs together under scrutiny, the flows a prospect will actually click through. That’s the mode where an hour replaces a day.

What to withhold here is openness. Leave the design negotiable and you get options when you wanted code. Every option is a decision I already made, handed back to me.

3. Iteration: Claude as product owner

What’s missing, what’s confusing, what would a real user hit that I’ve stopped being able to see.

Reply All read well and played badly, and I couldn’t tell from the inside. Two findings came out of this mode that I wouldn’t have reached on my own. Scene selection was technically correct and felt random, because nothing preferred to continue the thread you were already in. And runs ended far too early, because most endings had no pacing gate at all, so a decision you made three scenes in could drop you straight onto an ending screen.

Both were obvious once named. Neither was obvious to the person who wrote them.

What to withhold here is my defense of the thing I just built. The moment I start explaining why it works, I get compliance instead of critique, and I’ve wasted the only mode that can see what I can’t.

4. Review: Claude assembling a decision, not making one

This is the one I wouldn’t have listed a year ago, and it’s now the one I rely on most.

The output isn’t work. It’s the conditions for a judgment: facts assembled, real alternatives with their tradeoffs, what’s verified separated from what’s assumed, and the thing that would change my mind put where I can’t miss it.

What it takes in the prompt is mostly restraint. Make it surface what it doesn’t know instead of smoothing over it. Force genuine alternatives rather than one recommendation flanked by two strawmen. Label inference as inference.

What to withhold here is which way I’m leaning. Signal a preference and I get a decision package built to support it, which is the most expensive way to be wrong.

Why switching is the part that matters

The failure mode isn’t bad prompts. It’s using one voice for all four.

  • Stay in execution voice during iteration and you get compliance, because you’ve already cast yourself as the one who decides.
  • Stay in planning voice during execution and you get options when you wanted a build.
  • Skip review and you approve things you haven’t actually evaluated.

Most disappointing results come from never switching, not from poor wording.

A worked example: the score that was surest where it knew least

Current scores sleep efficiency from how much of the night you spent awake. The formula is 100 - (awakeHours / 1.5) * 100, and awakeHours was a plain Double. A sleep source that never reports wake time at all arrives as zero. Zero awake time scores a perfect 100.

So the app was handing out flawless efficiency scores to exactly the people it had the least information about.

What happened next is the whole argument of this post, so I want to be precise about it. Claude didn’t tell me to fix it. It told me what the bug actually cost, and the cost was bigger than I’d assumed.

Current already excludes signals it doesn’t have and redistributes their weight across the ones it does. So a fabricated 100 wasn’t just adding points it hadn’t earned. It was also eating the 10% that should have gone to HRV and resting heart rate, the two signals that were actually measured. The wrong number was crowding out the right ones.

Then the part I’d have gotten wrong on my own: the obvious fix doesn’t work. You can’t read zero as missing, because a real zero exists. Someone who slept straight through has genuinely zero awake time and has earned the 100. Absence had to be decided from the shape of the samples instead. Awake samples present, measure them. No awake samples but in-bed data, derive it. No wake evidence but staged sleep, trust it as a real zero. Nothing but unspecified asleep samples, return nothing and let the architecture drop the signal the way it already drops the others.

And the one I’d rather not print: I had looked at this before and cleared it. My own roadmap still has sleep efficiency filed under “not fixed, by design: defensible scoring choice, not a bug.”

None of that is a decision. It’s the map you need before you make one.

The decision was mine, and it wasn’t the code. Fixing this meant existing users’ scores would move. Anyone whose sleep data comes from something other than an Apple Watch would open the app one morning and see a different number. Current has a mechanism for exactly that: bump the algorithm version and every affected user gets a sheet explaining what changed. I’d used it in 1.6, when I pulled sleep stages out of scoring because wearables measure them inconsistently.

I chose not to use it. Current’s user base is small enough that I could count the people this would actually move, and firing a full-screen algorithm notice at everyone to explain a slight shift for a handful of them isn’t proportionate. It’s the same instinct as the fix itself: match the response to what the situation actually is, not to what would look most rigorous. The correction shipped with a plain line in the release notes instead.

That call cost something, and it’s worth saying so. The people whose scores moved are exactly the people least likely to read release notes. I decided a quiet, accurate fix beat a loud explanation nobody asked for. Someone else could reasonably decide the opposite, and that’s the point. There’s no technically correct answer here. Claude laid out what each option cost. I picked.

That’s the difference between a tool that does the work and a tool that makes you competent to decide. Going in, I didn’t know that a missing signal was quietly stealing weight from a measured one. I knew it before I chose.

Where it breaks

The process works because I can architect and review, not because the prompting is clever. I’m not hand-writing unassisted Swift and I don’t claim to be. Claude wrote most of the code in these projects, the same way Fable never wrote a line of Swift on a release it planned end to end. The architecture, the debugging, and every decision worth arguing about were mine.

Which means the process degrades exactly where my judgment does. If I can’t evaluate the output, mode four produces a well-organized decision package I’m not qualified to decide on, and I’ll approve it anyway because it looks thorough. The modes don’t add judgment. They put it in the right place.

Where the human belongs

I build human-in-the-loop systems for a living. On a matching engine I architected for a workforce nonprofit, the agent never finalized anything. It created a record, laid out its reasoning, and a counselor reviewed and approved before a single email went out. I’ve spent years deciding where the human belongs in agentic systems.

Then I noticed I run my own work the same way. The four modes are the same question asked four times: who owns this part, and what does the other party need to be useful?

The technology will keep changing. That question doesn’t.