On this page
The site you’re reading was rebuilt in one sitting. I opened a Claude Code session one morning with a list of things I didn’t like about the old site. Fourteen hours later the new one was live on Cloudflare, and by the time the post-launch fixes were in, the session had run 15.4 hours. In between, 30 pull requests merged and 227 commits landed on master.
I typed 33 messages in that time. The agent made about 5,000 tool calls, spread across its own thread and 49 other agents it spun up along the way. That ratio is the most interesting thing about the whole day. Most of what I contributed wasn’t instructions. It was taste, cost calls, and deciding when things could and couldn’t ship.
This is the same experiment as my Opus 4.8 post, just bigger and messier. A personal site is still the right kind of target. Nobody gets paged if the mascot falls through a card.
What I asked for
The brief was a list. I didn’t like:
- the WebGL backgrounds everywhere
- the homepage turning into an about page
- gradient cards on articles
- the phrase “case studies”
- word clouds
I liked the colors, the logo (though it needed redrawing) and the mascot, a little pixel diamond with a hand-built physics engine that wanders the page. I wanted the site to show projects and writing instead of reading like a resume. Also, while we were in there: upgrade Astro, better analytics, better SEO.
My second message was the most important one I sent all day: “I believe we have a lot of upfront planning to do before we get any work going ya know?”
Two hours before any code
The first pull request didn’t open until 2 hours and 20 minutes in. That time went on three things.
The interview came first. Six rounds of structured questions: nav, homepage contents, what happens to the portfolio, page depth for projects, logo direction, motion, typography, whether to keep a light mode, how to handle view counts. Every answer went into a decision log, docs/redesign/spec.md, before anything changed. That log paid for itself all day. Whenever a later question came up, the answer was usually already written down with a reason next to it.
While I was still answering questions, a research workflow ran six read-only agents in parallel. They covered a site audit, the Astro upgrade path, image models, analytics, an SEO audit and a visual baseline, and a seventh agent read all six sets of notes looking for gaps. The whole thing took 12 minutes.
It caught the thing that would have hurt most. The Cloudflare config had no pinned id for the KV namespace that stores public view counts. Upgrading the Cloudflare adapter could have quietly created a fresh, empty namespace and wiped every count on the site. The id got pinned before the upgrade. After launch, copilot-plugin-cc still showed 56 views and 53 unique, same as before.
Then came design. A second workflow pitched four logos, a set of type pairings and two whole homepage concepts, then handed everything to a judge agent that scored each pitch against the spec. Each logo also got rendered as a real 16px favicon, to check it at the size people would actually see. Both mascot-based logos lost. The judge’s line was that “the mascot works better as the mascot than as the logo,” which is correct and a little funny.
I didn’t take either homepage concept. I took the workbench look from one and the scout interactivity from the other. That’s a call a judge agent can’t make for you, because it depends on what you’ll enjoy looking at every day.
Building in waves
The build ran as milestones, M0 through M10, over about four and a half hours. The upgrade and cleanup went first and ran alone. The upgrade was Astro 5 to 7, and 7.3.4 turned out to be two majors further than expected, so it went through 6.4.8 on the way. Everything after that branched from the upgraded base.
Then the fan-out started. Each standalone agent got its own git worktree and opened its own PR. The biggest wave was five at once: the about, projects, writing and work pages, plus the cover image pipeline. There were four waves in all.
Serial work stayed serial on purpose. Running the upgrade and the content-model rename in parallel with anything would only have produced merge conflicts.
24 standalone agents shipped code. The busiest, the homepage build, made 245 tool calls. The longest, the cover pipeline, ran 98 minutes, mostly waiting on image generation and on me picking images.
Before launch a third workflow did an adversarial review. Five reviewers covered correctness, accessibility, performance, security and content. Then a skeptic per dimension tried to disprove each finding. Only findings that survived counted. 21 did, and one more agent fixed all of them in a single PR.
What my 33 messages did
Looking back at the prompts, very few were instructions in the “build X” sense. Most fall into four buckets.
Taste. The agent called the image gallery “AI art.” I killed that: “Don’t want to pretend it’s something they aren’t.” The images still show their model and prompt, and the page is just called Images. I also cut two bits of handwritten-style copy I had approved earlier, because once they were live on the page they read corny. I picked the headline from a numbered list by typing “4. / B.” I protected the Bible verse under the hero, which no interview round would have surfaced, because it isn’t a design question.
Cost. After a bake-off between three image models, the agent recommended Nano Banana 2 for consistency. I went with GPT Image 2.5 Flare, because it was good and cheap, and Nano Banana’s images “sucked and was expensive.” Later: “no more mai generations lol too expensive.” All in, every cover image on the site cost about three dollars.
Control. “Don’t complete the PR to main yet tho lol still working on this.” Then, hours later: “I think it’s ready. Get it merged.” The agent had full autonomy between those two messages, and none outside them.
Facts only I have. Three kids, six dogs, two cats, coaching the kids’ baseball team, which photo of Brindlee is the good one, and which of my repos are private ideas that must never show up on the site. Those became standing rules in the spec.
What went wrong on my side
I put the OpenRouter API key in .env.production. That file is committed to the repo. The agent caught it before anything was committed and told me to move it to .env, instead of moving it silently.
I approved copy in the abstract and rejected it in context. The corny notes and some static “scout says” speech bubbles only looked wrong on the real page. Feedback got cheap once there was a preview to look at. I should have asked for one earlier, per milestone, instead of saving an 18-item list for the end.
I flagged the duplicate-looking covers twice, and the second time I waved it off: “seems prompts are same again but whatever not important.” It was important.
And I asked for a rebase merge on the launch PR so all the commits would show up individually. GitHub can’t rebase a 204-commit branch that has merge commits in it. A merge commit kept the history anyway, which is what I actually wanted.
What went wrong on the agent’s side
The duplicate covers were the agent’s bug. The first round of image prompts only described style, motif and color, so every post sharing a preset got a near-identical prompt and a near-identical picture. The fix took three passes:
- Add the post’s own description to its prompt.
- Stop reusing motifs within a topic.
- Hand-set presets for the AI posts.
A cleanup PR said it removed four unused dependencies. The commit only staged one config file. The agent noticed on its own and shipped the real removal in a follow-up PR.
The site moved to trailing slashes, and the mascot’s page matching was exact-string, so it quietly lost its page-specific lines. The mascot agent caught that one in its own pass.
The hero’s social links took two tries. The first mobile pass left them in a detached column floating off to the right, and I had to point it out again after launch.
There was friction in the tooling too:
- The worktree safety guard refused complicated shell loops twice. Both times the agent rewrote the logic as a plain script file instead of fighting it.
gh pr editfailed on an unrelated GitHub deprecation, so it used the raw API.- On the Cloudflare dashboard, filling fields programmatically set values the React UI never registered. It caught that by re-checking, then clicked and typed by hand.
Green tests lied again
In the Opus 4.8 post it was 48 passing tests on a broken tool. This time three separate real bugs got past a green suite.
The view counter was supposed to ignore bots. The unit tests passed. But when the agent hit a preview build, its own request got counted. The local proxy sends an undici user agent, which the bot filter didn’t know. Seven of the new tests failed against the old code, which at least proved the fix was real.
The covers passed every check the pipeline had. They were still duplicates. It took a person looking at them.
Then after launch, on my phone, the scout sank about 20 pixels into the card it was standing on. It didn’t happen on desktop, and the agent couldn’t reproduce it in Chromium at all. It said so plainly, went looking, and found WebKit bug 297779: iOS 26 Safari shifts position: fixed layers when the toolbar collapses or the scroll changes direction. The fix moved the mascot’s drawing layer into the page itself and left the physics engine alone.
Every one of those was found by looking at the real thing.
The numbers
Lighthouse on mobile, one run per page, same version before and after. Single runs are noisy, so treat a few points either way as noise.
| page | perf | a11y | LCP | blocking time |
|---|---|---|---|---|
| home, before | 92 | 96 | 2.7 s | 110 ms |
| home, after | 92 | 100 | 2.6 s | 190 ms |
| writing index, before | 77 | 98 | 2.4 s | 770 ms |
| writing index, after | 91 | 100 | 2.8 s | 60 ms |
| a long post, before | 83 | 84 | 4.4 s | 0 ms |
| a long post, after | 92 | 100 | 2.9 s | 140 ms |
Accessibility and best practices hit 100 on every page, and layout shift went to zero. The writing index stopped blocking the main thread for most of a second; the old client-side search and tag chips were the cause. Home performance didn’t move. The writing index got heavier, because it has real cover images now instead of gradients.
The session itself, by the numbers:
- 15.4 hours and 33 prompts from me
- about 5,000 tool calls: 703 in the main thread, 3,259 across 24 standalone agents, 1,055 across 25 agents in three workflows
- 30 merged PRs and 227 commits
- 698 files changed, +18,789 and -17,676 lines
- tests grew from 236 to 462
- about 572,000 output tokens and about 879 million read back from cache
- one context compaction
That cache number is why a session this long was affordable.
The covers, and a cheap model making decisions
Every post now has a real cover image. The pipeline uses two models for two jobs.
Jev, a small structured-decision model from TypeSafe AI, reads each post and picks an art preset and motif. It also flags posts that need a real photo instead of a generated one. It costs about $0.0001 a post. It was nearly certain about art direction, and it correctly flagged the Brindlee post as real-photo-only. Its color picks were shaky, so color got tied to the post’s topic instead.
GPT Image 2.5 Flare does the drawing. It drew a terminal prompt onto an NVMe drive nobody asked for, and it returned the same resolution no matter what was requested. In round three it snuck in a Discord logo.
134 candidates got generated across 16 posts, and 18 covers shipped. Nothing else went to waste: the unpicked ones live on the Images page for anyone to download.
What I’d do differently
Ask for a preview after every milestone, not one at the end. Most of my feedback was only possible once I could see the page, and saving it for one big list made a big batch of rework.
Put the taste rules in the spec up front. “Don’t call AI images art” and “no cute copy” are rules I already held. I just didn’t know I needed to say them until something broke them.
Keep the planning. The two hours before any code were the best-spent part of the day.