Adil Islam
Sep 4, 2026 · Practice · 7 min ◉ Working note

Pull Requests, Now That Agents Write the Diffs

Push-to-main was the correct call for four years. Agents changed the arithmetic underneath it, and the review is now the only place a human meets the change.

The Setup This Argues Against

This site has no build step, no framework, and no CI. There is one rule and it fits in six words: push to main equals deploy. That rule has been right for as long as it has existed. One person wrote every line, the repository is static HTML, and a bad commit costs the ninety seconds it takes to write the next one.

Every argument for pull requests you have heard was written for a different situation than that one. Four engineers on a shared service. A release train. Compliance. None of it applies here, and adopting a process because it is what serious teams do is the most expensive kind of cargo cult, because the cost arrives daily and the benefit never does.

So the case has to be made on what actually changed. Something did.

The Thing That Changed

Push-to-main worked because the diff already lived in the author's head before it existed in the repository. Review was not skipped. It happened continuously, during authorship, by the person who was about to type the code. The commit was a formality confirming a decision that had already been reviewed by the only reviewer who mattered.

When an agent writes the diff, that step is gone. The code arrives fully formed, competent, plausible, and unread. Nobody has thought about it yet. If it lands on main without a human or a check in between, the first time anyone evaluates the change is after it is serving to the public internet, and typically because something looks wrong.

A pull request is not a bureaucratic step added to authorship. It is the place authorship used to happen, relocated to where it can still be observed.

That is the whole argument. Everything else is implementation.

What the Repository Already Knows

The strongest evidence for this is in the git log. Four commits, none of them interesting individually:

4ca6882 fix: absolute fetch URLs via base href (no-slash URL broke data loads)
c91c167 fix: harden JSON loading against stale cache / 404
c255581 Fix CSS path resolution for local file:// viewing
3937f93 fix: v3 column alignment via content-word overlap

Three of those four are the same bug wearing different clothes: a path that resolved in one context and not another. Each was found in production, by looking at the live site, because there was no earlier place to find it. Each cost a round trip of notice, diagnose, fix, push, wait for the deploy, check again.

Not one of them required a human reviewer to catch. A script that loads every page, waits for the network to go quiet, and fails on a console error would have caught all three before they were public. That check does not exist because there is currently nowhere to hang it. A pull request is a hook to hang it on.

The Three Checks Worth Having

The trap in adopting CI is asking for coverage. Coverage is a number that can be moved without moving the thing it stands for, and on a static site it is close to meaningless. The better question is the one this repository's own history already answers: what actually broke? Build the checks that would have caught those, and nothing else.

  1. Every page loads with zero console errors. Headless browser, every HTML file, wait for network idle, fail on any error or failed request. This catches the entire path-resolution family, which is most of the site's real bug history.
  2. Every JSON file parses, and every fetch target exists. Half the pages here render from JSON at runtime. A malformed file or a renamed path is a blank page in production and a one-line failure in CI.
  3. Every internal link resolves. Roughly two hundred HTML files with hand-written root-absolute paths. This is the cheapest check on the list and the one most likely to fail.

Three checks, no test framework, and they cover the observed failure modes rather than the imaginable ones. If a fourth is ever added, it should be because something broke that these three missed.

Not Everything Gets a Pull Request

The daily bulletin cron commits generated HTML, a regenerated sitemap, an RSS file, and a set of OG cards, every day, forever. Routing that through review is pure ceremony: nobody will read it, the reviewer becomes a rubber stamp, and a rubber stamp trains you to rubber-stamp the one that mattered. A process that is followed dishonestly is worse than no process, because it produces a record that says the change was reviewed.

Three tiers, by who wrote the change and what it can break:

ChangePathWhy
Generated content Straight to main Bulletins, sitemap, RSS, OG cards. Machine-authored from a template that was already reviewed. Run the three checks after the fact and alert on failure.
Agent-authored, isolated PR, auto-merge on green A new page, a new article, a benchmark data file. Blast radius is one URL. The checks are the reviewer; a human reads it only if the checks fail.
Anything shared PR, human reads it assets/styles.css, the nav stamper, scripts/, the benchmark renderers. One file, every page. This is the tier the whole policy exists for.

Most days that means nothing changes. The value is concentrated entirely in the third row, and the third row is exactly where an agent is most likely to make a change that is locally reasonable and globally wrong.

The Preview Is the Real Prize

Everything above is about catching mistakes. The larger benefit is different and easier to feel: a URL you can open before it is the live site.

On a site that is entirely visual output, a diff is a poor description of a change. Fourteen lines of CSS are unreadable as text and obvious as a rendered page. A preview deployment per pull request turns review from reading a patch into looking at the thing, which is the only review that works on a design system. That is worth more than the checks are.

What This Costs

Honest accounting, because this is where most process proposals lie:

  • Latency. A change goes live in minutes rather than seconds. On generated content that would be a real loss, which is why generated content is exempt.
  • A branch to name. Trivial for an agent, mildly annoying for a person fixing a typo. Keep push-to-main available and use it; the policy is a default, not a gate. Branch protection is the wrong tool here — it converts a good habit into a wall you will eventually route around, and routing around it is a worse habit than the one it replaced.
  • CI that will itself break. Three checks is a maintenance surface. It is small, but it is not zero, and a flaky check that gets ignored is worse than no check, since it teaches you to ignore red.

The proposal is deliberately the smallest thing that addresses the actual change: agents write diffs now, and there is currently no moment between authorship and production where anyone looks.

This Week

  1. Write the three checks as a single workflow. An afternoon, mostly the headless browser loop.
  2. Turn on preview deployments for pull requests.
  3. Set agents to open pull requests by default for anything outside the generated-content paths, and auto-merge on green.
  4. Leave branch protection off. Revisit in a month based on what actually got caught — and if the answer is nothing, delete the process rather than defending it.

Disclosure, in the spirit of the thing: this page was pushed straight to main. It is a new file in /writing/, which by the policy above sits in the middle tier and should have gone through a pull request with a preview URL. The irony is the point. It is exactly the class of change that is fine ninety-nine times and, on the hundredth, is a broken stylesheet path on a live page — noticed by whoever happens to look first.