the reader is live · scoring its own pull requests
Doug reads every pull request once, routes the few that need a human, and clears the rest. Every merge starts a clock against this repository’s own reverts. When Doug is wrong, it says so — in public.
#84 · Restore the WorkOS sign-in front door
Flagged · risk 0.58 · diff read
Needs you. Risk is above this repository's flag line, so Doug is asking for a human read. It does not block: this check is neutral and the merge button is unchanged.
Findings
adjudicated 20 · pending 114 · as of 2026-09-03
deep reads 43/200 this cycle
The open queue, pinned by risk
65% cleared without a human
One pull request, start to finish
Everything to the left of the merge is what a reviewer does. Everything to the right is what no reviewer does: keep the verdict, and find out whether it was right.
open
Or a push lands. The GitHub App takes the webhook and queues one job. Nothing else in your pipeline changes.
read
Title, files, patch — capped at 100k characters in a fixed order, and the check says when the cut fell short. No author, no dates.
route
Scored against the flag line you set for that repository. One neutral check, one sticky comment. Never a red X.
d0
The verdict becomes a dated row that nobody can edit — threshold, findings, and what the reader was shown, pinned.
d14 · d60
At each window the row is adjudicated: reverted, or survived. The scoreboard counts it either way.
publish
The miss rate publishes on its pre-committed date with its N, whatever it says. Until then it renders as a dash.
Three rules, in writing
Doug orders attention. It holds no merge hostage, gates no pipeline, and adds zero seconds to a cleared PR.
A reviewer that also writes is marking its own homework. Doug decides where eyes go — it never generates a fix.
Every escaped defect Doug cleared will be counted, dated, and published on the locked cadence. If the number is bad, you'll see it here first.
The cost of reviewing everything
Coding agents multiplied pull requests. Tools that answer with a full model review of each one — /code-review on every branch, a bot on every diff — scale their spend with exactly that number, and still leave a person reading comments on all of them. Doug spends one bounded read per PR, then spends your attention only above your flag line.
| Per pull request | A model review of everything | Doug |
|---|---|---|
| What a human reads | The comments, on 100% of PRs. | The PRs above your flag line. On this repository today: 59 of 167. |
| Model spend per PR | A full agentic read of the branch, as large as the branch is. Run it twice, pay twice. | One read of the diff, capped at 100k characters at a fixed effort. The look-up passes that follow run a cheaper model. |
| Where the spend shows | On a bill, later. | On the check run itself: “deep reads 143/200 this cycle.” The meter is the surface you already read. |
| When it runs | When someone remembers to run it. | On every push, as a GitHub check. Nobody has to remember, and nobody can forget. |
| The merge button | Whatever the tool decides that day. | Untouched. The check is neutral, every time, by design. |
| When the read fails | You re-run it, or it ships unread. | The deterministic tier scores it without a model and the check says so. A downgrade is never silent. |
| After the merge | Nothing. The comments were the product. | A 14- and 60-day clock, graded against this repository's own reverts, published on a date. |
Doug is not model-free. With the reader on, every PR costs a read; the deterministic tier is what runs when it is off or fails. The saving is bounded spend and routed attention, not a skipped model call.
What the reader is given
What the reader is not told
The judgment about the code is made without knowing who wrote it. That claim is narrow on purpose: it covers the read, not the whole of Doug.
Doug does see authorship elsewhere. The deterministic fallback — used when a read fails — scores a PR higher when a bot opened it, and the queue tells you who wrote each one, because you need that to route. What it never does is let the reader grade the code against the author’s reputation.
What’s actually measured
0.69/ 0.67
Ranking AUC on sentry and grafana, pre-registered before a single model call. The best deterministic baseline scored 0.59 and 0.52 — on grafana every metadata method we tried lands at or below random. Reading the diff is the first thing that survived a second repo.
That’s the 30,000-character probe reader, not the one running on your PRs — the shipped reader hasn’t been measured by it.
Published miss rate
—
Not yet. The live counters — adjudicated, pending, first due — sit on the scoreboard. The number lands here with a date next to it, good or bad.
Others learn what reviewers say
Every merge starts a clock. At 14 and 60 days the verdict is graded against what actually happened — reverted, or survived the window. The scoreboard starts at zero and says so.
Every verdict is a ledger row — dated, immutable, waiting to be graded. A repo running Doug for a year holds a calibrated risk record of itself that no point-in-time reviewer can replicate.
The graded history, served to coding agents before they type: “this migration shape reverted here 7 of 9 times — the two that survived used dual-write.” Ships when there is adjudicated data to serve, not before.
167 scored PRs, 59 worth your time, and the receipts behind every score.