KOEN
All notes
AI & REVIEW

Choosing what to read first does not clear the rest

A Jev wiki experiment raises a practical question: what evidence turns a review ranking into permission to skip a draft?

Two green trays hold documents under a magnifying glass and a stack still awaiting review

Now there are two trays

Imagine a desk covered in drafts awaiting review. Move the lowest-scoring ones to the left tray and you have decided where to start today. You have not learned anything new about the accuracy of the drafts on the right.

That distinction stayed with me while reading the Jev wiki experiment. How far is a useful reading order from a smaller review workload?

In the source’s current-pipeline sample, 4 of 62 pages had at least one critical or high defect. Jev’s ranking had an AUC of 0.491, with a 95% confidence interval of 0.20–0.79. Including older pages changes the picture, so the cohorts matter.

The labels came from AI reviewers, not human ground truth, and the private-wiki measurements are not externally verifiable. This small public experiment cannot settle Jev’s general performance.

I want to follow a narrower question: what work are we proposing to omit because of the score?

What counts as a bad draft?

Consider a hypothetical team reviewing customer notices. An incorrect application deadline can cost someone an opportunity. An awkward sentence deserves editing, but it does not create the same harm.

Put both in a single count of ‘bad drafts found’ and a review policy can look very different. It might find plenty of clumsy prose while missing the incorrect dates.

For this imagined team, I would define the defects that must be caught before running the comparison: wrong deadlines, missing eligibility conditions. Otherwise it is too easy to improve the result afterward by adding easier findings.

Then we need to know how common those defects are in this particular stack. If important errors are almost everywhere, the right tray still needs review. A useful order may remain useful without removing much work.

If the errors are rare, a few clean pages offer little reassurance. We need enough evidence to distinguish an empty sample from a selection method that missed the defects.

Check the drafts left behind

To compare policies in this imaginary team, I would freeze one set of draft versions. Give low-score-first, random order and the existing review order the same time budget.

Record important errors found, reading time and what remains in the unread drafts, alongside the number of pages reviewed. Include the time spent producing scores and labels in the cost.

In particular, independently inspect a sample of high-scoring drafts left at the back. Finding plenty of errors at the front does not establish that the back is clean. Which tray still contains the wrong deadline?

I would also separate notices produced under old writing rules from those made under the current rules. A score might be good at finding the old material without being useful for detecting errors in today’s drafts.

This is a proposed review procedure, not my reproduction of the wiki experiment. Human judgment still has a role. Models can agree because they share a blind spot; the date can remain wrong after that agreement.

Scores make the desk look organized. They can make today’s starting point clearer too.

Before covering the right tray at the end of the day, I would pause over whether its label can honestly say ‘review complete.’

Sources