Blue Pencil

From 31 wrong or unsure checks in every 100 to 4: how we tuned Blue Pencil for our own technical notes

This is how we at Spectacular AI use Blue Pencil internally, on the technical notes that document our own code.

At a glance

We tuned our checks until we could trust them, then used them to fix new drafts.

  • 314

    Checks wrong or unsure: 31 of every 100 at the start, and 4 of every 100 after tuning.

  • 202–3

    On one note: a review of one note runs 63 checks, so at those rates about 20 would be wrong or unsure at the start, and 2 or 3 at the end.

  • 37100

    Rules met by a new draft: 37 of every 100 on a first draft written from a one-line request.

    After one round of fixes: all 63 rules met, on each of three new notes.

  • $90.70$11.74

    Cost of a round of tuning: $90.70 for the first round, and $11.74 for the third.

The challenge

Our first rules for checking our own notes were wrong or unsure on 31 of every 100 checks.

The notes are part of a wiki about our own product. Each data-schema note describes one kind of record our code writes or reads, with every field, its type and the rules its values follow. People and AI coding agents read these notes before they change the code.

We wanted AI agents to write new notes, and we wanted every draft checked against one written definition of a good note before anyone read it. The first rules looked reasonable until we ran them on four notes we knew were right.

What we did

We tuned the rules in three rounds, and an AI agent did the tuning. It followed only the written guide that comes with Blue Pencil. After each round we improved the guide, and a new agent started from scratch.

Each round repeats three steps:

  1. Write the rules for what a good note has.

  2. Check the rules on the four notes we knew were right, and on copies of those notes with one planted mistake each.

  3. Rewrite the rules that came back wrong or unsure, and check again.

check again

The planted mistakes told us which rule should fail on each copy, so every check had a known right answer.

We got stuck at 88 of every 100 checks right and clear, from the end of round 1 to the start of round 3. The last rules were too vague to give the same answer twice. Once the agent rewrote them precisely, we reached 96.

Results

The checks became reliable. Over three rounds and eleven passes, the checks that were right and clear went from 69 of every 100 to 96 of every 100.

Checks wrong or unsure
31 of every 100 at the start
4 of every 100 after tuning
The checks became reliable.Over three rounds and eleven passes, the checks that were right and clear went from 69 of every 100 to 96 of every 100.60708090100round 1round 2round 3stuck at 88 of every 10068.777.483.787.687.988.28887.88896.9966996

New drafts met the rules after one round of fixes. We gave an agent three short fact sheets for new notes. Asked in one line to write each note, its first drafts met 37 of every 100 rules: each one left out 5 to 7 whole sections. After one round of automatic fixes, guided by Blue Pencil's checks, all three drafts met every rule.

Rules met by a new draftof every 100
Rules met by a new draftAsked in one line to write each note, its first drafts met 37 of every 100 rules: each one left out 5 to 7 whole sections. After one round of automatic fixes, guided by Blue Pencil's checks, all three drafts met every rule.first draft37After one round of fixes100

A detailed brief helped, and the checks still caught a gap. With a detailed written brief instead, the first drafts met 98 of every 100 rules. The one rule each draft missed was one the brief never stated.

"We use Blue Pencil internally on the notes that document our own code. Our first rules looked fine, and 31 of every 100 checks came back wrong or unsure. Three rounds later the checks were right and clear on 96 of every 100, and an agent's drafts from a one-line request met every rule after one round of fixes."
[Name], [job title], Spectacular AI

What it took

Tuning took three rounds. The first round cost $90.70 of AI usage and the third cost $11.74, because we improved Blue Pencil's tools and its guide between rounds.

Some rules still get an unsure answer, such as a rule about every row of a table or the exact form of a link. Blue Pencil's guide lists those rules, so you can leave them out or check them yourself.

These results cover one kind of document and three new notes. An AI agent fixed the drafts, and the same checker that guided the fixes scored them. So the results show that the drafts meet our definition of a good note, not that people prefer them.