Evals

How we measure

A rewrite is "in Paul Graham's voice" when six independent measurements say so at once. One discriminator is meaningless: style embeddings, LLM judges and stylometrics correlate near zero with each other, so a system that games one fails the others. The site renders a report the pipeline wrote; it never computes a number itself.

pending No published report yet. The first gate run will land here as JSON; the site never computes a number.

Style embedding pending

n/a
Pass: median position >= 0.5 and >= 90% of rewrites above the cross-author floor
StyleDistance cosine (LUAR as a second read) between each rewrite and the PG reference centroid, chunked to 512 tokens and pooled. Position = (cos - floor) / (ceiling - floor), where the ceiling is held-out PG essays and the floor is other authors.

Jury 2AFC pending

n/a
Pass: mean jury accuracy <= 0.60 (chance is 0.50); never a single judge
Each rewrite is paired with a real PG passage of similar length. Three judges from at least two non-generating model families pick which is real, in both orders. Markdown stripped and whitespace normalized on both sides.

Checklist pending

n/a
Pass: mean >= 8.0 of 10 items
Ten binary items per rewrite, 2 of 3 judges majority, anchored on a real PG paragraph: plain prose, conversational precision, concrete anecdote, counterintuitive turn, no purple prose, no AI-speak, short sentences, argument progression, plain diction, would pass as PG.

Meaning preserved pending

n/a
Pass: NLI >= 0.80, claim recall >= 0.90, claim precision >= 0.90
Bidirectional NLI entailment between input and output, plus atomic claim extraction on the input matched against the output (recall) and the reverse (precision: no invented claims).

AI detection pending

n/a
Pass: >= 95% of rewrites scored human
AI-detection document score. Pangram when the key exists; until then an open proxy (Binoculars) calibrated on real PG and GPT-style passages, reported under its own name.

Length and format pending

n/a
Pass: >= 90% within 25% of input length, 100% format clean
Output length relative to input, plus a format check: no markdown headers, bullets or bold, and no em dashes.

The evalset

120 inputs, 10 per category: email, memo, blog, tweet, product, essay, announcement, apology, pitch, explainer, review, public domain. 60 to 600 words each, with tweet-length rows where the category allows. The set is frozen; every run rewrites the same inputs.

Rules we hold ourselves to

Contamination

Contamination caveat: PG's essays are in every pretraining set. The rewrite task (new input text, not continuation) limits regurgitation credit, and claim precision catches fabricated PG positions.

Detector

The AI-detection family uses the Pangram API document score when the key exists. Until then an open proxy (Binoculars, two small causal language models) calibrated on 50 real PG passages and 50 GPT-style paragraphs stands in, and the report names the proxy explicitly. No report has been published yet.

Promotion

A prompt version or distilled checkpoint is promoted when all six families pass on the full evalset with the baseline arm beside it, the human check passes, and the report is committed with a dated decision entry in the same push.

Published reports

None yet. Each published report is listed here with its run id, prompt version, model and gate result.

Raw report: /metrics/pg.json. Back to the site.