How we measure
A rewrite is "in Paul Graham's voice" when six independent measurements say so at once. One discriminator is meaningless: style embeddings, LLM judges and stylometrics correlate near zero with each other, so a system that games one fails the others. The site renders a report the pipeline wrote; it never computes a number itself.
Style embedding pending
Jury 2AFC pending
Checklist pending
Meaning preserved pending
AI detection pending
Length and format pending
The evalset
120 inputs, 10 per category: email, memo, blog, tweet, product, essay, announcement, apology, pitch, explainer, review, public domain. 60 to 600 words each, with tweet-length rows where the category allows. The set is frozen; every run rewrites the same inputs.
Rules we hold ourselves to
- The generating model never judges alone. Output from one model family is never judged only by that family; the jury always spans at least two other families.
- Metric values are not compared across prompt versions as if calibrated. We compare pass/fail and position, and re-run the baseline beside any candidate (paired bootstrap, 1000 resamples, 95% CI).
- Holdout essays (2025 and later) are never in prompts, exemplars, references or training data. The style reference centroid is built from non-holdout essays of at least 800 words.
- Judges prefer markdown and penalize plain prose, which is the whole PG style. Both sides of every comparison are stripped of markdown and whitespace-normalized before judgment.
- A human check runs once per promoted version: a blind 2AFC sheet of 12 items plus 3 boundary items where a real PG opening is spliced into a rewrite. Pass is at most 65% human accuracy.
Contamination
Contamination caveat: PG's essays are in every pretraining set. The rewrite task (new input text, not continuation) limits regurgitation credit, and claim precision catches fabricated PG positions.
Detector
The AI-detection family uses the Pangram API document score when the key exists. Until then an open proxy (Binoculars, two small causal language models) calibrated on 50 real PG passages and 50 GPT-style paragraphs stands in, and the report names the proxy explicitly. No report has been published yet.
Promotion
A prompt version or distilled checkpoint is promoted when all six families pass on the full evalset with the baseline arm beside it, the human check passes, and the report is committed with a dated decision entry in the same push.
Published reports
None yet. Each published report is listed here with its run id, prompt version, model and gate result.
Raw report: /metrics/pg.json. Back to the site.