Journal

Research

·

10 min read

Evaluating taste

Evaluating taste

Evaluating taste

Benchmarks measure correctness. We built a panel of poets, editors and support agents to measure something harder.

Lena Moreau

Benchmarks measure correctness. We built a panel of poets, editors and support agents to measure something harder. In this note we walk through the thinking behind the work, the experiments that failed, and the one idea that finally made it click.

Where it started

Every project at Machine Poetry begins with a sentence that sounds wrong. For this one it was: “A model should be allowed to be quiet.” We collected thousands of moments where our systems spoke too soon, too confidently, or too much — and treated each one as a line of poetry that needed editing rather than a bug that needed patching.

What we measured

Accuracy alone told us very little. We tracked cadence, clarity and consequence — how a reply reads out loud, whether a busy person can act on it, and what happens downstream when they do. Small gains in restraint produced large gains in trust.

What comes next

We are opening this work to every customer on the Studio plan next month. If you would like early access, or simply want to argue with us about semicolons, our inbox is always open.

Create a free website with Framer, the website builder loved by startups, designers and agencies.