Yazmyne Ortega / verdah / accuracy audit · n=100 · measured August 2026

When it says AI, it has never been wrong. It also misses more than half.

I labelled 100 LinkedIn posts by hand, half written by a model and half by people, then ran them through my own product. Every one of the 21 posts it flagged as AI really was AI, and it never once flagged a person. It also let 28 machine-written posts through. Both numbers are below, with the full matrix, because a percentage you cannot check is worth nothing.

The confusion matrix

Rows are what the post actually was. Columns are what Verdah said.
ActuallySaid AISaid HumanSaid MixedTotal
AI2128150
Human050050

71 right out of 100. Every mistake runs the same direction: it lets AI through, and it leaves people alone. That lopsidedness is on purpose. It's the whole design.

Where it fails

56%of AI posts called human
28machine-written posts missed, out of 50
1AI post landed in Mixed, not flagged

Twenty-eight of fifty AI posts came back as human. If a machine-written post has real names, real numbers, or something unflattering in it, this will miss it. That's the most common way it fails. Read a Human verdict as "nothing jumped out", not as a clean bill of health.

Where it holds

0%false accusation rate
21/21AI calls that were correct
50/50human posts correctly left alone

Not one of the fifty human posts got called AI, and all 21 of its AI calls were right. That isn't proof of a perfect false-positive rate though. It's a hundred posts. The honest reading is that none showed up here, not that none can happen.

What retuning can and cannot buy

The verdict comes from a 0 to 100 score and two cutoffs. I swept every possible pair against this same set. Detection barely moves. There's no setting that catches meaningfully more AI without it starting to accuse people.

Same 100 posts, same model, thresholds swept across the full range.
ThresholdsAI caughtPrecisionFalse accusationsExact
33 / 32 · current21/50 · 42%21/21 · 100%0/5071/100
30 / 29 · more aggressive22/50 · 44%22/23 · 96%1/5071/100
53 / 33 · more cautious17/50 · 34%17/17 · 100%0/5067/100

Two more points of detection buys you your first false accusation. Not worth it. Detection gets better by changing what the model looks for, not by nudging a number, so that work belongs in the next version of the prompt.

What I will not claim

Method