Evidence — read the label first

The benchmark we ran ourselves, published with its limits.

Method

One variable separates the three arms

One generated corpus of 10,000 documents with known ground truth, about 1 in 10 of them a scan. One set of 148 questions across the five things people ask of a document set. Three arms — Citenda, the same AI model over the same files in a folder, and the same model with no corpus at all — with the same hardware, the same budgets and the same grader. 444 runs in total; every one returned an answer. The only thing that differs between arms is what we sell.

Results

What Citenda answered — and what the same model answered without it

Share of questions answered correctly, by time budget On the 10,000-document generated corpus, the share of the 148 questions answered correctly and finished rises with the time budget. Citenda reaches 81.1% inside a 60-second budget and 93.2% at 10 minutes. The same model on the raw files in a folder reaches 11.5% at 60 seconds and 56.8% at 10 minutes. The no-corpus control stays flat at 17.1%. The time axis is logarithmic. 100% 75% 50% 25% 0% 10s30s60s2min5min10min 60-second budget 81.1% 11.5% Citenda engine 93.2% at 10 min Raw folder 56.8% at 10 min No corpus 17.1% — flat
Share of all 148 questions answered correctly and finished within each time budget, on the 10,000-document generated corpus. Time axis is logarithmic. Synthetic run — we generated the corpus; not a client accuracy claim.

Reading rules

These figures travel under rules

Limits

The limits are part of the result

Publishing these limits is the point, not a concession: a number you can check, with the label that qualifies it, is the product working on its own marketing.