Evidence — read the label first
The benchmark we ran ourselves, published with its limits.
Method
One variable separates the three arms
One generated corpus of 10,000 documents with known ground truth, about 1 in 10 of them a scan. One set of 148 questions across the five things people ask of a document set. Three arms — Citenda, the same AI model over the same files in a folder, and the same model with no corpus at all — with the same hardware, the same budgets and the same grader. 444 runs in total; every one returned an answer. The only thing that differs between arms is what we sell.
Results
What Citenda answered — and what the same model answered without it
- On the 10,000-document generated corpus, Citenda answered 93.2% of the 148 questions at the 10-minute budget — the same AI model, given the same files in a folder, answered 56.8%.
- With no corpus at all, the same model answered 17.1% across all 148 questions of that run — the control that shows the questions are not answerable from general knowledge.
- Inside a 60-second budget on the same generated corpus, Citenda was correct and finished on 81.1% of questions; the raw-folder arm finished 11.5%. Median time to answer across all 444 runs: 28 seconds.
- On counting and totalling across that whole generated corpus, Citenda's margin over the raw-folder arm was +68 points; on confirming something is absent, +76 points.
- 99.3% of Citenda's answers on that run resolve to a stable citation id you can reopen months later.
Reading rules
These figures travel under rules
- The corpus is named in the sentence, every time. At a different scale, or on a corpus with different properties, these numbers would be different — a figure that omits its corpus is false in one direction or the other.
- The tripwire is non-removable. It rides above the fold here and travels with every quotation of this run, anywhere.
- Citenda is the subject; the arms are contrast. We publish what Citenda answered alongside what the same model answered without it. We do not benchmark or review anyone else's product.
- The control is quoted in its generous form. The no-corpus arm passes some questions by refusing to answer; we quote the all-questions figure rather than the harsher variant that excludes those refusal-passes.
- No figure on this page enters a quote or a contract. A benchmark number is not a commitment about your archive.
Limits
The limits are part of the result
- Synthetic corpus. We generated it; that is what makes ground truth checkable, and what keeps these from being client numbers.
- One domain, one question set, one grader.
- We ran it ourselves. No third party has repeated it. We publish that fact rather than soften it.
- No figure here is a claim about your archive. Accuracy on a real client corpus is unmeasured until a real engagement measures it — which is exactly what the Estate Assessment exists to start.
Publishing these limits is the point, not a concession: a number you can check, with the label that qualifies it, is the product working on its own marketing.