Accuracy and known limits
Use the recorded benchmark to plan examiner review, without treating historical field scores as proof of a new deal’s ownership.
AI Title Examiner is a review assistant. The benchmark uses project-maintained answer keys. It is not independent landman approval, and it does not provide an answer key for a new live deal.
The recorded benchmark
These figures are the repository’s results refreshed on September 3, 2026, recorded in results/benchmark_report.md. They are not a new evaluation performed for this guide.
The table reports unweighted means across 17 full surveys, including older handwritten records. Each survey contributes equally to the mean, regardless of document count. Some field scoring uses a model-based judge and can vary between evaluations.
| Metric | Recorded mean | Survey range |
|---|---|---|
| Lease Run Sheet coverage | 96.7% | 88.0–100% |
| Book type | 81.4% | 54.5–100% |
| Instrument type | 73.9% | 42.9–97.1% |
| Instrument date | 85.0% | 63.6–96.6% |
| File date | 81.8% | 57.9–100% |
| Grantor | 86.8% | 81.0–97.4% |
| Grantee | 84.6% | 68.9–100% |
| Acreage | 69.8% | 26.9–96.6% |
| Conveyance type | 74.9% | 31.4–100% |
Lease Run Sheet coverage measures representation of keyed instruments. It does not mean that the represented rows have correct ownership conclusions.
Of the 14 surveys with a Letter of Opinion on Title answer key, ownership fractions were exactly correct on 6, partially correct on 3, and failed on 5. Three of the 17 surveys had no ownership key. Do not describe the result as 6 of 17, or count partial cases as exact matches.
The survey range shows why one average cannot predict your case. Modern clean scans and nineteenth-century handwritten packets present different reading problems. The repository has no paired evidence supporting a quantitative accuracy comparison with Perplexity.
Grounding is a different measurement
The repository records approximately 98.0% name grounding in an August 29, 2026 cached-extraction sweep: 98 run directories, 3,862 document instances, and 13,742 names.
Grounding measures whether an extracted name occurs in that document’s OCR transcript, which is the text read from its scan. It does not measure correct ownership, legal interpretation, or complete source coverage.
The sweep re-read available source PDFs and scored cached extracted values. It excluded 1,022 cached extractions whose source files could not be located. For 81 PDFs that could not be re-transcribed, it retained stored text. Repeated runs can include the same underlying PDF, so document instances are not unique documents.
A name can be grounded in an incorrectly transcribed word. A fully grounded record set can still omit the deed that changes the owner. In a current job, read the faithfulness coverage warning and ungrounded entries alongside the source pages.
What needs close review
| Review area | Question to resolve with source evidence |
|---|---|
| Subject-tract acreage | Does the figure describe the tract under review or every tract in the deed? |
| Fraction basis | Does the clause transfer part of the whole tract or part of the grantor’s existing interest? |
| Probate and reservations | Do the heir shares, community-property interests, life estates, and reserved interests support the proposed distribution? |
| Dates | Does the file date use the appropriate filed or recorded stamp? |
| Name continuity | Do aliases, initials, suffixes, and spouse references identify the same parties? |
| Missing records | Is a necessary incoming or outgoing conveyance absent from the packet? |
| Lease status | Is there separate evidence of expiration, production, or held-by-production status? There is no Texas Railroad Commission integration. |
| Liens and releases | What county database research or additional records are needed? |
A correct arithmetic total does not resolve these questions. Use Read and review the results to create an evidence-backed worklist.
Leasing and mineral buying
A leasing review can use a structured first pass to organize examiner work. Mineral buying depends directly on reliable present interests and tract acreage. The recorded ownership performance does not support closing a mineral purchase without independent review.
The deterministic ownership calculation makes supported arithmetic exact. It cannot supply a missing instrument or correct a mistaken interpretation of an operative clause. Keep arithmetic verification separate from examiner approval.
Cost and duration
The repository gives approximately $150 per standard full run as planning guidance. Historical live sourcing work reached roughly $300–450 per deal when repeated ownership recomputations increased cost. These figures are estimates, not a quote or spending cap.
Document count, scan quality, quality-review work, source availability, and recomputation affect cost and duration. There is no verified completion-time guarantee for a new case.
The request’s budget tracks estimated document-sourcing spend. Model inference, browser sessions, search calls, and rejected paid deliveries can incur costs beyond that estimate. Agree on the case and operational spend plan with the owner before starting a full examination.
Evaluate improvements honestly
For a comparable evaluation, you need the same case, source packet, expected outputs, and saved artifacts from both runs.
- Record the source set and configuration for each run. Identify any changed inputs before comparing results.
- Compare fields and ownership against the appropriate answer key or a qualified reviewer’s source-backed findings.
- Separate reading errors, missing records, scorer vocabulary, and ownership-reasoning errors.
- Repeat comparisons that depend on model-generated outputs. A single run is one observation, not a reliability estimate.
- Save reports, review labels, and denominators with the findings so another reviewer can check the claim.
A useful improvement report states the measured task, case count, exact versus partial results, and unresolved limitations. Do not convert synthetic arithmetic success or name grounding into a general title-accuracy claim.
To collect reviewer evidence, follow Pilot testing and feedback.