Construction litigation

Every document in every case. One memory the firm can question.

Hundreds of thousands of contracts, letters, expert reports and drawings, read once, linked together, and checked against the firm’s own quality bar. Every answer cited to the page. Built in your environment, owned by your team.

texo-law/Riverside Tower v. Halden/Ask prod · us-west-2⌘KJA
When did the owner first receive notice of the slab deflection, and what did the structural engineer say about it?⌘⏎
Answer · verified 3 citations · 2.2s · 1 sub-agent

The owner's representative first received written notice of the slab deflection on 14 March 2023, in RFI-212 from Halden's superintendent1. The structural engineer of record replied on 21 March 2023 that the measured deflection at gridline C-7 was "within L/360 but outside the tolerance stated in S-501 note 4" and recommended a survey before placing topping2. No document in the matter records a response from the owner before Change Order 31 on 2 May 20233.

citation-exists 3/3 quotation-match 3/3 matter 24-118 only
Abstained "Did the owner waive the liquidated-damages clause?"

No document in matter 24-118 supports an answer. 2 passages mention liquidated damages without waiver language. Show every reference after 1 Jan 2023 →

Building chronology from 6 cited events…

The problem

Three reasons this is harder than “put it in a chatbot.”

Legal AI tools are good at reading text and most now work inside your matters. None of them read construction drawings as drawings, and none of them check every quotation against the page or show you how often each check is right.

Drawings

A drawing is not a document

When an AI turns a sheet into text, the title block, the revision clouds and the callouts that point to other sheets are lost. The best systems today read the text on a sheet well and the graphics badly. The drawings need their own pipeline.

Scale

Finding is harder than reading

At hundreds of thousands of documents the model never sees most of them. What it answers depends on what the search brought back, so the search has to be exact on clause numbers and sheet IDs and good at meaning at the same time, and it has to stay inside one matter.

Proof

The answer has to survive a courtroom

Courts have sanctioned lawyers for AI-invented citations even at firms with written AI policies. Every answer needs its sources attached and checked, and every flag needs a record of who looked at it and what they decided.

The product

Three screens the firm will live in.

Ask a matter a question, work the review queue, and read a drawing the way the system reads it. The agent’s work is visible on every screen, and every matter is walled from every other.

Ask

Ask the matter a question. Get an answer you can file.

Every sentence carries a citation to the page and paragraph it came from, and a check that the quote really is on that page. The work the system did to get there sits beside the answer. When nothing supports an answer, it says so.

texo-law/Riverside Tower v. Halden/Ask prod · us-west-2⌘KJA
When did the owner first receive notice of the slab deflection, and what did the structural engineer say about it?⌘⏎
Answer · verified 3 citations · 2.2s · 1 sub-agent

The owner's representative first received written notice of the slab deflection on 14 March 2023, in RFI-212 from Halden's superintendent1. The structural engineer of record replied on 21 March 2023 that the measured deflection at gridline C-7 was "within L/360 but outside the tolerance stated in S-501 note 4" and recommended a survey before placing topping2. No document in the matter records a response from the owner before Change Order 31 on 2 May 20233.

citation-exists 3/3 quotation-match 3/3 matter 24-118 only
Abstained "Did the owner waive the liquidated-damages clause?"

No document in matter 24-118 supports an answer. 2 passages mention liquidated damages without waiver language. Show every reference after 1 Jan 2023 →

Building chronology from 6 cited events…
Ask: every sentence cited and verified; the agent's work shown beside the answer

Review queue

Candidates, not verdicts. A person has the last word.

Each run shows which checks ran, what they found, and the evidence for every candidate. Reviewers accept or dismiss from the keyboard; precision and recall sit on the screen with confidence intervals, so the firm can show a court how the bar was set.

texo-law/Riverside Tower v. Halden/Review queue prod · us-west-2⌘KJA
Flag run #412 started 00:41 · 318 sheets · 2,140 documents
Open 23Review 4Accepted 118Dismissed 9Severity ↓
DocumentCandidateDetectorConfStatus
A-401 CCallout 7/A-401 has no detailrule1.00open12ms
S-201 ↔ S-501Schedule S2 ≠ callout 4/S-501 at C-7judge0.91open3.1s
03 30 00Spec ACI 318-14; drawings 318-19rule1.00open9ms
RFI-212Body 14 Mar; RFI log 16 Marreconcile0.84review1.4s
M-301Title block rev blank; index Brule1.00accepted8ms
Log 04-02OCR 0.62; verify figuresrule1.00dismissed6ms
Evidence · S-201 ↔ S-501spec_drawing_judge@v7 · prompt 3e1f…a9 · rubric conflict.v3
S-201 schedule S2#5 @ 12" o.c. T&B, note 4
4/S-501 callout#5 @ 8" o.c. TOP · 1½" cover
Accept ADismiss DOpen both sheets
Review queue: candidates with evidence, detector version and disposition; precision published with intervals

Drawings

Sheets read as sheets, not as broken text.

A framing plan is split into its plan, callouts, schedule, notes, revision table and title block, read at full resolution, and linked to the details, specs, RFIs and change orders that cite it. The trace shows each step the index took.

texo-law/Riverside Tower v. Halden/Drawings/S-201 prod · us-west-2⌘KJA
S-201 Level 4 framing plan · Rev C · 2023-02-10
SheetZones 8Fields 41Refs 14Flags 2
Zone · callout 4/S-501
Refers to
4/S-501 · slab edge @ C-7
Schedule
S2 · #5 @ 12" o.c. T&B
Note
5 · coord. w/ RFI-212
Confidence
0.93
Reference graph
S-201→4/S-501→RFI-212→CO-031
4/S-501→RFI-230 · Subm. 07
Flags
S2 ≠ callout 4/S-501 at C-7 judge 0.91
Note 2 cover vs 03 30 00 rule
Meridian Structural · Riverside Tower · 1400 River St1/8" = 1'-0"·zoom 42%·revisions A B C
Drawing viewer: a framing plan as the index read it, with zones, extracted fields and the graph to RFIs and change orders

How it works

How it works, top to bottom.

Files come in from wherever they live, get sorted and checked, are read as text or as drawings, and land in one memory walled off by matter. From there two things happen: lawyers ask questions and get answers with their sources, and the system keeps a running list of what fails the firm’s standards for a person to review.

WHERE THIS SYSTEMIS DIFFERENTYour files, wherever they areDocument system, email, project platform exports, drawing sets, expert filesSorted and checked on the way inDuplicates removed, email threads kept together, versions put in order, tagged to the right matterUnreadable scans caught here, not laterRead as textContracts, letters, reports and emails read page by pageEach passage keeps a note of where it came fromRead as drawingsSheets are looked at, not flattened to textTitle block, schedules, revisions read at full sizeSearchable by words and by meaningExact matches and matches by meaningFiltered by matter, type, date, partyEvery sheet linked to what it referencesSearchable by what a sheet showsSheet → detail → spec section → RFI → change orderOne case memory, in the firm’s own cloud accountEach matter walled off from every other; closed matters kept cold and cheapAsk a questionBreaks the question into searchesOpens documents, keeps digging within a budgetAnswers with the source behind every sentenceFind what fails the quality barCertain checks first: missing sheets, broken refsThen judgment calls, with the evidence quotedThen conflicts between documentsAnswer with its sourcesCitations checked to exist; says "not found" if soA review list for a personRanked by severity and certainty; a lawyer decidesA record of everythingWhat was searched, what was shown, who decided what, and how accurate it has been

The piece most people have not seen before is the right-hand column: drawings are read as pictures, zoned and cropped so the small text is legible, and every sheet is linked to the details, spec sections, RFIs and change orders it references.

The agent

It works a question the way an associate would.

Plans, searches several ways at once, reads the strongest documents in full, follows pointers into the drawings, builds the timeline, then checks every citation before it answers. Each run has a budget and signs the audit log.

  1. It plans
  2. It searches several ways at once
  3. It reads, not skims
  4. It walks the drawings
  5. It builds the timeline and writes the answer
  6. It checks itself and signs the log

The quality bar

What it flags, and how sure it is.

Most flags come from checks that are simply true or false once the documents are read properly. The judgment calls are kept separate, always show their evidence, and always go to a person.

What it catchesHowWho decides
Unreadable or blank scans, duplicates, superseded versions filed as currentChecked on the way inAutomatic, logged
Missing or incomplete title blocks; sheet index that does not match the sheetsRead the title blocks, compare the listsAutomatic, logged
A detail callout that points at a detail or sheet that does not existFollow every referenceAutomatic, logged
Revision letters and dates out of sequence across a set; changes outside revision cloudsCompare revisions, overlay sheetsReviewer confirms
Spec says one thing, the drawing says anotherThe system quotes both sidesReviewer decides; often the contract says which governs
Defined terms never defined, broken cross-references, numbers that disagree within a contractRules, like a proofreaderAutomatic, logged
Dates, amounts or parties that disagree across documentsFacts pulled from every document and comparedReviewer decides
A citation that does not exist, or does not say what the brief says it saysLooked up in the citation database; the passage comparedReviewer confirms before filing

Questions firms ask

The questions a managing partner asks first.

Will this hold up if opposing counsel asks how it was done?

Yes, because it is validated the same way courts already accept for technology-assisted review: a sample is checked by people, and the system reports how much it found and how often it was right, with confidence intervals. Every answer and every flag keeps a record of which documents it used, which version of the rules ran, and who decided.

Does our data leave our control?

No. The system runs inside the firm’s own cloud account if the firm prefers, with models reached over private connections that do not retain data. Nothing trains a shared model. Each matter is walled off in the database itself, so a user on one matter cannot see another, and that wall is tested on every release.

Can it actually read construction drawings?

That is the part most tools get wrong, and the part we built differently. Each sheet is split into its details, schedules and title block, read at full resolution, and linked to the specs, RFIs and change orders that reference it. A benchmark of real drawing sets shows that generic text extraction fails on most of them; purpose-built parsing recovers most of the loss.

What does it get wrong?

It will miss things no document records, and it will sometimes raise a candidate that a reviewer dismisses. The design accepts that: it says “no document supports an answer” rather than guessing, judgment calls always go to a person, and the review queue shows the system’s precision so the firm knows how much to trust each kind of check.

The 90-day plan

Ninety days to a working system on real matters.

Three months, one fee, measured gates. We start on two live matters, add drawings in month two, and turn on the quality checks in month three. Each month ends with a test the firm can see, not a demo.

01Days 1–30
Read the text of two matters
  • Weeks 1–2: bake-off on 500 of the firm’s own pages and 200 sheets to choose the readers
  • Documents pulled from the document system, de-duplicated, threaded, sorted by matter
  • Ask-a-question over two live matters, every answer with sources
  • Matter walls and the audit record from day one
Gate to pass200 questions written by the firm’s associates, answered with sources at an agreed accuracy floor, and a matter-wall test that fails to cross.
02Days 31–60
Read the drawings
  • Every sheet zoned and its title block, schedules and revisions read
  • Sheets linked to the details, spec sections, RFIs and change orders they reference
  • Drawing questions answered with sheet-level sources
  • First automatic checks: missing sheets, broken references, revision order, incomplete title blocks
Gate to passTitle-block fields and sheet-index checks scored against a labelled set; the automatic checks run with near-zero false alarms.
03Days 61–90
Turn on the quality bar and hand over
  • Judgment-call flags: spec-vs-drawing conflicts, contradictions across documents, citation checks
  • Review queue in the lawyers’ workflow (Word add-in or document-system panel, the firm’s choice)
  • Validation report with recall and precision per check
  • Runbook, training, and the remaining matters queued in priority order
Gate to passReviewer agreement with the system’s judgment calls at the target on a held-out set; the firm’s team runs the nightly job without us.
$50,000 / monthOne fee for the Texo team across all three phases. No licenses, no per-seat pricing.
$150,000Total for the ninety days, invoiced monthly, each month gated by the test above.
At costThe fee covers the Texo team. Cloud, model and parsing costs run in the firm’s own accounts and are passed through at cost; for the first ninety days on two matters they are small next to the fee.

What the firm has on day 91.

  • Two matters fully readable by question, with sources, and the rest queued
  • Every drawing in those matters indexed and linked
  • A nightly job that reads new documents and keeps the review list current
  • A written validation report the firm can show a court or a client
  • The code, the infrastructure and the runbook, owned by the firm
  • A team trained to run it, and Texo on call as it grows

Next step

Bring one matter. We will run the bake-off on its documents and tell you honestly what the system will and will not catch.

Schedule a conversationjack@texoadvisors.com

Construction litigation

Matter-scoped retrieval, a vision-first drawing index, and a validated flagging engine.

Hybrid retrieval with an agentic query loop and sentence-level citations for text; a separate zone-detect, extract, graph and page-image index for drawings; tiered flagging validated the way e-discovery validates TAR. Runs in the firm’s cloud account behind private model endpoints.

texo-law/Riverside Tower v. Halden/Ask prod · us-west-2⌘KJA
When did the owner first receive notice of the slab deflection, and what did the structural engineer say about it?⌘⏎
Answer · verified 3 citations · 2.2s · 1 sub-agent

The owner's representative first received written notice of the slab deflection on 14 March 2023, in RFI-212 from Halden's superintendent1. The structural engineer of record replied on 21 March 2023 that the measured deflection at gridline C-7 was "within L/360 but outside the tolerance stated in S-501 note 4" and recommended a survey before placing topping2. No document in the matter records a response from the owner before Change Order 31 on 2 May 20233.

citation-exists 3/3 quotation-match 3/3 matter 24-118 only
Abstained "Did the owner waive the liquidated-damages clause?"

No document in matter 24-118 supports an answer. 2 passages mention liquidated damages without waiver language. Show every reference after 1 Jan 2023 →

Building chronology from 6 cited events…

The problem

The three hard parts, and why each needs its own design.

Text retrieval is a known pattern. Validated review exists in e-discovery. Nothing in the legal stack reads drawings as drawings. The design treats each as a separate problem with its own index, its own verifier and its own validation.

Drawings

Text-flattening is the dominant failure

Generic agents fall back to text extraction on drawing sets and lose the title block, revision clouds and callouts. Frontier VLMs read sheet text well and count symbols badly. The fix is structural: zone the sheet, read crops at full resolution, and build the reference graph before retrieval.

Retrieval

Partition by matter or pay in latency

A matter is tens of thousands of documents inside a firm-wide corpus of millions of chunks. Matter-partitioned indexes are both the security wall and the latency fix. Contextual chunk prefixes, hybrid BM25 + dense search and a reranker remove most retrieval failures before the model sees anything.

Proof

GenAI buys recall with precision

Independent studies of AI-assisted review show higher recall than TAR at lower precision, and commercial research tools still misground citations. The design assumes human false-positive triage, deterministic citation checks after generation, and TAR-style validation with confidence intervals.

The product

Three surfaces over one tool layer.

Every screen is a client of the same API the agents use: ask, review and drawing views call the MCP tools (search_text, search_drawings, verify_citation, review_dispose) with a matter-scoped context, and show the run trace that the audit log records.

Ask

Agentic retrieval with deterministic verification.

Orchestrator plans, fans out to research and drawing sub-agents under a budget, and returns citation spans; verify_citation runs string-match checks after generation; abstention is a first-class output. The run trace is the audit record. Same call is an MCP tool and a REST endpoint.

texo-law/Riverside Tower v. Halden/Ask prod · us-west-2⌘KJA
When did the owner first receive notice of the slab deflection, and what did the structural engineer say about it?⌘⏎
Answer · verified 3 citations · 2.2s · 1 sub-agent

The owner's representative first received written notice of the slab deflection on 14 March 2023, in RFI-212 from Halden's superintendent1. The structural engineer of record replied on 21 March 2023 that the measured deflection at gridline C-7 was "within L/360 but outside the tolerance stated in S-501 note 4" and recommended a survey before placing topping2. No document in the matter records a response from the owner before Change Order 31 on 2 May 20233.

citation-exists 3/3 quotation-match 3/3 matter 24-118 only
Abstained "Did the owner waive the liquidated-damages clause?"

No document in matter 24-118 supports an answer. 2 passages mention liquidated damages without waiver language. Show every reference after 1 Jan 2023 →

Building chronology from 6 cited events…
Ask: every sentence cited and verified; the agent's work shown beside the answer

Review queue

Three-tier flagging with calibrated judges.

Tier 1 rules at precision ≈ 1.0, Tier 2 rubric judges with Trust-or-Escalate thresholds, Tier 3 cross-document reconciliation. Each flag carries evidence spans, detector@version, prompt hash, dedupe key and disposition, written to an insert-only audit log.

texo-law/Riverside Tower v. Halden/Review queue prod · us-west-2⌘KJA
Flag run #412 started 00:41 · 318 sheets · 2,140 documents
Open 23Review 4Accepted 118Dismissed 9Severity ↓
DocumentCandidateDetectorConfStatus
A-401 CCallout 7/A-401 has no detailrule1.00open12ms
S-201 ↔ S-501Schedule S2 ≠ callout 4/S-501 at C-7judge0.91open3.1s
03 30 00Spec ACI 318-14; drawings 318-19rule1.00open9ms
RFI-212Body 14 Mar; RFI log 16 Marreconcile0.84review1.4s
M-301Title block rev blank; index Brule1.00accepted8ms
Log 04-02OCR 0.62; verify figuresrule1.00dismissed6ms
Evidence · S-201 ↔ S-501spec_drawing_judge@v7 · prompt 3e1f…a9 · rubric conflict.v3
S-201 schedule S2#5 @ 12" o.c. T&B, note 4
4/S-501 callout#5 @ 8" o.c. TOP · 1½" cover
Accept ADismiss DOpen both sheets
Review queue: candidates with evidence, detector version and disposition; precision published with intervals

Drawings

Rasterize → zone → crop → extract → link → embed.

RF-DETR zones each sheet; crops go to a VLM for typed extraction; sheet, detail, spec, RFI and change-order references become graph edges; ColQwen-class page embeddings, binary quantized, handle "find the sheet that looks like this."

texo-law/Riverside Tower v. Halden/Drawings/S-201 prod · us-west-2⌘KJA
S-201 Level 4 framing plan · Rev C · 2023-02-10
SheetZones 8Fields 41Refs 14Flags 2
Zone · callout 4/S-501
Refers to
4/S-501 · slab edge @ C-7
Schedule
S2 · #5 @ 12" o.c. T&B
Note
5 · coord. w/ RFI-212
Confidence
0.93
Reference graph
S-201→4/S-501→RFI-212→CO-031
4/S-501→RFI-230 · Subm. 07
Flags
S2 ≠ callout 4/S-501 at C-7 judge 0.91
Note 2 cover vs 03 30 00 rule
Meridian Structural · Riverside Tower · 1400 River St1/8" = 1'-0"·zoom 42%·revisions A B C
Drawing viewer: a framing plan as the index read it, with zones, extracted fields and the graph to RFIs and change orders

How it works

Reference architecture.

Two parse paths and two indexes behind one router, a store partitioned by matter with row-level security, and two consumers: an agentic query loop with sentence-level citations and a deterministic verifier, and a tiered flagging engine feeding a human review queue. Everything both consumers do is written to an audit log the evaluation program samples from.

DIFFERENTIATOR:VISION-FIRST PATHSourcesDMS, email archive, Procore or Aconex exports, drawing sets, expert filesIngest and triageExact and MinHash near-dup, inclusive email threading, version lineage, matter taggingLegibility gate: OCR confidence and image quality score; route by document classText parsing and enrichmentDocling, Textract or Azure DI; stronger OCR on failuresStructure-aware chunks with contextual prefixesDrawing parsing (vision first)Rasterize; detect title block, legend, schedulesFields read at full resolution; VLM sheet captionsText indexBM25 plus dense embeddings, reranker on topFilters: matter, document type, date, partyDrawing index and reference graphPage-image late-interaction embeddings, pooledSheet → detail → spec section → RFI → change orderMatter-scoped store in the firm’s own cloud accountPostgres row-level security by matter and user; object storage for page images; cold tier for closed mattersQuery: agentic retrieval and reasoningDecompose, parallel hybrid sub-searches, rerankOpen documents and iterate within a budgetPrivate-endpoint LLM; sentence-level citationsFlagging engineTier 1: deterministic set-level rulesTier 2: rubric LLM judge, calibrated abstentionTier 3: cross-document fact reconciliationCited answer plus verifier passCitation-exists check; abstains when retrieval is emptyReview queueRouted by severity and confidence; reviewer decidesAudit and evaluationRetrieval sets, flags, model and prompt hash, reviewer outcomes; sampled recall and precision with 95% CIs

Parsing is routed by document class (Docling, Amazon Textract or Azure Document Intelligence Layout for clean PDFs and schedules; Mistral OCR 3 or Reducto for scans) with an OCR-confidence gate. Text chunks are paragraph groups with contextual prefixes and parent-child retrieval. The drawing path rasterizes, runs a layout detector (title block, legend, schedules, revision table, viewports; published mAP50 0.95), extracts fields from full-resolution crops, builds a reference graph from deterministic IDs, and indexes page images with a late-interaction visual embedding, token-pooled. Retrieval fuses BM25, dense and visual results and reranks the top 50–100.

The agent

The agentic loop, in six steps.

Classify and budget; decompose into parallel hybrid searches; open, extract and iterate through tools; traverse the reference graph for drawings; build the chronology and draft with cited spans; verify deterministically, abstain on empty retrieval, log the run.

  1. Classify and budget
  2. Decompose and search in parallel
  3. Open, extract, iterate
  4. Graph traversal for drawings
  5. Chronology and drafting
  6. Verify, abstain, log

The quality bar

Flagging engine: issue taxonomy by detection tier.

R = deterministic rule over extracted structure. M = trained model. L = rubric-based LLM judge with evidence spans and calibrated abstention (Trust-or-Escalate: 80% human-agreement guarantee at 79% coverage using the strong model on 17.5% of items). H = human required. Every flag carries detector@version, prompt hash, evidence spans with bounding boxes and reviewer disposition.

IssueTierEvidence
Illegible scan, blank page, exact or near duplicate, superseded version citedM, RDocument AI quality score (hard flag < 0.5), mean OCR confidence < 0.8; SHA-256 and MinHash
Title block incomplete; sheet index vs actual sheetsM then RLayout detector mAP50 0.95 (FELD); GPT-4o field extraction 95% value match (arXiv 2504.08645)
Dangling detail or section calloutR, MReference graph with dangling edges; callout bubbles need VLM or symbol extraction
Revision inconsistency; changes outside revision cloudsR, M, LSequence rules; raster or vector overlay diff; vendor precedent (Trunk Tools, Bluebeam), no published accuracy
Spec-vs-drawing conflictL, HRetrieve the pair; judge must quote both sides; labelled “conflict requiring resolution” because precedence clauses often say specs govern
Undefined terms, broken cross-references, number-word mismatchesRLitera Check class of deterministic checks
Cross-document contradiction (dates, amounts, parties)R + L, HTyped fact store; same-key different-value candidates; hybrid NLI + LLM F1 89.5 self / ~71 pairwise (LegalWiz); humans found 43% of injected contradictions
Citation nonexistent or unsupportiveR; L, HNon-generative database lookup (Clearbrief pattern); semantic sentence-to-source scoring; quotation match stays string-based

Questions firms ask

The questions an IT and risk committee asks first.

How is the matter wall enforced?

Postgres row-level security keyed on tenant, matter and principal, set once per transaction from a verified token. The application role cannot bypass it; application WHERE clauses are for performance only. A penetration test with two synthetic tenants runs in CI and blocks deploys.

How are answers verified?

Two deterministic checks after generation: citation-exists (the cited span is in the retrieval set) and quotation-match (string comparison against the page text). Model-reported support scores are advisory. Empty retrieval yields abstention, which is a first-class output.

Why not a managed RAG service or a vector database?

Matter-scoped RLS, audit of every retrieval set, and in-account deployment rule out managed stores. Postgres with pgvector handles the sizing (≈9M chunks, 500K sheets) with HNSW on halfvec; OpenSearch is deferred until golden-set evidence demands it.

How do model and prompt changes ship?

Prompts, rubrics, detectors and model ids are content-hashed versions. A change is a new version plus a frozen regression run that must pass floors before promotion. A model-id change is a prompt change. ZDR tenants are restricted to ZDR-eligible generations at config load.

The 90-day plan

Build plan: three 30-day phases with measured gates.

Phase order follows risk: text retrieval first (known pattern), drawings second (the hard part), LLM-tier flagging last (needs labelled data from the first two). Durations assume a corpus sized by the bake-off; a predominantly raster drawing corpus shifts effort into month two.

01Days 1–30
Text brain on two matters
  • Bake-off: three parsers, two VLMs on 500 firm pages and 200 sheets; vector-vs-raster census
  • Ingest: DMS connector, SHA-256 + MinHash dedup, inclusive threading, version lineage, matter tagging
  • Paragraph-group chunking with contextual prefixes; hybrid BM25 + dense index in Postgres/pgvector with RLS; Cohere or Voyage rerank
  • Single-shot and agentic query classes; Citations API; citation-exists verifier; audit log
Gate to pass200-pair associate-written golden set: faithfulness and citation precision at agreed floors; RLS penetration test; P95 latency target.
02Days 31–60
Drawing pipeline and Tier 1 flags
  • Rasterize at 200–300 DPI; layout detector for title block, legend, schedules, revision table, viewports; full-res crop extraction to a fixed schema
  • Reference graph (sheet, detail, spec section, RFI, change order) in a Postgres adjacency table
  • Page-image late-interaction embeddings, token-pooled; VLM reranking on crops; sheet captions with calibrated abstention
  • Tier 1 rules: sheet index diff, dangling callouts, revision sequence, title-block completeness, undefined terms, broken cross-refs, cite-exists
Gate to passTitle-block field accuracy and sheet-index consistency on a labelled set; Tier 1 precision near 1.0 per rule on the seed set; drawing-navigation questions answered with sheet citations.
03Days 61–90
Tier 2/3 flagging, validation, handoff
  • Rubric LLM judge with self-consistency confidence and escalate/abstain thresholds; typed fact store and cross-document reconciliation; revision-cloud diff
  • Review queue routed by severity x confidence; Word add-in or DMS panel surface
  • Stratified recall validation with 95% CIs; frozen regression sets; drift monitoring from reviewer dispositions
  • Runbook, IaC, nightly ingest and flagging job, training, remaining matters queued
Gate to passHuman-agreement target met on a held-out set with the escalation procedure; per-detector recall and precision published to the firm; the firm operates the nightly job unassisted.
$50,000 / monthOne fee for the Texo team across all three phases. No licenses, no per-seat pricing.
$150,000Total for the ninety days, invoiced monthly, each month gated by the test above.
At costFee covers the Texo team (engagement lead, architect, two engineers part-time, a reviewer-labelling coordinator). Infrastructure and model usage run in the firm’s accounts at cost: parsing ~$1–10 per 1,000 pages by path, drawings ~$30–130 per 1,000 sheets per VLM pass, queries ~$0.05–0.50 each by class.

Deliverables at day 91.

  • Infrastructure as code in the firm’s cloud account; Postgres + object storage; private model endpoints
  • Ingest, parse, index and flagging services with MCP tool layer and audit tables
  • Golden sets, seed sets and frozen regression suites with per-detector metrics and CIs
  • Validation report in TAR-style format (recall, precision, elusion, sampling method)
  • Runbook, on-call procedure, model/prompt change protocol
  • Priority-ordered backlog for the remaining matters and the integrations not yet built

Next step

Bring one matter. We will run the bake-off on its documents and tell you honestly what the system will and will not catch.

Schedule a conversationjack@texoadvisors.com