GeoLassoSPATIAL EXPLORATION INTELLIGENCE ← Back to the overview

The trust mechanism · in numbers

Data quality, measured

Our error rates, with their denominators — and one rule underneath: suspect data is flagged, never silently corrected, because the geologist who signs a model is accountable for every number in it.

Evidence base: pilot corpus — 2 companies · 13 drill-table PDFs · 30 hand-verified intervals (Mt Isa district, Jun–Jul 2026). Expanding Aug–Sep 2026 to 20+ companies and 300+ verified intervals; ground-truth fixtures and precommitment hashes publish with it.

Measured results

What we measuredResultDenominator / method
Full-corpus extraction run377 PDFs (two companies, 5 years) → 13 carried drill tables → 41 collars, 218 assay intervalsevery announcement fetched — none hand-picked
Audit findings on that corpus73 problems → 0 errors, 1 warningsix-dimension audit: accuracy · completeness · timeliness · consistency · validity · uniqueness
The surviving warninga real grade difference between two announcements of the same interval — left flagged for a geologistflag-never-correct, working as designed
Cross-announcement duplication136 → 95 intervals after merge/dedupthe same holes republished across announcements; provenance retained
Import-safe identity84 duplicate-ID collisions → 0incl. one hole ID shared by two collars 17 km apart; delivery-unique IDs, source ID kept as an alias
Template extraction recall30/30 intervals · 95% CI [88%, 100%] (n=30)hand-verified fixtures, regression-tested; measures the template on its own format family. Unseen layouts are loud-flagged, never silently parsed — cross-format performance is what the expansion measures
Machine-readability of the recordcore intercepts + grades ~95%; full collar geometry ~50% per announcementmeasured against a hand-verified corpus; geometry assembled across announcements
Legacy open-file reports0/18 matched our modern templatespre-JORC layouts — which is why legacy extraction runs as a separate, human-verified tier, not a template product
Open-file ingest, deduplicatedone legacy project: 764 collars / 7.7 k assays; audit problems 893 → 7interval dedup on ingest
Gaps we publish, not hideone prolific operator: 196 announcements, 0 collars in the showcase area — and that is correctall 196 spatially checked: their drilling clusters 20–70 km north of the box. Holding a permit ≠ drilling it

The same accounting, drawn

Documents — where the corpus narrows

One run · 2 companies · 5 years. Units: announcements (PDFs).

Announcements downloaded 377 announcements — every one fetched, none hand-picked 377 Headline-classified as drill results (a pre-filter) 81 announcements — headline classifier; kept broad on purpose 81 Carried extractable drill tables (the router's truth) 13 PDFs — the rest loud-flagged, none silently parsed 13

Most announcements carry no drill table at all — the narrowing is the record's shape, and every drop is accounted for, never silently skipped.

Intervals — from extracted to delivered

From those 13 PDFs. Units: assay intervals.

Extracted 218 assay intervals extracted 218 Inside the target area (82 outside — filtered out) 136 intervals inside the 5 km area of interest 136 Unique, delivered (41 republished duplicates merged) 95 unique intervals after cross-announcement dedup — provenance kept 95

The delivered number is smaller than the extracted one — because the same hole announced three times must arrive as one.

Problems found → problems left

found by the auditleft after dedup + fixesthree different corpora — scales are independent
ASX corpus audit 73 findings on the first pass 73 1 warning left — a real cross-announcement grade difference, kept flagged 1 the 1 stays on purpose Duplicate-ID collisions 84 colliding collar IDs found 84 0 — resolved by the delivery-unique identity model 0 identity model, not deletion Open-file ingest 893 problems on raw legacy ingest 893 7 left after interval dedup 7 one legacy project

The boring correctness layer

Same hole, announced three times

Initial release → quarterly → investor deck: recognised as one hole via a natural-key identity model; intervals merged with provenance to every source document.

A company corrects an earlier result

Both values survive. The conflict is flagged for geologist review — history is never silently replaced.

Subsidiary → listed-parent resolution

A curated, human-maintained alias register. It does not invent links; unresolved holders are flagged. Automatic inference is deliberately not trusted yet.

Provenance travels with the data

Every delivered value links to its source announcement and page. The trust columns — confidence, holder match, completeness regime, source page — ship inside the export, not only on screen.

Grade plausibility, geology-calibrated

A two-tier check separates "above range — verify (possible bonanza)" from "implausible — likely a unit or source artifact", calibrated to published district grades. A real high-grade hit is not flattened by the rule that catches a units error.

Two guarantees, kept apart

Per-delivery: your pack carries its own audit

Every delivered dataset ships with the full audit run on that data — its flags, its completeness ledger, its provenance. You are never asked to trust a global average about your own ground.

Global: the generalisation evidence, versioned

The table above, at its current pilot scope. It grows by protocol: sampling fixed before extraction runs, confidence intervals on every rate, fixtures downloadable, hashes published before verification.

The review ledger

Extraction is assisted; acceptance is human. Every batch below was locked with a SHA-256 hash of the raw predictions before review — so neither the sample nor the results could be quietly adjusted afterwards. Corrections found are published, because a review that never finds anything isn't reviewing.

First reviewed batch lands here (Aug–Sep 2026), with its precommitment hash and correction count.

Have a say in what gets verified next. Suggestions feed our public verification queue — name a company or draw an area and mention "verify" in the notes. Need an area compiled for your own work, guaranteed and model-ready? That's an area request.

Take our data apart

A demo can be faked; a hostile review cannot. We invite sceptical geologists to take a delivered dataset apart against the source announcements. Findings are published here regardless of outcome — including the errors found.

First external review: being arranged (Aug–Sep 2026) — results will appear on this page.