Picture a single drill intercept: a grade, an interval, a depth, a coordinate. Four numbers that, between them, tell a working geologist almost everything that matters about one line in a corpus.

Now picture how that same intercept actually arrives. The grade is in volumetric units, because the report is from the late 1970s. The interval is a horizon-average, not a sampled length. The depth references a regional datum that is no longer current. And the coordinate is referenced to the Pulkovo Meridian, the reference point of the Russian Empire, sitting 30°19′34″ east of Greenwich.

Get one of those four wrong and the asset lands in the wrong country. This is the median morning at any team trying to ingest a real corpus of historical, multilingual mining disclosures, and it is the work almost every off-the-shelf platform quietly fails on.

Most platforms were built around the cleanly-disclosed quarterly figures of major operators in the major reporting jurisdictions. The rest of the industry has been left to fend for itself. This white paper is about the rest of the industry.

What the full white paper covers

A field report, written for chief geologists, technical directors, exploration leads and CTOs, and for the heads of data science, data scientists and data engineers who support them. Drawn from live engagements, anonymised under client agreements:

  • The five normalisation disciplines: units of measure, currency conventions, reporting-standard reconciliation (JORC ↔ NI 43-101 ↔ GKZ), georeferencing, and 150 years of shifting measurement convention. Every one solvable; none of them casually.
  • Why 90% accuracy is worthless, and what it actually takes to push past 99% on a real Cyrillic, multi-format, mid-twentieth-century corpus.
  • A worked example, anonymised: a tier-one producer's pilot on six complex Soviet-era geological reports, benchmarked against the producer's own ground truth.
  • The path from corpus to working geologist: translate, search, extract, spatially link, and visualise natively in ArcGIS, QGIS or LeapFrog.
  • What the alternatives actually do: generic OCR, generic LLMs, and Copilot inside Office 365, measured against a real technical report.
99.21%
Element accuracy on the scored ground-truth corpus
11,249
Drill-hole assay rows extracted across 37 drill-log images
0
Hard processing failures across 164 input files

The short version of a long answer: it took two years of blood, sweat and tears, and an architecture that compounds rather than degrades. The full paper is the field report: the disciplines, the failure modes, and the numbers behind them.

Download the white paper

Enter your details and we'll send the full paper straight over. No sales call, just the work.


Pulse Intelligence

Less Searching. More Strategising.™

See the platform running on real geological data. Scoped proof-of-concept engagements available, benchmarked against your own ground truth.