CuratedData

Use cases

Where messy data costs the most.

Examples of the kinds of work we take on, each described by the problem, our approach and what you would receive. Plus ClarionThinking, our own product built on curated data.

Example engagements

01 Tabular

Reconciling figures across systems

Problem
Revenue from the ERP, the billing system and finance spreadsheets disagrees, and nobody can say which number is right.
Approach
Match entities and periods across sources, normalise units and currencies, tie each figure back to the ledger and explain every variance.
You receive
One reconciled dataset with a variance report: each difference resolved or explained.
02 Semi-structured

A product catalogue from scraped pages

Problem
Thousands of product pages from different sites list the same items under different names, units and prices.
Approach
Extract attributes from the HTML, normalise units and naming, and cluster listings that describe the same product.
You receive
A canonical catalogue in which every attribute links back to its source page.
03 Documents

Contracts and invoices into structured data

Problem
Key terms and amounts sit in PDFs and scans; manual extraction is slow and error-prone.
Approach
Layout-aware extraction, cross-checks against totals and related documents, human review of low-confidence fields.
You receive
Structured fields with page references and a review status for each extracted value.
04 Tabular

One customer, one record

Problem
The CRM holds the same customer several times with conflicting addresses, owners and histories.
Approach
Entity resolution with explicit match evidence and survivorship rules agreed with your team before anything is merged.
You receive
A merge plan to approve, then a deduplicated CRM with the full merge history.
05 Text

Making support text usable

Problem
Years of tickets and reviews with inconsistent tags, boilerplate and personal data you cannot analyse safely.
Approach
Strip boilerplate, redact personal data, build a category taxonomy, and audit and correct labels with human sign-off.
You receive
A clean, labelled, privacy-safe corpus ready for analysis.
06 Audio

Call recordings you can rely on

Problem
Automatic transcripts get the important parts (names, amounts, dates) wrong.
Approach
Domain vocabulary, targeted re-checks of numbers and names against the audio, speaker review.
You receive
Corrected transcripts with timestamps, speakers and confidence marks.
07 Images

Auditing an image set before training

Problem
A model underperforms and nobody knows whether the cause is the model or the data.
Approach
Find duplicates leaking between training and test sets, mislabelled images and corrupt files.
You receive
A corrected dataset and an audit report quantifying each issue found.
08 Mixed

Training and evaluation data for AI

Problem
Fine-tuning or evaluation data of unknown quality: duplicates, contamination, inconsistent labels.
Approach
Deduplicate, check for overlap between training and evaluation data, audit labels and document the dataset.
You receive
A curated dataset with a datasheet describing its sources, cleaning and known limits.

In production · our product

ClarionThinking

A free investment-research platform for US equities, built and run by CuratedData. It applies our method to market data every day.

ClarionThinking provides information only and is not investment advice.

01

Ingest

Prices, fundamentals and filings arrive daily through automated, monitored pipelines.

02

Curate

Point-in-time history, designed so figures use only information available at the time. Delisted companies are kept to avoid survivorship bias, and stale prices are flagged instead of being shown silently.

03

Insight

Valuation percentiles, z-scores and peer benchmarks. Invalid inputs are excluded rather than forced into a number.

04

Serve

A live web application that documents the data sources and methods behind its figures.

Visit clarionthinking.com ↗

Your use case

Don’t see yours? Most aren’t on a list.

Tell us about your data →