What we clean
Six kinds of data. One standard.
Tables, semi-structured feeds, text, documents, images and audio. For each one: the problems we catch, what we do about them, and what you receive.
Tabular
Spreadsheets · CSV exports · database tables · warehouse views
Problems we catch
- Duplicate and near-duplicate records under slightly different names
- Units, currencies and scales mixed in one column
- Joins that silently drop or multiply rows
- Totals that no longer tie out after aggregation
- Stale values carried forward as if they were current
- Implausible values: negative quantities, future dates, decimal shifts
What we do
We profile every column, reconcile figures back to source totals, resolve entities, normalise units and investigate each outlier before deciding what it is.
What you receive
- Curated tables in your format
- Row-level change log
- Data dictionary
- Reconciliation report
revenue = "1.31" (unit: M)revenue = 1,310,000 USD · unit normalisedSemi-structured
JSON · XML · HTML · API responses · logs
Problems we catch
- Schema drift between versions of the same feed
- The same field under different names or paths
- Nested arrays flattened inconsistently
- Scraped pages carrying navigation, ads and markup noise
- Mixed date formats, time zones and character encodings
- Truncated or malformed records hiding inside large files
What we do
We infer and version the schema, map every source field to one canonical model, extract the content from the markup and quarantine malformed records with the reason attached.
What you receive
- Canonical schema
- Clean JSON or flattened tables
- Field mapping document
- Quarantine log
<span class="price">$1,299<sup>.00</sup></span>price = 1299.00 · currency = USDText
Support tickets · reviews · notes · chat · survey answers
Problems we catch
- Exact and near-duplicate entries
- Several languages mixed in one field
- Signatures, templates and boilerplate drowning the content
- Personal information where it should not be
- Labels and categories applied inconsistently
- Broken encoding (’ instead of an apostrophe)
What we do
Language models read the entries to deduplicate by meaning, strip boilerplate, detect personal data and audit labels. A person reviews proposed relabels and redactions.
What you receive
- Clean corpus
- Redaction log
- Label corrections with reasons
- Category taxonomy
I’ve emailed jane@acme.com twiceI’ve emailed [EMAIL] twice · redactedDocuments
PDFs · scans · forms · contracts · invoices · filings
Problems we catch
- OCR misreads: O for 0, l for 1, lost decimal points
- Tables split across pages
- Values extracted into the wrong field
- Missing pages, amendments or versions
- The same document stored several times under different names
What we do
We extract fields with the page layout in view, cross-check figures against totals and related documents, and send low-confidence fields to a human reviewer.
What you receive
- Structured fields with source page references
- Confidence and review status per field
- Document inventory and version map
Total: 12,98O.00 (OCR)total = 12,980.00 · matches line itemsImages
Product photos · scans · field images · training sets
Problems we catch
- Corrupt or truncated files
- Exact and visually identical duplicates
- Labels that do not match what is in the image
- Metadata (dates, dimensions, location) inconsistent with the content
- Blurred, dark or wrongly oriented images
What we do
We check file integrity, cluster visual duplicates, verify labels with vision models and reconcile metadata. Disputed labels go to a person to decide.
What you receive
- Cleaned image set
- Label corrections
- Duplicate clusters
- Quality report
img_0412.jpg · label = "cat"label = "dog" · corrected, reviewer approvedAudio
Calls · interviews · meetings · voice notes
Problems we catch
- Transcription errors on names, numbers and jargon
- Speakers attributed to the wrong person
- Long silences, crosstalk and background noise
- Timestamps drifting away from the audio
- Partial and duplicate recordings
What we do
We transcribe against your vocabulary, re-check numbers and names against the audio, review speaker turns and mark segments that cannot be used.
What you receive
- Corrected transcripts with speakers and timestamps
- Segment quality labels
- Domain glossary
"the rate is fifty percent"15% · confirmed against audio at 02:14