By · Updated

Team reviewing data provenance and quality for a digital projectData & AI · 12 min

Prepare reliable data before an AI project

A model cannot repair an unclear definition or an unknown source. Before choosing a tool, establish which decision or task needs support, which data describes it, what the data misses and how an error will be detected.

Key points

What to remember.

  • Define a task and success measure before collecting more data.
  • Document source, rights, freshness, limitations and ownership.
  • Test errors and human review on representative cases.
01

Describe the task rather than proposing AI in general

Replace “use AI with our data” with a concrete task such as finding an applicable document, classifying an incoming request or detecting a catalogue inconsistency. Name the user, decision point and costly errors. Compare the idea with a simple rule, better search or clearer documentation before choosing a more complex system.

02

Inventory sources and accountable owners

For each file, database or API, record its producer, operational owner, coverage period, format, update frequency, reuse rights and the person who can explain its fields. Agree on the meaning of business terms such as “active customer” or “closed case”. Without shared definitions, teams can derive incompatible answers from the same records.

03

Measure defects that affect the outcome

Check accuracy, completeness, consistency, uniqueness and freshness for the intended use. A missing value may be tolerable in an aggregate trend but unacceptable for an individual decision. Segment checks by date, channel, region or case type so averages do not hide gaps. Correct root causes where possible rather than concealing them in a pipeline.

04

Clarify access, licence and personal data

Availability online does not grant unrestricted reuse. Read the licence, access terms and original purpose before importing an external dataset. Where personal data is involved, assess necessity, permitted use and access controls. Use an appropriate test dataset instead of copying confidential production documents into an experiment without a defined framework.

05

Build a repeatable preparation process

Document cleaning, matching, deduplication and transformations. Keep provenance and version the field definitions. A changed source schema should trigger a check, not an undocumented manual fix. Separate development data from evaluation cases and look for fields that accidentally reveal the answer. A smaller documented collection can outperform a large opaque assembly.

06

Test difficult cases and real errors

Create evaluation cases reflecting actual use, including incomplete records, rare situations, ambiguity and recent changes. Define criteria before reviewing results: correctness, false positives, false negatives, review cost and effect on the final decision. Have knowledgeable people inspect consequential outputs; escalation or abstention can be better than a confident wrong answer.

07

Assign maintenance and an exit route

Name who monitors new data, reviews evaluation criteria, handles reported errors and approves changes. Record dependencies on an API, supplier or licence. Compare benefits with the full cost of preparation, operation, human review and correction. A pilot should explain what happens if the source disappears or results deteriorate.

Sources

Check the reference material.

GOV.UK · Government Data Quality Framework

CNIL · Collect and qualify training data

Continue

Turn the method into a clear project.

Use the directory to explore relevant resources, or describe your context so the right questions can be identified before a conversation begins.