By · Updated

Data & AI · 12 min

Prepare reliable data before an AI project

A model cannot repair an unclear definition or an unknown source. Before choosing a tool, establish which decision or task needs support, which data describes it, what the data misses and how an error will be detected.

Key points

What to remember.

  • Define a task and success measure before collecting more data.
  • Document source, rights, freshness, limitations and ownership.
  • Test errors and human review on representative cases.
01

Describe the task rather than proposing AI in general

Replace “use AI with our data” with a concrete task such as finding an applicable document, classifying an incoming request or detecting a catalogue inconsistency. Name the user, decision point and costly errors. Compare the idea with a simple rule, better search or clearer documentation before choosing a more complex system.

02

Inventory sources and accountable owners

For each file, database or API, record its producer, operational owner, coverage period, format, update frequency, reuse rights and the person who can explain its fields. Agree on the meaning of business terms such as “active customer” or “closed case”. Without shared definitions, teams can derive incompatible answers from the same records.

03

Measure defects that affect the outcome

Check accuracy, completeness, consistency, uniqueness and freshness for the intended use. A missing value may be tolerable in an aggregate trend but unacceptable for an individual decision. Segment checks by date, channel, region or case type so averages do not hide gaps. Correct root causes where possible rather than concealing them in a pipeline.

04

Clarify access, licence and personal data

Availability online does not grant unrestricted reuse. Read the licence, access terms and original purpose before importing an external dataset. Where personal data is involved, assess necessity, permitted use and access controls. Use an appropriate test dataset instead of copying confidential production documents into an experiment without a defined framework.

05

Build a repeatable preparation process

Document cleaning, matching, deduplication and transformations. Keep provenance and version the field definitions. A changed source schema should trigger a check, not an undocumented manual fix. Separate development data from evaluation cases and look for fields that accidentally reveal the answer. A smaller documented collection can outperform a large opaque assembly.

06

Test difficult cases and real errors

Create evaluation cases reflecting actual use, including incomplete records, rare situations, ambiguity and recent changes. Define criteria before reviewing results: correctness, false positives, false negatives, review cost and effect on the final decision. Have knowledgeable people inspect consequential outputs; escalation or abstention can be better than a confident wrong answer.

07

Assign maintenance and an exit route

Name who monitors new data, reviews evaluation criteria, handles reported errors and approves changes. Record dependencies on an API, supplier or licence. Compare benefits with the full cost of preparation, operation, human review and correction. A pilot should explain what happens if the source disappears or results deteriorate.

Sources

Check the reference material.

GOV.UK · Government Data Quality Framework

CNIL · Collect and qualify training data

Put the method into practice

An illustrative application case

An import contains different dates and units. Define target formats and retain original values needed for verification. Test missing values and matching before a full import. Silent conversion can produce plausible but incorrect data.

Document an API before integration

Use this worksheet in a review with the person responsible for delivery. Keep the evidence alongside the decision, rather than marking a task complete on trust alone.

Actions and evidence to retain
ActionEvidence
Document the business definition, publisher, licence and period of each source.A field dictionary and provenance record.
Measure missing values, duplicates and inconsistencies by relevant segment.A defect log stating the impact on the intended decision.
Reserve an evaluation set and test ambiguous cases.A result compared with a baseline, recording errors and abstentions.

Download the worksheet to fill in (CSV)

A question to resolve before acting

Is a larger dataset necessarily better?

No. Coverage, relevance, quality and usage rights are essential. Extra volume that repeats errors or omits a population can degrade analysis. Test the added value for your actual use case.

Continue

Turn the method into a clear project.

Use the directory to explore relevant resources, or describe your context so the right questions can be identified before a conversation begins.