Prepare reliable data before an AI project
A model cannot repair an unclear definition or an unknown source. Before choosing a tool, establish which decision or task needs support, which data describes it, what the data misses and how an error will be detected.
What to remember.
- Define a task and success measure before collecting more data.
- Document source, rights, freshness, limitations and ownership.
- Test errors and human review on representative cases.
Describe the task rather than proposing AI in general
Replace “use AI with our data” with a concrete task such as finding an applicable document, classifying an incoming request or detecting a catalogue inconsistency. Name the user, decision point and costly errors. Compare the idea with a simple rule, better search or clearer documentation before choosing a more complex system.
Inventory sources and accountable owners
For each file, database or API, record its producer, operational owner, coverage period, format, update frequency, reuse rights and the person who can explain its fields. Agree on the meaning of business terms such as “active customer” or “closed case”. Without shared definitions, teams can derive incompatible answers from the same records.
Measure defects that affect the outcome
Check accuracy, completeness, consistency, uniqueness and freshness for the intended use. A missing value may be tolerable in an aggregate trend but unacceptable for an individual decision. Segment checks by date, channel, region or case type so averages do not hide gaps. Correct root causes where possible rather than concealing them in a pipeline.
Clarify access, licence and personal data
Availability online does not grant unrestricted reuse. Read the licence, access terms and original purpose before importing an external dataset. Where personal data is involved, assess necessity, permitted use and access controls. Use an appropriate test dataset instead of copying confidential production documents into an experiment without a defined framework.
Build a repeatable preparation process
Document cleaning, matching, deduplication and transformations. Keep provenance and version the field definitions. A changed source schema should trigger a check, not an undocumented manual fix. Separate development data from evaluation cases and look for fields that accidentally reveal the answer. A smaller documented collection can outperform a large opaque assembly.
Test difficult cases and real errors
Create evaluation cases reflecting actual use, including incomplete records, rare situations, ambiguity and recent changes. Define criteria before reviewing results: correctness, false positives, false negatives, review cost and effect on the final decision. Have knowledgeable people inspect consequential outputs; escalation or abstention can be better than a confident wrong answer.
Assign maintenance and an exit route
Name who monitors new data, reviews evaluation criteria, handles reported errors and approves changes. Record dependencies on an API, supplier or licence. Compare benefits with the full cost of preparation, operation, human review and correction. A pilot should explain what happens if the source disappears or results deteriorate.
Check the reference material.
Related guides
Put the method into practice
An illustrative application case
An import contains different dates and units. Define target formats and retain original values needed for verification. Test missing values and matching before a full import. Silent conversion can produce plausible but incorrect data.
Document an API before integration
Use this worksheet in a review with the person responsible for delivery. Keep the evidence alongside the decision, rather than marking a task complete on trust alone.
| Action | Evidence |
|---|---|
| Document the business definition, publisher, licence and period of each source. | A field dictionary and provenance record. |
| Measure missing values, duplicates and inconsistencies by relevant segment. | A defect log stating the impact on the intended decision. |
| Reserve an evaluation set and test ambiguous cases. | A result compared with a baseline, recording errors and abstentions. |
Download the worksheet to fill in (CSV)
A question to resolve before acting
Is a larger dataset necessarily better?
No. Coverage, relevance, quality and usage rights are essential. Extra volume that repeats errors or omits a population can degrade analysis. Test the added value for your actual use case.