
Reuse open data while preserving its context
A public portal is a starting point, rather than evidence that every file suits your project. Before reusing a dataset, identify its publisher, reuse terms, observation period and covered people or places. Keep that context alongside transformed data.
Define the question and analytical unit
State the decision: comparing service areas, validating an address or contextualising business activity. Specify whether you analyse people, establishments, enterprises or buildings. A business establishment and its parent enterprise are different units.
Inspect field definitions, granularity and identifiers before importing. Check joining keys and their versions. A postcode does not always correspond to a single municipality; similar addresses do not prove that records describe the same building.
Read the specific dataset licence
Locate the licence on the dataset page. France’s Open Licence 2.0 includes attribution of the information producer and identification of the information’s latest update. Retain the licence and resource URLs, and examine the requirements before publishing.
Other licences may impose different terms. Personal data and access restrictions require separate assessment. Avoid assigning an open licence to an entire catalogue or equating free access with permission to redistribute.
Check what the observations represent
Distinguish publication date, file-update date and observation period. Examine missing values, duplicates, units, geography and definition changes. A blank value is not zero; a missing area does not demonstrate absence of the phenomenon.
Retain an identifiable working copy and a transformation log: filtering, exclusions, aggregation and correction rules. Measure rejected records and explain their effect. Comparisons require compatible populations and periods and an appropriate denominator.
Publish results with their limits
Alongside the result, identify the producer, dataset, period, version, licence and material transformations. An available source alone does not let the reader reconstruct how observations were selected. Explain what the analysis does not cover.
Plan refreshes around actual need, a responsible person, schema-change checks and unavailable sources. An API also requires inspection of quotas, pagination and service terms. Try a bounded sample before building a production pipeline.
A record to keep with the decision
| Item | Evidence |
|---|---|
| Identification | Publisher, dataset and file URLs, version or retrieval date. |
| Rights | Dataset licence and attribution or sharing requirements. |
| Coverage | Period, geography, analytical unit and missing populations. |
| Processing | Transformations, checks, rejected records and interpretation limits. |
Download the worksheet to fill in (CSV)
Frequently asked questions
Does Open Licence 2.0 allow commercial reuse?
The licence permits commercial exploitation under its terms. Read the version applicable to the specific dataset and provide the required attribution. Other restrictions, including those concerning personal data, need separate assessment.
Is an API always better than a CSV?
An identifiable file supports reproducible one-off analysis. An API may suit regular refreshes. Compare available history, volume, quotas, stability and maintenance against the actual need.
Reference material
data.gouv.fr — Licence Ouverte 2.0
data.gouv.fr — Comprendre les conditions de réutilisation
API Recherche d’Entreprises — Documentation
The practical checklist is an editorial synthesis to adapt to your service. It does not constitute a certification or an audit result.