Primary and secondary data: which to collect for which line
Modelling everything with primary data is neither possible nor necessary. The foreground-background split and allocating a data collection budget well.
By clca Editorial TeamLast updated
Data collection is the largest effort line in an LCA, and where that effort goes decides the study's quality. There are two common extremes: trying to model everything with primary data, and collecting nothing and relying entirely on a database. Both are wrong.
Foreground and background
The foreground system is the processes you control or can directly influence: your own site, your own recipe, your own energy consumption, your direct suppliers. The background system is the generic inputs you buy from the market: grid electricity, commodity chemicals, standard transport services.
The rule is simple: model the foreground with primary data and the background with a consistent database. A study without foreground primary data describes the sector average rather than your product, and provides no decision support.
What EN 15804 expects
When preparing an EPD that distinction is not optional. EN 15804 expects data on the production processes — energy consumption, raw material quantities, production waste, direct emissions — to be site-specific. Background database use is legitimate for the upstream chain of purchased inputs, not as a substitute for foreground data.
- Your own site's energy and material flows — always primary
- Main raw materials — request an EPD or product footprint from the supplier
- Grid electricity — a country or regional factor; secondary data is appropriate
- Transport — distance primary, vehicle emission factor secondary
- Minor additives — secondary data suffices, but check environmental relevance
- Waste treatment — depends on regional infrastructure; secondary but must be regional
Allocating the data collection budget
If time is limited, sensitivity analysis should decide which lines get primary data. Build a rough first model, run it entirely on secondary data and see which lines carry the result. Spend the primary data collection effort only on those.
Beyond optimising effort, this approach delivers a second benefit: you have documented why each piece of data was collected. When the verifier asks why a line was modelled on secondary data, you already have the answer.
Tags
- primary data
- secondary data
- data quality
- ecoinvent
- EN 15804
- data collection
Related reading
All articles in Data & Methods- Data & Methods
LCI databases: choosing between ecoinvent, EF and sector datasets
Choosing a background database is not a technical detail but a methodological decision that changes the result. Coverage, system models and why you must not mix.
11 min read - Data & Methods
Sensitivity analysis: which assumption is actually carrying the result?
An LCA result is only as solid as its weakest assumption. One-at-a-time scans, hotspot analysis and how this differs from uncertainty.
10 min read - Data & Methods
Midpoint vs endpoint LCIA: which level for which decision?
Midpoint analysis offers 18+ categories; endpoint collapses to 3 main areas. The practical impact of characterisation, normalisation and weighting steps.
9 min read - Data & Methods
CML vs ReCiPe vs IMPACT World+: which LCIA method when?
Differences between the three leading LCIA methods: midpoint vs endpoint, geographic resolution, EN 15804 and PEFCR alignment, academic use.
11 min read