Primary and secondary data: which to collect for which line
Modelling everything with primary data is neither possible nor necessary. The foreground-background split and allocating a data collection budget well.
By clca Editorial TeamLast updated

Data collection is the largest effort line in an LCA, and where that effort goes decides the study's quality. There are two common extremes: trying to model everything with primary data, and collecting nothing and relying entirely on a database. Both are wrong.
Foreground and background
The foreground system is the processes you control or can directly influence: your own site, your own recipe, your own energy consumption, your direct suppliers. The background system is the generic inputs you buy from the market: grid electricity, commodity chemicals, standard transport services.
The rule is simple: model the foreground with primary data and the background with a consistent database. A study without foreground primary data describes the sector average rather than your product, and provides no decision support.
What EN 15804 expects

When preparing an EPD that distinction is not optional. EN 15804 expects data on the production processes — energy consumption, raw material quantities, production waste, direct emissions — to be site-specific. Background database use is legitimate for the upstream chain of purchased inputs, not as a substitute for foreground data.
- Your own site's energy and material flows — always primary
- Main raw materials — request an EPD or product footprint from the supplier
- Grid electricity — a country or regional factor; secondary data is appropriate
- Transport — distance primary, vehicle emission factor secondary
- Minor additives — secondary data suffices, but check environmental relevance
- Waste treatment — depends on regional infrastructure; secondary but must be regional
Allocating the data collection budget
If time is limited, sensitivity analysis should decide which lines get primary data. Build a rough first model, run it entirely on secondary data and see which lines carry the result. Spend the primary data collection effort only on those.
Beyond optimising effort, this approach delivers a second benefit: you have documented why each piece of data was collected. When the verifier asks why a line was modelled on secondary data, you already have the answer.
Tags
- primary data
- secondary data
- data quality
- ecoinvent
- EN 15804
- data collection
Related reading
All articles in Data & Methods- Data & Methods
LCI databases: choosing between ecoinvent, EF and sector datasets
Choosing a background database is not a technical detail but a methodological decision that changes the result. Coverage, system models and why you must not mix.
11 min read - Data & Methods
Sensitivity analysis: which assumption is actually carrying the result?
An LCA result is only as solid as its weakest assumption. One-at-a-time scans, hotspot analysis and how this differs from uncertainty.
10 min read - Data & Methods
Transport in LCA (A4): tkm, load factor and empty return
A4 is the transport of the product from factory gate to site, calculated as mass × distance × a per-tkm factor. Vehicle class, utilisation, empty return and the trap of volume-limited products; the EN 15804 scenario table and a worked example.
11 min read - Data & Methods
Discount rate in LCC: real or nominal, and which rate?
The discount rate is the single assumption that moves an LCC result most. The real-versus-nominal distinction, the Fisher relation, typical ranges for public and private projects, the advantage of real terms under high inflation and a sensitivity example.
8 min read