Data & Methods10 min read

Primary and secondary data: which to collect for which line

Modelling everything with primary data is neither possible nor necessary. The foreground-background split and allocating a data collection budget well.

By Last updated

Data collection is the largest effort line in an LCA, and where that effort goes decides the study's quality. There are two common extremes: trying to model everything with primary data, and collecting nothing and relying entirely on a database. Both are wrong.

Foreground and background

The foreground system is the processes you control or can directly influence: your own site, your own recipe, your own energy consumption, your direct suppliers. The background system is the generic inputs you buy from the market: grid electricity, commodity chemicals, standard transport services.

The rule is simple: model the foreground with primary data and the background with a consistent database. A study without foreground primary data describes the sector average rather than your product, and provides no decision support.

What EN 15804 expects

When preparing an EPD that distinction is not optional. EN 15804 expects data on the production processes — energy consumption, raw material quantities, production waste, direct emissions — to be site-specific. Background database use is legitimate for the upstream chain of purchased inputs, not as a substitute for foreground data.

  • Your own site's energy and material flows — always primary
  • Main raw materials — request an EPD or product footprint from the supplier
  • Grid electricity — a country or regional factor; secondary data is appropriate
  • Transport — distance primary, vehicle emission factor secondary
  • Minor additives — secondary data suffices, but check environmental relevance
  • Waste treatment — depends on regional infrastructure; secondary but must be regional

Allocating the data collection budget

If time is limited, sensitivity analysis should decide which lines get primary data. Build a rough first model, run it entirely on secondary data and see which lines carry the result. Spend the primary data collection effort only on those.

Beyond optimising effort, this approach delivers a second benefit: you have documented why each piece of data was collected. When the verifier asks why a line was modelled on secondary data, you already have the answer.

Tags