Skip to content
Signal · Data

Your 2030 Target Is Hiding in Dark Data

By Lewis Howard · Founder, Co-CEO BRAE13 August 20269 min read

You will miss your 2030 carbon target. Not because you cannot decarbonise, but because you cannot measure it.

And you cannot measure it not because the data is missing, but because it is dark - it already exists, inside your own business and your suppliers', simply never connected to the question.

Dark data is a term borrowed from data management, and in its usual sense it undersells itself.

There, it tends to mean the information a business collects and stores in the ordinary course of operating but never looks at again - dormant logs, dusty archives, records kept for compliance that cost more to hold than they will ever return. The image behind the name is dark matter: unseen, unmeasured, and yet most of the universe by mass. Framed that way, dark data is waste, and the task is to find it and clean it up.

In sustainability the term means something sharper, and close to the opposite. Here, dark data is the data an organisation already holds, already trusts, and often already shares - which has simply never been connected to the question of its climate and nature position. It is not waste. It is the organisation's best data.

It is already systematised. It already carries provenance, because it is trusted enough to run real decisions on. It is already used - for commercial choices, for operational planning, for reporting up and down the value chain to trading partners who rely on it. One or more of those things is almost always true of it. It is not dormant, and it is not low quality. It is live, high-grade, and in motion. It is simply dark to one function. It is not dark to the business. It is dark to sustainability.

The unlock: leaving spend-based behind

Once you see it that way, the size of the prize changes - and it has a name. The most valuable thing connecting dark data does is let a company move off spend-based carbon footprinting.

Spend-based accounting is the method of last resort: multiply what you paid a supplier by a sector-average factor, and book the result. It rests on an assumption - that the real activity data does not exist. I have written elsewhere about how wrong the number it produces can be: in work with a large multinational client, gaps between the spend estimate and suppliers' own verified footprints ran from a factor of two or three to, in the worst case, a thousand. But the deeper problem is not only that spend-based is inaccurate. It is that the assumption underneath it is usually false. The real data mostly does exist. It is simply dark.

Consider what already moves between organisations for reasons that have nothing to do with carbon. Freight and logistics data passes between shipper and carrier - real weights, real distances, real modes, in many cases real fuel burned. Bills of materials pass between buyer and supplier for procurement and quality, carrying the actual composition of the thing being bought. Metered energy passes between a site and its utility. Telematics pass between a fleet operator and the company that leases it the vehicles. This is structured, verified, provenance-carrying data, exchanged every day, trusted enough that commercial contracts settle on it - and far richer than any factor. The sustainability team reaches for the sector average while the activity itself flows past, recorded, unread.

Why the best data goes dark

If this data is so good, why is it dark? The answer is almost always the same: the sustainability function is looking in one place, and the data lives somewhere else.

Scope 3 programmes fish in the ERP, because that is where spend sits, and spend is what the method needs. But the record of what was actually made, moved, delivered and consumed mostly lives outside it, in systems the sustainability team rarely opens. Four are worth knowing by name.

Suppliers keep better records of their sales than buyers keep of their purchases. A supplier's revenue depends on knowing exactly what it sold, to whom and in what quantity, and that record is reconciled and often audited. The buyer's record of the same transaction is frequently just a cost line. So the supplier's sales ledger is usually a better account of what you consumed than your own purchasing data is. A technology reseller knows the precise units, models and configurations it shipped you, where your procurement system may show only "IT equipment" and a figure. The data you need is sitting in your supplier's sales system, kept to a higher standard than your own.

A great deal of performance-management data never touches the ERP at all. The operational detail of what happened lives in performance systems, and only a summarised cost reaches finance. A media agency's performance reports describe placements, impressions and delivery in fine grain; the ERP sees "media" and a number. The same gap runs through IT, facilities, logistics and professional services - the invoice is the shadow the activity casts on the finance system, and the substance is elsewhere.

Supply-chain performance data is often already automated. Operations depend on tracking volumes, movements, deliveries and throughput continuously, so that data is high-frequency, systematised and machine-generated - a close description of physical activity, collected for operational reasons and never routed to carbon accounting.

Business-continuity and regulatory systems record where things come from. Disaster-recovery and compliance systems exist to answer where a product originates, where a service is delivered from, and what happens if a site fails. To do that they hold precise origin and location data - exactly what a footprint needs to attach a grid factor, a water-stress figure or a nature impact, and exactly what the sustainability team is so often told does not exist.

The common thread is that spend-based footprinting is what you get when you only fish the ERP. Because the ERP holds spend, spend is what you build the footprint from. Connecting dark data is not a measurement programme; it is a routing problem. High accuracy, low incremental effort, because you are not creating anything - you are pointing the reporting function at data the organisation, or its suppliers, already holds and already believes. That is how you leave spend-based behind: not with a better factor, but by fishing where the real data is.

That is the unlock, and on its own it would be worth the article. But it is only half the story, and the smaller half.

The hierarchy nobody climbs

There is an old model in knowledge management called the DIKW hierarchy: data, information, knowledge, wisdom, stacked in that order. Raw data sits at the base, meaning nothing on its own. It becomes information when you give it context, so that it starts to answer a question. Information becomes knowledge through analysis and synthesis, and - this is the part that matters - knowledge is characteristically the product of experience, expertise and judgement, not of data alone. At the top sits wisdom, which the model describes as knowledge applied to the question of what is actually best to do.

The model has its critics, and the criticism is fair: it implies a tidy upward ladder when real understanding is provisional, looping, revised through experience rather than climbed in clean stages. But that critique cuts in a useful direction here, because it locates precisely the thing sustainability assessment keeps leaving out.

Look at where the discipline actually operates. Reporting sits at the bottom of the hierarchy - and, in the spend-based case, not even on the bottom rung but on a proxy for it, a modelled stand-in for data the organisation could have had directly. The unlock I have just described is really an argument for climbing one rung: from a proxy, to real data, to that data placed in context as information. Worth doing. Necessary, even. But the serious failure is not at the bottom of the ladder. It is at the top.

Amputated at the top

Sustainability assessment is stuck in data and information, and it has cut off the two rungs above - experience and wisdom - as though they were not admissible evidence.

The reason is structural, and it is worth being precise about, because it is not stupidity and it is not bad faith. Assurance credits what is documentable. A spend figure multiplied by a published factor is fully traceable: every step has a source, every number has a citation, the whole chain can be audited and signed. Data and information live comfortably inside that world. Experience and wisdom do not. The knowledge held by the person who receives the service, manages the contract and signs off the invoice - who can tell you without hesitation what a supplier does and where - leaves no audit trail. Judgement does not come with a reference. So the very mechanism built to create rigour quietly pushes the assessment down onto the lowest rungs, where a number can sit comfortably precisely because it is auditable, and rules out the rungs where a person would simply say: that is obviously wrong.

Which produces the perverse result at the centre of all this. The more seriously an organisation takes assurance, the more firmly it can entrench a figure it knows to be false. The spend-based number survives not despite the contradicting information but because the framework only credits the rung the spend number sits on. When a supplier's own life-cycle assessment says one thing and the spend factor says something a hundred times different, and the organisation reports the spend factor anyway, that is not a data problem. The data has already spoken. It is a refusal, or an inability, to let experience and judgement into the assessment at all.

The inversion of wisdom

There is a sharp irony to end on.

Wisdom, in the model, is knowledge applied to the question of what is best. Clinging to a number you know to be wrong because it is the assurable one is the exact inverse of that. It is letting the process stand in for the judgement the process was meant to serve. The discipline has, in effect, sawn off its own top two rungs and then wondered why its assessments keep defending figures that everyone close to the work knows are wrong.

I made a related point in the last piece, from the reporting side: the gap between what you can prove to an auditor and what you know to be the case. This is the same fault seen from the assessment side. There, the answer was to keep two numbers - the assured one for the external report, the truer one to manage the business on. Here the answer is the same instinct, generalised: assessment has to make room for evidence that is real but not yet audit-grade, and for the experience that no dataset holds, or it will keep grading confidence in exactly the wrong direction.

None of this means abandoning assurance. It means being honest that assurance answers one question - is this documented? - and not the question we actually care about, which is: is this true? Those come apart more often than the discipline admits, and where they come apart, a good assessment follows the truth and records its confidence honestly, rather than following the paperwork and calling the result rigour.

What this asks for

So the unlock and the indictment turn out to be the same argument, read at two heights.

At the bottom: connect the data you already hold and already share, and use it to leave spend-based footprinting behind. It is systematised, it has provenance, and pointing the reporting function at it is low effort for a large gain in accuracy. At the top: build an assessment that can climb - that admits real-but-unassured data and the tacit knowledge of the people who do the work, each carrying its own honest grade of confidence, rather than admitting only what survives an audit and mistaking that for the whole truth.

Doing both is a knowledge problem, not a data problem, and that distinction is the whole point. You are not assembling a bigger pile of numbers. You are building a foundation that runs all the way up the hierarchy - data connected, placed in context, informed by experience, and used, finally, to decide what is best to do. Provenance-grading is what makes the climb honest: it lets a number and a judgement sit in the same assessment, each labelled for exactly how much weight it can bear. That grading, and that connecting, is what Reverberate is built to do - and it is why we treat this as a knowledge foundation rather than a reporting pipeline.

Most of what determines a company's real climate and nature position is already known - to the business, and to the partners it trades with every day. The failure was never that the information did not exist. It was that we built a discipline that could only see the rungs it could audit, and called the rest dark.

Sources

Data-Information-Knowledge-Wisdom (DIKW) Pyramid - Springer / ISKO Encyclopedia of Knowledge Organization. The DIKW hierarchy as used in information science and knowledge management: data as the base, information as contextualised data, knowledge as the product of experience and synthesis, wisdom at the apex.

The DIKW Model: A Useful but Flawed Map - Conversational Leadership. The standard critique of DIKW: it implies a tidy upward ladder, whereas real understanding is provisional and revised through experience rather than climbed in stable stages.

Topics
dark dataScope 3spend-based emissionsactivity-based dataDIKW hierarchyknowledge managementassuranceprovenancevalue chain data
← All of Signal