Definitional disagreement is the binding constraint
Two functions that cannot agree what a customer is will never reconcile their reports, and no architecture resolves that. The governance work is the project; the engineering is the implementation.
Consolidation projects are usually justified on efficiency and delivered on hope. This paper examines where the value actually is and what the honest cost looks like.
Every institution above a certain size has a data consolidation programme, a business case built on analyst productivity, and a set of extracts that continue to circulate regardless.
This paper argues that the productivity case is the weakest available justification, that the real value is in traceability and decision speed, and that the cost most programmes underestimate is not engineering but the governance work of agreeing what words mean. It is written for institutions deciding whether to start, and for those wondering why a programme already underway has not yet changed anything.
If you read nothing else, read these. The analysis that follows sets out the evidence for each.
Two functions that cannot agree what a customer is will never reconcile their reports, and no architecture resolves that. The governance work is the project; the engineering is the implementation.
An examiner asks who approved this, on what information, and whether the control operated throughout. A platform that answers those instantly is worth more than one that answers ordinary questions faster.
Parallel spreadsheets persist until the governed number has survived being disputed. Building trust is a sequence of small public wins rather than a launch.
Capacity-based compute rewards efficient design and punishes inattention in a way perpetual licensing never did. Monitoring belongs with the first workload, not after the first invoice.
Microsoft Fabric puts every analytics workload over one logical lake, with data held in open table format so each workload reads the same copy rather than maintaining its own extract. Power BI reads it directly through Direct Lake without an import step, removing the refresh window that made reporting stale before anyone read it.
These are genuine architectural improvements and they solve the problem institutions least often have. The problem institutions most often have is that finance, risk and the front office each define margin differently, and each is defensible.
A single copy of the data does not settle that. It makes the disagreement more visible, which is useful, and it does not resolve it.
We would put it in three places, in order.
First, traceability. In a regulated institution the ability to trace a published figure back through every transformation to the originating record converts an internal report into something a board committee can act on. That capability is architectural and it is difficult to retrofit.
Second, decision latency. Not query speed — decision speed. A figure that arrives continuously rather than monthly changes which decisions are possible, and the value is entirely in what somebody does differently as a result.
Third, and least, analyst productivity. It is real, it is the easiest to model, and it is the argument most likely to be disputed by anybody who has seen a previous programme fail to deliver it.
Not engineering. Agreement.
Every institution we have worked with has underestimated the elapsed time required to get finance, risk and the business to sign up to one set of definitions. It is unglamorous, it involves people whose incentives differ, and it cannot be delegated to the data team without producing a platform the business disputes.
The second underestimate is capacity cost. Consumption-based compute is efficient when designed well and expensive when not, and the feedback arrives on an invoice a month later. An inefficient notebook or an over-refreshed model costs money in a way a perpetual licence never did.
Every paper in this series ends with a framework you can run internally. We would rather you used it and reached your own conclusion than took ours on trust.
Five stages. Institutions that skip the first reliably rebuild it later.
Definitions and owners for every metric that will be published, signed by the functions that use them.
One workload end to end — ingestion to a report somebody uses — before the second starts.
Lineage and reconciliation built in from the first workload, not added when an examiner asks.
Row-level entitlement sourced from the directory, tested with real role accounts including one with no access.
Capacity monitoring, refresh alerting and a change process for definitions.
The same argument lands differently across an executive team. These are the three versions worth separating.
Microsoft's own documentation for the product behaviour described above. We would rather you verified the basis than accepted our summary of it.
On these references: each entry names a Microsoft Learn article or documentation area by title, because deep links change while titles are stable. Searching the title on learn.microsoft.com will reach the current version. Where we have cited a figure or a product behaviour, it is Microsoft's statement rather than ours; where we have given a number of our own it is labelled as such in the text.
If you are deciding whether to start, we will run the definitions workshop for one subject area and build that single workload end to end, so the decision is made on evidence from your own data.
Your core system knows about a checking account, a mortgage and a commercial loan. It does not know they belong to the same household.
The decisions were made properly. The evidence just was not captured at the time, so it has to be reconstructed.
Describe the situation in your own words.