Which sites received the suspect batch, and which customers have it now? The answer exists. It sits in the ERP, the manufacturing system, the quality register and the CRM, under four different part numbers. Someone opens a ticket. Three weeks later a spreadsheet arrives, reconciled by hand, and the batch has already shipped. The remedy proposed at the next steering committee is a data warehouse.
The warehouse is the pilot’s big brother
A pilot copies the business once and asks the copy questions. A warehouse copies the business continuously, at ten times the cost, and calls the copying a strategy. The two share a foundation, and the foundation is the copy.[1]
The programme runs the same way every time. A year of pipelines. A schema frozen in month three by a team that never handled the part. A reconciliation layer, then a team whose career is the reconciliation. At the end, the question that started the programme is still open, because the warehouse holds what was loaded, in the shape someone chose before the question was asked.
IBM describes the condition the warehouse is meant to cure: data scattered across systems, applications, clouds and documents, no single source of truth, and stale by the time it reaches its downstream use. It lists consolidation into a warehouse or lake through ETL pipelines among the remedies.[2] The description of the disease is exact. The remedy reproduces it: one more system, one more format, one more version of the part number.
Every format is a customs post
The frontier that matters is the format.[3] Inside a single organisation, every proprietary format levies its toll in friction, delay and reconciliation. Data does not flow. It is extracted, transformed, degraded, reconciled at quarter-end, and finally trusted by no one.
A warehouse does not remove the customs posts. It builds a central one and routes everything through it. The four systems still disagree about the part. The disagreement has moved to a nightly job, where it fails silently.
The ticket between the question and the answer
Those who understand the work cannot execute. Those who execute do not understand the work. Between them sits a ticket queue.[4] The warehouse is where that queue lives. The quality engineer knows what “batch” means on the floor. She writes it down as a request. The request becomes a specification, then a queue position, then a pipeline, validated by people who never held the original knowledge. Fidelity is lost at every conversion, and nobody holds both ends of the chain to notice.
Friction now resides in information: in the days required to establish which of several systems is telling the truth about a single part number, and the further week required to persuade the directors to accept the answer.[5] Whoever compresses that interval prevails. What decides an engagement is completing the loop faster than the other side, so that his picture of reality is always one iteration old.[6] Your competitor reads the same consultants’ deck. What defeats him is that your organisation answers a question about itself in the time it takes to ask it, while his requires a steering committee and a warehouse programme.
Go to the data
The alternative is old and has a name. Data virtualisation lets an application query sources where they stand: the data remains in place, and access is given to the source system rather than to a copy.[7] Nothing is extracted. Nothing is stale by construction. The same article lists the cost: every source must be reachable, and the approach records no historic snapshot.[7:1]
Name it, use it, then say what it does not give you. A federated query across four systems that call the same part four things returns four answers faster. It reconciles nothing, because reconciliation is not a data problem. It is a vocabulary problem. One cannot observe that for which one possesses no concept, and an institution with no explicit model of itself has no orientation, only reflexes and folklore.[8] The model has to be declared first, in the words of the people who do the work. A programme that schedules it as documentation after the load has inverted the order of operations.[8:1]
A machine of consequence needs three things beyond the query. A declared model of what the part is, across all four systems. A path from every figure back to the row it came from. And a name at the foot of the decision the answer produces.
What Galahad refuses to do
We do not migrate. We do not build a copy of your business and ask the copy questions. We do not hold your data: the software installs inside your infrastructure, under your accounts and your jurisdiction, and reads across the systems you already run.[9] A copy is stale by construction, a copy is a fifth system with a fifth opinion, and a copy that leaves your network has left your jurisdiction with it. Every step auditable, every decision traceable, and nothing moved.
The questions to put to the vendor
- Does your product need my data moved, copied or indexed before it can answer? Where does the copy live, and when was it taken?
- When I ask about a part number, which system’s version do I get, and who decided that?
- Can the person who knows the trade change the vocabulary, or does that need a ticket?
- If one source is unreachable, does the answer say so, or does it answer from what it had?
- Who holds the model of my business, and does it leave with me?
Monarch reads across the systems you already run, inside your infrastructure. Nothing is migrated. See it on your data.
IBM, What is data fragmentation?: data scattered across systems, applications, clouds, databases and documents; stale by the time it reaches its downstream use; consolidation into a warehouse or lake through ETL/ELT pipelines among the remedies. Source ↩︎
Wikipedia, Data virtualization: the data remains in place and real-time access is given to the source system, unlike ETL; among the drawbacks, every source must be reachable and the approach is not suited to historic snapshots. Source ↩︎ ↩︎
Galahad, the company: installed inside your infrastructure, under your accounts, your keys and your jurisdiction; we operate no hosting for your data. Read it ↩︎