Skip to main content

Industry examples

LeadershipDomain expertsEngineers
In one minute

DataHub has no vertical editions. The same four building blocks and the same platform serve every industry, the difference between deployments is the model you load into it, and how much of one you need on day one.

Six deployments that share nothing but the platform. The first is worked through in detail, because it shows the whole arc, from a number nobody can defend, to a measured one, to a process that costs less and discharges less.

Oil and gas: cutting chemical use at a processing facility

The clearest example of what a model buys you, because the same work delivers an environmental result and a financial one from a single change, and because it starts from a number almost everybody reports and almost nobody can substantiate.

The situation

An oil and gas processing facility runs on chemicals. They are injected continuously, all over the process, each doing a specific job:

ChemicalInjected intoDoing what
DemulsifierThe separatorBreaking the oil–water emulsion so the two separate
Scale inhibitorThe gas trainStopping mineral scale building up in pipework
BiocideProduced waterKeeping bacteria out of the system
GlycolSubsea linesPreventing hydrate ice forming and blocking flow
Oxygen scavengerWater systemsRemoving dissolved oxygen that would corrode steel

Every one of these ends up somewhere, discharged to sea, emitted to air, or retained in the product. All of it must be reported.

Why today's reported number is not really a measurement

Here is the part that surprises people outside the industry: discharge figures are usually built from what was purchased, not from what was actually used.

Tank deliveries are invoiced, so procurement records exist and are accurate. What happens between the tank and the sea is estimated, dose rates assumed from design values, run hours assumed from schedules, losses assumed from convention. The resulting figure is defensible as an accounting exercise and is not a measurement of anything.

Two consequences follow, and they point in opposite directions:

  • You cannot prove the number is right, which is a compliance exposure that grows as reporting obligations tighten.
  • You cannot find the waste, because a figure derived from purchase records is structurally incapable of showing that a particular pump has been over-dosing for a year.
The same total, twice. Only one of them can be acted on.A figure assembled from what was bought cannot show that one pump has been over-dosing for a year.What was purchasedinvoiced, accurate, and silentWhat was meteredthe same quantity, with structureone number, no structure=per chemical, per injection pointGlycolScale inhibitorOxygen scavengerDemulsifierBiocidethe shaded band is dose the process never neededBoth bars are the same height and both are honest. Only the right-hand one says where any of itwent, which is the difference between a number you can defend and one you can improve.
Both bars carry the same quantity, and the invoices behind the left one are accurate. The difference is that only the right one is attached to the pumps, the process stages and the discharge routes, which is what turns a reported total into something you can act on rather than only defend.

The environmental classification makes it sharper

Offshore chemicals are not interchangeable from an environmental standpoint. They are classified by hazard, conventionally by colour:

Green
Considered to pose little or no risk, glycol, for example
Yellow
In use and permitted, but tracked, scale inhibitors
Red
Substitution actively expected, biocides
Black
Discharge prohibited outright; use only under specific permit

So "how much chemical did we discharge" is really four questions, and the red and black ones carry far more weight than the total. An aggregate figure built from purchase records cannot answer them separately, which means the number that matters most is the one you have least evidence for.

What gets modelled

LayerModelled as
AssetsChemical tanks, dosing pumps, injection valves, the separator, the gas train, produced-water system, metering skid
FunctionsSeparation, gas treatment, produced-water treatment, reinjection, each chemical's injection duty
Business knowledgeDischarge permits and their per-chemical limits, the environmental classification, substitution obligations, the reporting cycle
Time seriesTank levels, pump strokes, valve positions, flow rates, separator performance, water quality
EventsDose changes, pump faults, tank deliveries, permit periods, sampling

The critical relationships are the ones connecting a dosing pump to the part of the process it injects into, and that process to the discharge route and the permit it falls under. Those relationships are exactly what nobody writes down today, and they are what turns a pump stroke into a reportable discharge figure.

Measuring what cannot be measured directly

One obstacle: some of the values you most want are impractical to measure continuously. Oil content in water sent for reinjection is the classic case: it is a slow laboratory test, so the process is corrected after the fact rather than kept in specification.

A soft sensor closes that gap. It is a virtual meter: it watches the cheap signals you already have, pressure, flow, temperature, and a model trained against historical lab samples infers the value you actually want, second by second. You get a live reading where before there was a delayed lab result. Today that computation runs beside the platform, against the same APIs, fed by subscription; running it in-platform is exactly what functions are being built for.

That only works if the model knows which signals belong to which vessel, which lab samples correspond to which stream, and what the process configuration was at the time. In other words, the soft sensor is downstream of the knowledge graph, which is why this is not simply a sensor purchase.

What changes

On the roadmap

The metering, the model and the reproducible figure work today. The traceability that makes a submission self-evidencing, here and in the grid and water examples below, depends on lineage, which is planned.

Injection becomes metered rather than assumed

Tank levels, pump strokes, valve positions and the separator's response are fused into a continuous account of how much of each chemical entered each part of the process, and how much left, to air or water.

The reported figure becomes a measurement

Discharge per chemical and per classification, traceable through lineage to metered injection and measured flows, with data-quality flags where a reading was gap-filled. The submission carries its own evidence instead of resting on purchase records.

Doses get trimmed to what actually works

Once you can see dose against process response, the over-dosing becomes visible. The target is the smallest dose that still does the job, which is almost never the design dose that has been running unchanged for years.

Substitution becomes evidence-based

With per-chemical usage measured, a proposal to move a duty from a red chemical to a yellow one can be argued from data, how much is actually used, where, and what the process response has been.

Why the return is unusually good

Most efficiency projects trade one benefit against another. This one compounds in three directions at once:

Less
Discharged to sea and air
A real environmental reduction, weighted toward the chemicals that matter most
Less
Chemical purchased
Dosing to what works rather than to a design assumption cuts consumption directly
Longer
Asset life
Correct dosing protects pipework and equipment rather than over- or under-treating it
Defensible
Reported figures
A measurement with an audit trail, instead of an estimate from invoices

The third one is easy to overlook and often the largest. Chemical treatment exists to protect steel; getting it right adds years to pipes and equipment, and that shows up as deferred capital rather than as a line in an operating budget.

Why it needs the model

Every step above depends on knowing which pump feeds which process, which process discharges by which route, and which permit that route falls under. That is not data any single source system holds, the tank levels are in one place, the pump telemetry in another, the permits in a document, and the connection between them in somebody's head.

That connection is the model. Build it, and the environmental report, the cost reduction and the equipment-life benefit all fall out of the same work. Building your model →

This sector is also where the model stops being only an analytical asset. Offshore, installations designed to run with no permanent crew are operated through the model itself, so the description built for a discharge report becomes, eventually, the operating surface. Where this is heading →

Starting play: audit-ready reporting on the discharge submission, it already exists, already costs effort, and already cannot be substantiated, which makes it an unusually easy case to fund.

The rest of the sector

The same modelling approach carries across oil and gas generally:

LayerModelled as
AssetsWells, separators, compressors, pipelines, valves, instruments
FunctionsSeparation, compression, export, flaring
Business knowledgeProduction targets, emissions obligations, integrity policies
Time seriesPressure, temperature, flow, vibration
EventsAlarms, work permits, isolations, interventions

This sector has the strongest head start, because the naming is already standardised, ISA-5.1 instrument tags, NORSOK Z-DP-002 coding, IEC/ISO 81346 reference designations, CFIHOS handover classes. A large part of the model is already specified and in daily use; DataHub mirrors it rather than replacing it. Standards →

If a capital project is in flight, capital project handover is the cheapest complete model you will ever get.


A grid operator

LayerModelled as
AssetsGenerating units, substations, battery storage, wind farms
FunctionsBalancing, delivery, load management
Business knowledgeTariffs, power purchase agreements, regulated reporting obligations
Time seriesSCADA telemetry
EventsGrid events, trips, switching operations

What it makes possible. "Which assets contributed to yesterday's capacity shortfall, and what events were raised against them?" becomes one query. The hourly emissions figure submitted to the regulator is reproducible today, and once lineage ships it traces back to individual generating units with data-quality flags attached, an auditable artefact rather than a claim.

A shortfall is a number. Which assets made it is a traversal.One afternoon, one axis: the gap on top, and what was happening underneath it.Demand against delivereddemanddeliveredthe shortfallWhat the model says was underneath itGT-2tripped 13:58, back at 18:20GT-5derated to 60% on ambientBESS-1fully discharged 15:2012:0015:0018:0021:00
The gap in the top chart is what a shortfall report contains. The rows underneath are the same afternoon read off the model: which units were unavailable, for how long, and what was raised against each. An outage and a derate cost different amounts, which is why they are drawn differently.

Starting play: audit-ready reporting, because the regulated submission already exists and already costs real effort.


A municipal water utility

LayerModelled as
AssetsTreatment plants, pumping stations, network segments
FunctionsTreatment stages, delivery zones
Business knowledgeThe regulatory framework each site reports under
Time seriesFlow, pressure, quality parameters
EventsExcursions, maintenance, sampling

What it makes possible. When an effluent reading goes out of spec, the team traverses from the excursion to the contributing treatment stages to the upstream events and maintenance history. Compliance submissions stop being spreadsheets assembled by hand and become queries against the model, each number reproducible, and carrying its full audit trail once lineage lands.

The reading that fails is downstream, and lateSame morning, five stages. What the model adds is the chain and the delay between them.IntakePrimary settlingAerationClarifierOutfallblower B swapped, 06:10sludge blanket rising, 09:30turbidity above limit, 11:40five and a half hours, and three stages, between cause and reading06:0008:0010:0012:0014:00Without the chain, the excursion belongs to the outfall, which is the one place it did not come from.
The excursion is recorded at the outfall, which is the one stage it did not come from. The model supplies both missing pieces at once: the chain of stages upstream of it, and the hours between the cause and the reading that finally failed.

Starting play: faster investigations, then audit-ready reporting once the model covers the treatment train.


A manufacturer

LayerModelled as
AssetsLines, stations, machines, quality instruments
FunctionsProduction steps, quality gates, changeovers
Business knowledgeOEE targets, product specifications, customer commitments
Time seriesThroughput, cycle time, quality signals, energy per unit
EventsStoppages, defects, changeovers, maintenance

What it makes possible. Traceability from a defect back through the stations, settings and material that produced it. Energy per unit compared across lines, with the differences attributable to specific equipment rather than assumed. Downtime attributed to cause rather than to a category somebody picked from a dropdown.

The same defects, plotted against what made themA defect count says there is a problem. Attached to a station and a material lot, the dots say whose.material lot BMon AMon BTue ATue BWed AWed BST-1ST-2ST-3ST-4ST-5one station, one lotOn a weekly report these are one defect rate. Against the station that made each unit, and the lot itcame from, they are a question with an answer.
The same defects, twice over. Counted by week they are a rate somebody has to explain. Attached to the station that made each unit and the lot the material came from, they land in one row and two shifts, and the explanation is already in the picture.

Starting play: data liberation across the MES, maintenance system and energy meters, the fastest route to a visible win on a shop floor.


An IT operations team

LayerModelled as
AssetsServers, storage arrays, switches, virtual machines
FunctionsThe services and pipelines they host
Business knowledgeService levels, business criticality, ownership
Time seriesLatency, utilisation, error rates
EventsDeploys, alerts, incidents

What it makes possible. When a service degrades, the investigation traverses the graph: service, to host, to the array showing elevated error rates, to the switch that dropped packets, to the deploy that went out twenty minutes earlier. "What is the blast radius of taking down this switch?" is a relationship query rather than tribal knowledge.

Twenty minutes apart, and four hops awayThe chart says when. The chain says what, and it is the same walk every time.Checkout latencydeploy 918, 14:00latency steps up, 14:20Checkoutdegradedhost-14hosting itarray-3errors climbingsw-04dropping packetsdeploy 918went out at 14:00start at the symptomEvery hop is a relationship somebody recorded once. The alternative is four people in a call, eachholding one hop of it in their head.
Two facts that are useless apart. The chart knows something changed at 14:20 but not why; the graph knows what the service depends on but not when. Together they close the question, and the walk is the same one every time: start at the symptom and step outward.

This one is not hypothetical, IntelliStream runs DataHub on its own infrastructure this way.

Starting play: faster investigations. The data is already accessible and the relationships are already known, so the model comes together unusually fast.


A defence organisation

Defence is the clearest case of an operation judged on availability rather than output. The questions are the ones any fleet operator asks, with the consequences of a wrong answer rearranged, and there are several distinct applications sitting on one model.

LayerModelled as
AssetsPlatforms (vessels, vehicles, aircraft, radar and communications installations), their sub-systems, and the bases and depots supporting them
FunctionsSustainment, maintenance and overhaul cycles, supply and resupply, training and trials
Business knowledgeUnit ownership, readiness definitions, maintenance doctrine, support contracts
Time seriesPlatform and sub-system telemetry, engine and generator hours, fuel and power consumption, conditions recorded during trials
EventsFaults, defects and rectifications, sorties and sailings, exercises, inspections, part replacements

What it makes possible, in five directions that share the same model:

  • Readiness computed rather than assembled. How much of a fleet is available now, and what is holding the remainder back, taken from maintenance events, usage hours and spares on hand rather than from a week of staff work reconciling returns. The value is not the number, it is that it can be recomputed on the morning somebody needs it. The same question, answered by an agent →
  • Sustainment against usage instead of the calendar. Military platforms are used in bursts. A hull, an airframe or a vehicle that spent the quarter alongside has not aged like one that deployed, and once hours, cycles and conditions hang off the platform itself, maintenance can be planned against what was actually done to it. Condition monitoring →
  • Spares, and where the supply is thin. Which parts each platform consumes under which usage, so demand comes from the fleet's real profile rather than from an average. The same model answers the uncomfortable version of the question: which critical items depend on a single supplier, a single factory or a single shipping lane. Why that question got harder →
  • The estate, not only the platforms. Bases are industrial sites: power and backup generation, fuel storage, water, cooling, runways and quaysides. They fail the way any facility fails, and they carry exactly the same monitoring case as a processing plant.
  • Trials and exercise data that stays comparable. Telemetry linked to the configuration it ran with and the conditions it ran in, so this year's result can be set against last year's instead of being a folder of files somebody has to interpret. Reproducibility is what turns exercise data into evidence, and it is also what makes a digital twin usable for rehearsal rather than decoration.
Two identical vehicles. One of them deployed.A calendar cannot tell them apart. Hours, cycles and conditions can.Cumulative running hours600 h1200 hVehicle A, deployedVehicle B, in reserveA crosses it hereand again herecalendar service, four times, both vehicles, whatever they didA was serviced on neither occasion that mattered, B four times without needing it, and on paperthey are the same vehicle with the same maintenance record.
Cumulative hours, because that is the quantity the decision lives in. Both vehicles were serviced on the same four dates; only one of them crossed the interval that actually governs wear, and it crossed it twice. On a calendar-based record the two are indistinguishable.

On sovereignty and separation, stated plainly. The platform is open source and runs on infrastructure you control, which is normally the first requirement in this sector. Separation between classification domains, though, is a deployment boundary rather than a data set one: in the reference stack every tenant shares one graph, so material that must not mix belongs in separate deployments. What the shipped deployment actually isolates →

Starting play: condition monitoring on one platform class, or data liberation where the maintenance and supply systems already disagree with each other.


A plant that just wants its own data back

Not every deployment starts with an ontology. Sometimes the value is plain data liberation.

The data sits in a historian, a maintenance system, a work-permit system and an ERP, and each guards it behind a data model that takes vendor training to query. DataHub liberates it by simplifying it: whatever arrives from a silo becomes one of a few primitives. A work permit opening or closing, a new purchase order, a sensor alarm, a system state change, all become events. Signals become time series. The things they describe become resources.

Liberation simplifies, it does not copyWhatever a silo sends becomes one of a few primitives, without copying its schema.What the silos sendWhat you read, one interfaceMaintenance systemequipment master recordHistoriantag 21-PT-1034.PVWork-permit systempermit 4471, hot workERPpurchase order 4500171Resourcesthe things it all concernsTime seriesmeasurements over timeEventsthings that happened, at a timeOne interface reads all of it, and nobody has to learn a vendor schema again.No new cage: the model is tiny and documented, and data leaves in open formats.
Four vendor record shapes, three primitives, one interface. Liberation is a simplification rather than a copy: the historian tag, the permit and the purchase order keep their meaning while losing the schema you had to learn to read them.

You never need to understand a source system's schema to use its data, and there is one interface for querying and reading all of it. Lineage, the knowledge graph and everything above can be layered on later. Having every silo readable through one simple model is worth doing by itself, and because the platform is open source, the liberated data has not simply moved into a newer cage.


The common thread

Same platform, same building blocks, no vertical editions. What changes between these six is the vocabulary in the model, and that vocabulary comes from your own people rather than from a taxonomy anybody imposed on you. Each of the six, grown to whatever size its first question needed, is a digital twin of a different operation.

The starting plays differ, but the sequence rarely does: liberate the data, model one area, answer one question, then widen.

Go deeper