GeoBusinessIQGeoBusinessIQ

Manufacturing analytics: joining data that was never designed to be joined

What this answers

Can we connect a quality result back to the machine, batch, tooling and conditions that produced it, without manual reconciliation?

Analysis in a factory is rarely limited by technique. It is limited by whether a scrap record can be tied to the machine, tool, material lot and shift that produced it, and whether the words used for those things mean the same in every system. Once that is true, useful questions become answerable in an afternoon. While it is not, every investigation begins with a fortnight of manual reconciliation and ends with a conclusion nobody fully trusts.

Written for: manufacturing data analysts, process and quality engineers investigating variation, operations leaders funding analytics work.

Everything rests on a key that survives between systems

The works order number, the batch or serial identity and the equipment identifier are the joins that make cross-system analysis possible, and they are routinely broken. Production splits an order and the child inherits a new number; the laboratory records a batch code with a different prefix; maintenance renumbered its assets during a system change. Add clocks that disagree between machines, shift boundaries recorded differently by different systems, and timestamps stored without a stated time zone, and a genuine correlation becomes invisible. Fixing identity is unfashionable groundwork and it determines whether anything built on top will function.

Definitions belong in one place, not in each report

Ask three departments what scrap means and you will get answers that differ on rework, on material lost in setup, and on parts written off after a customer return. Each department is internally consistent and the figures will never reconcile. The remedy is a small agreed layer where each measure is defined once, with its rule visible, and every report drawing from it. Building this is a governance exercise more than a technical one, since somebody with authority has to adjudicate between departments that both have reasons for their version. Without it, meetings are spent reconciling rather than deciding.

The valuable questions are comparative

Little insight comes from a single aggregate. It comes from contrast: this batch against that one, night shift against day, the same product on two lines, before and after a supplier changed material, one operator's setups against another's. Structuring data so those comparisons are easy, with the grouping attributes attached to every record, is what turns a data store into an investigative tool. It also disciplines the questions, because a comparison forces you to state what you think is different, which is a far more productive starting point than asking a system to reveal something interesting.

A model proposes a hypothesis that the shop still has to test

Statistical models and machine learning applied to process data will find associations, some of them real and some artefacts of how the data was collected. A pattern showing that a defect rises with a particular material lot may reflect the lot, or the machine that lot happened to run on, or the period during which an operator was covering for someone absent. Treat the output as a candidate explanation and test it deliberately on the floor before changing a process. Plants that skip verification eventually implement a change based on a coincidence, and the credibility cost of that lasts a long time.

Analysis dies when nobody owns the follow-up

The characteristic failure is not a bad model, it is a good finding with no route to action. An analyst identifies a recurring loss, presents it, and the meeting agrees it is interesting. Nothing is assigned, the next month brings a new analysis, and after a while the function is regarded as reporting. Attaching each finding to a named owner with a decision date, and reporting back on what changed, converts analysis into an operating practice. It also improves the analysis, because knowing something will be acted on changes the standard of evidence people apply.

Frequently asked questions

How is this different from the dashboards we already have?
A dashboard displays an agreed set of measures to a defined audience so they can see the current position. Analysis is exploratory and retrospective: it joins records that no dashboard was designed around, asks why something happened, and often produces a one-off answer that never becomes a screen. The two need each other. Analysis without display leaves findings stuck with the analyst; display without analysis shows movement in a figure while leaving everyone guessing at the cause.
Do we need a data scientist to get value from process data?
Not at the start. A large share of useful manufacturing analysis is careful counting, sensible grouping and honest comparison, which an engineer who knows the process can do well. Specialist modelling becomes worthwhile when relationships are genuinely multivariate, when volumes exceed what inspection by eye can handle, or when prediction is the goal. Hiring specialists before the identity and definition groundwork is done tends to produce sophisticated work built on data that will not support the conclusions.
How long should we keep detailed production data?
Long enough to cover the questions you will actually ask, which is usually longer than an IT retention default and shorter than forever. Warranty periods, customer investigation windows, regulatory retention obligations and the seasonality of your process are the guides. A common approach keeps full detail for a recent window, then thins older data to summaries and to the events that matter, so a multi-year trend remains available without holding every raw reading indefinitely.

Data limitations

  • Manufacturing figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no factory costs, production volumes, yields, cycle times, tooling prices or capacity data and does not estimate them — every result reflects only the figures you enter.

Explore the graph

Sources

  • National Institute of Standards and Technology NIST (accessed )
    Covers: Measurement science, manufacturing technology research, cybersecurity frameworks, and industrial standards support.
    Does not cover: Certification of products, endorsement of vendors, or costs for any specific implementation.
    Why it matters: A United States federal research institute whose public material covers measurement, manufacturing technology and control-system security.
    Review cadence: annual
  • OECD OECD — economic and tax statistics (accessed ; reviewed )
    Covers: Comparable corporate tax, statutory rate, and economic indicators across member and partner economies.
    Does not cover: Effective tax rates, deductions and incentives, local surtaxes, and personal residency rules.
    Why it matters: Used as a cross-country baseline to sanity-check rates against primary tax-authority figures.
    Review cadence: Annual, plus on major statutory changes.

Educational and operational information only — not legal, engineering, safety, customs, tax, or financial advice. Requirements vary by jurisdiction, product, process, and contract; confirm with the relevant authority or a qualified professional before acting.

Last updated: