"Our OT and IT data are separate, and we want to integrate them." This is the starting point for many digital transformation projects. The common next step is to select a platform and import data from both sides into the same location.
The data indeed enters the same database. However, one or two years later, the original question—"Which upstream process parameters are related to this batch of quality anomalies?"—often remains unanswered.
The problem is: there is more than one type of silo, and changing systems only solves one of them.
Three Types of Silos
1. Physical Silos
Data resides on different machines, in different network segments, with firewalls or even physical isolation between OT networks and office networks. This is the easiest type to identify and the only one that can be solved by implementation.
Most integration projects address this layer. After completion, the data is in the same place, and the problem appears to be solved.
2. Semantic Silos
Both sides have different definitions for the same event.
A "batch" in the production system refers to the input batch; a "batch" in the quality control system refers to the inspection batch. The boundaries between the two do not necessarily align. "Work order completion" in ERP is the point of financial recognition, while "finished" on the shop floor is the actual time of off-line, which could differ by several hours.
When these two types of "batches" and two types of "completions" are placed in the same table with the same column name, the discrepancy disappears into visual orderliness. Subsequently, all analyses based on this table inherit this error.
This layer cannot be solved by migration. After migration, you have merely placed two inconsistent definitions closer together.
3. Responsibility Silos
This layer is discussed the least. After data integration, who has the authority to define the meaning of a column? Who is responsible for updating definitions when process changes occur? Who decides under what circumstances a certain piece of data should not be used for judgment?
If these questions remain unanswered, then even if the first two layers are resolved, the trustworthiness of the data will still degrade over time—because no one is responsible for maintaining it.
Why "Just Import It First" Usually Isn't Cost-Effective
Collecting all data first and then processing semantics sounds like a logical sequence. The practical difficulty is: context can be lost during migration, and once lost, it cannot be recovered.
When a numerical value leaves its original system, its sampling method, time baseline, equipment status at the time, and who entered it under what circumstances—if these are not carried along, then only the number itself remains. To reconstruct it, one must go back and ask people; the precondition for getting an answer is that the person is still there and remembers.
This is also why many data lake projects encounter the same dilemma in their second year: there is a lot of data, but no one dares to use it for important decisions.
Another Approach
Compared to "move everything over first," another approach is to start from a specific problem, connect only the data required for that problem, and bring the context along at the moment of connection.
For example, first select a real traceability requirement: "The defect rate for a certain model is high in a specific shift; we want to know which process parameters it is related to." To answer this, you need to know which data sources, which columns, who is responsible for their definitions, and how time is aligned. The scope is small, but each item is complete.
After answering this one question, you will get two things: a usable answer, and a small, clear data foundation with well-defined semantics and responsibilities. The next question can build upon it.
This is slower than migrating everything at once and then organizing it, but each step can be validated, and it will not accumulate unseen errors.
X·Neurons' Role and Boundaries
Our chosen role is to connect rather than replace: not starting with a complete replacement of PLCs, MES, ERP, or shop floor programs, but building governable connections on top of existing investments, allowing data to enter decision-making and actions with reason and responsibility.
Currently, all X·Neurons capabilities are in a candidate state, describing expected behavior in joint validation, not released, directly purchasable functions. Enterprise Intelligence is our working theory, not an established science; its boundaries are written on the methodology page.
Limitations that need to be clarified: semantic definitions cannot be generated automatically, and responsibility attribution cannot be determined by the system. These two things still require people within the organization to make choices. What tools can do is ensure that once these choices are made, they can be clearly recorded, consistently applied, and leave a trace when changes occur.
A Judgment Question
If you are evaluating integration solutions, there is a question that can quickly reveal which layer it addresses:
After implementation, when the definition of a certain column needs to be modified, what is the process? Who approves it? How is existing data marked?
If the other party's answer stops at "it can be changed," then it addresses the first layer. If the answer includes who is responsible, how traces are left, and how existing data is handled, then it has at least considered the third layer.