Print

Is Your Data Ready for AI?

Manufacturers often have large amounts of data from their ERP systems, equipment, maintenance reports, and quality systems. Yet these data are not automatically usable by artificial intelligence. This article explains how to assess their relevance, consistency, context, and representativeness before starting a project.

Security Aerospace Manufacturing Medical Devices Life sciences
Franck Boulbes
Date  June 2026

Manufacturing companies have been accumulating data for years.

Their ERP systems retain histories of orders, inventory, and operations. Equipment continuously generates measurements. Maintenance teams record downtime and interventions. Quality departments document defects, while supervisors produce their own reports and tracking tables.

Faced with this abundance of information, one conclusion may seem natural: the company must surely have everything it needs to develop an artificial intelligence solution.

The reality is more nuanced.

The real question is not simply:

Do we have data?

It is rather:

Do our data truly tell us what is happening in our operations?

Data can be structured enough to support day-to-day operations without being precise, consistent, or contextualized enough to train a model. They may indicate that an event occurred without explaining why. They may be distributed across several systems or documented differently by different teams. Some essential knowledge may never be recorded because it remains in the minds of experienced employees.

The question, therefore, is not simply whether your company has data. It is whether it has the right data, in an appropriate format and with enough context to address the problem at hand.

Summary

Having a large volume of data does not mean it is ready for artificial intelligence. To be truly usable, data must be relevant to the problem, complete, consistent, contextualized, and associated with a reference outcome. A history of failures, defects, or interventions loses much of its value if the causes, conditions, and corrective actions are not documented. Before developing a model, you therefore need to assess what the data actually tell you and address the gaps that could limit the reliability of the solution.

Having data does not mean having a use case

Before assessing the quality of a dataset, you need to know what you are trying to accomplish.

A company may have twenty years of historical data on its customers, orders, or equipment. That history does not automatically become useful for AI simply because there is a lot of it.

To determine its value, you first need to formulate a specific question.

Are you trying to predict downtime? Detect a defect? Recommend an intervention? Reduce material losses? Plan an operation more effectively? Make it easier to search technical documentation?

Without a defined question, the analysis is like searching for a needle without knowing its shape or location. The volume of data increases the space to explore, but it does not indicate what should be found there.

This distinction is important because the same data may be useful for one objective and insufficient for another.

Production history might make it possible to calculate the number of downtime events per month without making it possible to predict their causes. A defect log might meet traceability needs without containing the details required for a visual recognition system.

The first step in an AI project is therefore to connect the data to a business problem:

Which decision, prediction, detection, or recommendation do we want to improve?

Only once that question is defined is it possible to determine what information is required.

Data designed for an ERP system are not necessarily designed for AI

Management systems are developed to meet specific objectives.

An ERP system must, among other things, make it possible to record transactions, track orders, manage inventory, and coordinate operations. It was not necessarily designed to explain in detail the causes of a failure or the exact conditions in which a defect appeared.

Data can therefore be perfectly valid for their original purpose while still being insufficient for an AI project.

Consider a downtime code. For a monthly report, it may be enough to know that a piece of equipment stopped for 45 minutes because of a “mechanical problem.” For a diagnostic assistant, that category is probably too broad.

To develop a diagnostic assistant, the system would instead need to connect several pieces of information: the observed symptoms, measurements taken before the stoppage, machine parameters, production conditions, the component involved, the confirmed cause, the intervention performed, and the outcome.

If this information has never been collected, it cannot be reconstructed automatically from the duration of the stoppage alone.

Artificial intelligence can identify relationships, recognize patterns, and produce recommendations from the available data. However, it cannot reliably infer a cause that has never been observed or documented.

Consistency in data entry is a determining factor

Manufacturing data are created by systems, but also by people. Their quality therefore depends on day-to-day documentation practices.

Within the same company, observed defects may be documented in detail by some teams, while others group them into a more general category. One shift may document the defect type, the machine involved, and the corrective action, while another routinely selects the first option offered in the form to complete data entry more quickly.

The data exist in both cases, but they do not have the same precision or analytical value.

These differences can be difficult to detect in standard reports. A generic category may appear valid in a database even though it actually groups several different phenomena together. The problem becomes visible only when you try to use these data to train a model.

The system may then learn the teams’ data-entry habits rather than the actual characteristics of the process. It may also produce different results depending on the shift, plant, or person who documented the event.

Before launching a project, you therefore need to examine not only the fields present in the database, but also how they are completed:

  • Are the categories understood in the same way by everyone?
  • Are some fields regularly left blank?
  • Are generic values used by default?
  • Do practices vary by shift, site, or team?
  • Is information entered at the time of the event or several hours later?
  • Are there checks in place to identify inconsistencies?

The quality of data does not depend solely on their presence. It also depends on how consistently they represent reality.

A failure with no documented cause teaches the model very little

Machine downtime data clearly illustrate the difference between an operational history and an asset that can be used for AI.

A company may know precisely when each stoppage began, how long it lasted, and which equipment was involved. This information can be used to measure machine availability and overall equipment effectiveness. However, it is not necessarily sufficient to develop a diagnostic or predictive maintenance solution.

For a model to establish useful connections, the event must be associated with context and an outcome.

Ideally, each incident should make it possible to connect several elements:

  1. the conditions preceding the event;
  2. the observed symptom or defect;
  3. the identified cause;
  4. the intervention performed;
  5. whether or not normal operation was restored.

If the database only indicates that a machine stopped, the model may learn to recognize certain signals preceding a stoppage. It will be much more difficult for it to distinguish among the different possible causes or recommend an appropriate action.

Documenting the intervention is equally important. Two stoppages may show similar symptoms while requiring different corrective actions. Without knowing what action was taken and what the outcome was, it is impossible to determine which intervention actually resolved the problem.

The issue, therefore, is not to keep accumulating more downtime data. It is to better document the full chain from observation to resolution.

To explore this topic further, read the article Unplanned downtime: the hidden cost of diagnostic time.

Context turns data into useful information

An image, measurement, or note takes on its full meaning only when it is placed in context.

A photograph of a defective part may seem very useful for developing an inspection system. But what exactly does it represent? Was the defect confirmed by a specialist? Was the part rejected? Is it an acceptable variation? What were the material, lighting, batch, machine, and production parameters?

Without this information, a collection of images can be difficult to interpret. Some images may represent the same defect under different conditions. Others may show several defects at the same time or include variations that do not compromise part conformity.

The same principle applies to machine data. A temperature variation may indicate an anomaly in one context and be perfectly normal in another. To interpret it, you need to know, among other things, the product being manufactured, the stage of the cycle, the production speed, and the settings applied.

Context makes it possible to understand what happened, under what conditions the event occurred, and what the verdict or outcome was. The more precisely these elements are documented, the more the data can support relevant analysis.

The Bridgestone example: data connected to a specific problem

The project carried out by Bridgestone Canada at its Joliette plant illustrates why data must be connected to a specific operational problem to become truly useful.

The company was seeking to reduce production stoppages caused by imperfect joints between layers of rubber during tire assembly. The challenge was complex: stoppages occurred frequently, raw material varied according to conditions such as temperature or humidity, and there were many manufacturing parameters to adjust.

In this context, production history alone was not enough. To produce useful recommendations, the data had to make it possible to understand the relationships between production conditions, manufacturing parameters, material variations, and the quality of the resulting joints.

Using more than a year of historical data made it possible to analyze these relationships and identify the variables influencing quality. Predictive models and time-series analyses were then used to design a recommendation system capable of guiding operators in their adjustments.

This case shows that data become usable when they do more than simply record that a stoppage or defect occurred. They must also help explain why it occurred, what factors influence it, and how teams can act to prevent it.

Quality can matter more than quantity

For a long time, AI projects were often associated with the need to accumulate very large quantities of data. This perception can discourage some manufacturing companies, particularly when the defects of interest are rare or critical events occur infrequently.

Technological advances now make it possible to consider some projects with more limited datasets. This does not mean that a small number of examples is always sufficient. Requirements vary depending on the problem, the diversity of conditions, the desired level of accuracy, and the method used.

Rather, it means that a smaller set of properly documented data can sometimes be more useful than a large quantity of ambiguous examples.

One hundred images whose context, defect, and verdict have been validated by an expert can provide a stronger foundation than thousands of images without reliable labels. Similarly, a shorter maintenance history that clearly connects each symptom to its cause and resolution may be preferable to several years of incomplete notes.

Context matters: it is not enough to provide images or events to the system; you must also document what they represent and the conditions under which they were observed.

However, this idea should not be turned into a universal rule. The required volume cannot be determined before examining the use case, the variability of the data, and the expected level of performance.

The right question, therefore, is not:

Do we have enough data?

It is rather:

Do we have enough representative and reliable data to answer our question with the required level of confidence?

Parallel Excel files often reveal important data

In many companies, some operational information remains outside official systems. It can be found in Excel files, local documents, personal notes, or tables created by employees to meet needs that existing tools do not fully address.

These files can reveal essential information that is not always found in official systems: indicators actually used on the shop floor, manual steps between two systems, categories created by teams, exceptions that existing tools cannot handle, or decisions that still rely on an employee’s interpretation.

The proliferation of siloed systems, manual re-entry, and parallel files is a common reality in manufacturing environments. Employees then spend time copying, correcting, or reorganizing information so they can do their work.

Before considering centralizing or eliminating these files, you need to understand why they exist. They may be a symptom of an integration problem, but they may also contain some of the context needed for a future AI project.

However, they have significant limitations. Their structure may vary from one person to another. Multiple versions may proliferate. Some columns may be interpreted differently, and formulas can be modified without traceability.

The goal, therefore, is not to automatically feed every file into a model. It is to determine what useful information the files contain and how to integrate it into a more consistent data process.

Undocumented knowledge is also a data gap

Not all useful data are already stored in a database.

In many plants, experienced employees recognize certain problems based on a sound, vibration, smell, appearance, or combination of signals that are difficult to formalize. They know which checks to perform first and which interventions are most likely to work.

This expertise is a form of data, even if it has not yet been digitized.

When a failure is resolved without documenting the reasoning, cause, and intervention, the company loses an opportunity to enrich its organizational memory. The problem is resolved in the short term, but the information remains inaccessible to other employees and to any future diagnostic support system.

Preparing the data may therefore require adapting work practices. A technician could, for example, record more systematically:

  • the symptoms observed;
  • the hypotheses considered;
  • the checks performed;
  • the confirmed cause;
  • the repair performed;
  • the outcome obtained.

Technology can facilitate this capture through a simplified form, a tablet, annotated photos, or voice notes. However, the tool chosen matters less than the consistency and quality of the documentation.

When all expertise remains “between the ears” of a few people, neither new employees nor AI systems can truly benefit from it.

Data analytics lays the groundwork for AI

Data analytics plays a central role in the use of AI by businesses in Canada. In the second quarter of 2026, 19.2% of Canadian businesses reported using AI to produce goods or provide services during the previous 12 months. Among these businesses, data analytics was the most frequently cited application, at 36.6%.

Another Statistics Canada study shows that businesses already using data analytics were 15 percentage points more likely to adopt AI than those that were not.

This association does not demonstrate that analytics alone causes AI adoption. It nevertheless suggests that data-related skills, practices, and infrastructure are important complementary capabilities.

In other words, preparing data is not an administrative step separate from the AI project. It is an integral part of the company’s ability to use this technology.

Organizations that already know how to locate their data, verify their quality, connect them, and use them to support decisions generally have a stronger foundation for exploring advanced applications.

How can you assess whether your data are ready?

Data readiness is not a binary outcome. A company may be ready for one use case and not for another. 

An assessment should begin with a representative sample of the data related to the problem. It is rarely necessary to immediately examine the entire history. The initial objective is instead to determine what the data actually make it possible to understand.

1. Do the relevant data exist?

First, verify whether the variables directly related to the problem are available.

To predict a failure, does the company have measurements from before the event? To recognize a defect, does it have representative images? To recommend an intervention, are the causes and corrective actions documented?

A large quantity of peripheral data does not compensate for the absence of essential information.

2. Are the data sufficiently complete?

Some fields may exist in the system without being filled in regularly. You need to measure the proportion of missing values and determine whether the omissions are random or related to certain teams, periods, or situations.

Missing data during the most complex events can create significant bias.

3. Are they entered consistently?

Categories, units, and conventions must be examined. Does the same defect have several names? Are measurements always recorded in the same unit? Can timestamps from different systems be synchronized?

Normalization rules can correct some discrepancies, but they cannot always reconstruct lost meaning.

4. Is the context preserved?

The data must be linkable to the product, batch, machine, shift, process parameters, and, where relevant, environmental conditions.

Without this association, it becomes difficult to distinguish a normal variation from an anomaly.

5. Is there a reference outcome?

For many projects, you need to know the “right answer” used to evaluate the model. This may be an inspector’s verdict, the confirmed cause of a failure, or the actual outcome of an intervention.

If this reference does not exist, you may need to start by creating it.

6. Are the data accessible and secure?

Information may be distributed among the ERP system, machines, maintenance software, Excel files, and control systems. You need to determine whether it can be extracted and brought together without disrupting operations or compromising security.

Before connecting a solution to an ERP, MES, or production equipment, it may be preferable to work in a separate environment with a data warehouse. This approach makes it possible to explore the potential of a use case without compromising day-to-day operations.

7. Are the data representative of real conditions?

The data must cover the diversity the solution will encounter once deployed: different products, settings, equipment, shifts, seasons, or lighting conditions.

A model trained only on the most common situations may fail precisely in the rare cases where its assistance would be most useful.

What should you do when the data are not ready?

Discovering that the data are insufficient does not necessarily mean the project should be abandoned. On the contrary, that conclusion can provide a concrete roadmap.

The company can begin by improving data collection within a limited scope: one machine, one family of defects, or one category of failures. Forms can be simplified to encourage more consistent data entry. Categories can be clarified with users. Certain required fields can be added, and employees can be informed about how the data will be used.

It is also possible to bring together a small group of experts to review and annotate a sample. This work can help determine whether the existing information contains enough value to continue exploring the opportunity.

Data preparation can therefore progress in stages:

  1. define the problem and the information required;
  2. examine a sample of the current data;
  3. identify gaps and inconsistencies;
  4. improve the documentation method;
  5. collect new data consistently;
  6. verify their quality before developing the model.

This approach creates value even if the AI project is not launched immediately. Better-structured data can already make it easier to analyze operations, monitor quality, train employees, and support decision-making.

Read also

Conclusion

Preparing the data means preparing the project

Manufacturing companies do not necessarily need to wait until they have a perfect database before exploring AI. However, they must understand the limitations of the data they have and avoid confusing quantity, quality, and relevance.

A large historical dataset can be an important asset, but only if it makes it possible to connect events to their context, causes, and outcomes. The data must represent operations consistently enough for the observed relationships to be reliable.

Data preparation therefore begins well before a model is trained. It begins with how a failure is documented, how a defect is classified, how an intervention is linked to its outcome, and how employee expertise is preserved.

Before asking what AI can discover in your data, ask yourself a more fundamental question:

Do our data truly tell us what is happening in our operations?

When the answer is uncertain, a targeted assessment can help determine which information is already usable, which gaps need to be addressed, and which practices need to be put in place.

Luqia’s AI experts can help you assess the quality and relevance of your data in relation to a specific manufacturing need, then determine realistic next steps for your project.

Assess your data maturity - Contact us!

Statistical sources

Statistics Canada, “Analysis of the use of artificial intelligence by businesses in Canada, second quarter of 2026,” June 11, 2026: https://www150.statcan.gc.ca/n1/pub/11-621-m/11-621-m2026010-fra.htm

Statistics Canada, “Adoption of artificial intelligence and productivity in Canadian businesses,” April 22, 2026: https://www150.statcan.gc.ca/n1/pub/36-28-0001/2026004/article/00002-fra.htm

About the author

Franck Boulbes

Director, AI Business – Manufacturing & Industry

Franck Boulbes is an expert in artificial intelligence, digital transformation, and industrial technologies. With more than 20 years of experience in electronic engineering, industrial computing, and technological innovation, he has also supported more than 100 companies in their digital transformation and automation projects. An entrepreneur for several years within the startup ecosystem, he holds an engineering degree as well as a master’s degree in business and technology from Université Savoie Mont Blanc, complemented by training in financial and management accounting at McGill University.

View LinkedIn Profile

Subscribe to the blog

Stay tuned for our latest articles.

Contact