Machines That Understand Maintenance: How LLMs and RAG Are Rewriting Diagnostics, Troubleshooting, and Field Support

The sentient computer HAL 9000 in Stanley Kubrick’s 1968 film 2001: A Space Odyssey remains one of cinema’s most memorable visions of artificial intelligence. HAL knew the spacecraft around it, monitored systems, interpreted events, related information to the operational world in which it was acting, and relayed information to the crew.
They are not HALS, but today’s machines can talk convincingly about engineering. A Large Language Model (LLM) can explain cavitation, bearing defects, or motor overheating and produce a plausible troubleshooting sequence. It gets more interesting when we move from asking our LLM “What causes high vibration in a centrifugal pump?” to asking it “Why is P-204 vibrating today?”
Answering the first requires general knowledge about pumps. Answering the second requires more specific knowledge of a particular physical asset: its configuration, current operating state, previous failures and interventions, alarms, measurements, and manufacturer recommendations. Knowing about pumps is not the same as knowing this pump.
To be useful to maintenance, our LLM must have the ability to determine what knowledge and evidence apply to a particular asset, in a particular condition, at a particular moment, and what actions legitimately follow from that knowledge and evidence. The challenge is to get it to relate knowledge and evidence to this machine, in this condition, now.
The knowledge we already have. Maintenance organizations have spent decades becoming digital, accumulating condition monitoring systems, CMMS/EAM databases, and predictive models. Yet when a difficult failure occurs, one person still searches a manual, another hunts for an old work order, and eventually somebody asks, “Who worked on this machine last time?”
The problem is fragmentation. Manuals, CMMS, condition monitoring, FMEA/RCM, alarms, procedures, and human experience collectively form an enormous technical memory, but the engineer must still reconstruct the scenario. Organizations do not simply need more data; they need better ways to convert distributed technical memory into decision-relevant evidence when it matters.
Generative AI offers a way to access this memory, allowing technicians to interrogate technical knowledge in ordinary language. The trap is that model fluency can make us confuse the ability to formulate an answer with the possession of evidence needed to defend it.
Fluency is not evidence. An LLM may explain bearing lubrication perfectly without knowing which lubricant was actually applied. It may generate a beautifully structured maintenance instruction while being unaware that the OEM revised the procedure six months ago. It may provide a tightening torque that looks entirely plausible but has no defensible source
LLMs are exceptionally good at producing plausible language, but linguistic plausibility is not engineering validity. Once generated, language influences physical decisions; hallucination stops being a natural language processing quality issue and becomes a question of decision integrity, reliability, and operational risk.
This is where Retrieval-Augmented Generation (RAG) becomes important. Rather than relying exclusively on information encoded inside the model, RAG retrieves relevant material from authorized sources before constructing an answer. In maintenance, those sources may include manuals, procedures, work orders, inspections, historical failures, and standards.
But retrieval alone does not amount to understanding. Consider our vibrating P-204. A general LLM can explain why pumps vibrate. A RAG system can retrieve what the OEM says about this pump family. A contextual system can determine whether P-204 itself has experienced this pattern before and under what conditions. An investigative system can ask which measurement would best discriminate between competing hypotheses. A trustworthy system can then take the accumulated evidence and determine whether it is sufficient for a recommendation and whether it has authority to do anything with that recommendation.
Together, these capabilities form what I call the Maintenance Understanding Stack: Language, Retrieval, Context, Investigation, and Authority.
Language covers general engineering knowledge; retrieval accesses technical memory; context determines what applies to P-204 now; investigation identifies what evidence is needed next; authority asks what the evidence justifies and what the system may do.
This is not an AI maturity staircase climbing towards autonomy. The levels represent increasingly demanding claims about knowledge and legitimate action. A system may retrieve the correct manual without establishing that it applies to P-204; an agent may investigate effectively while still lacking enough evidence for a decision, and a justified recommendation may remain outside its authority.
Retrieval is not context. Investigation is not authority. Capability does not imply permission.
Maintenance knowledge is relational. Connecting an LLM to thousands of PDFs is not enough because maintenance knowledge is relational, not merely textual. A useful chain looks like asset • condition • symptom • evidence • failure hypothesis • action. Semantic retrieval can find related material, while exact identifiers, alarm codes, and structured relationships help determine what actually belongs to the case at hand.
For the engineer, the requirement is simple: a recommendation must be traceable backwards through an evidentiary chain: recommendation • reasoning • evidence • original source or measurement. A fluent explanation is not enough. If the system recommends checking the coupling of P-204, the engineer should be able to inspect the spectrum, procedure, inspection result, or work order supporting that recommendation. And if several causes remain plausible, the system should identify what observation would discriminate among them, not jump from similarity to diagnosis.

Sometimes the most intelligent answer is “I don’t know”. Suppose the technician asks whether P-204 can continue operating until a planned shutdown the following morning. The system retrieves the manual, previous interventions, and similar cases, but discovers that the latest vibration spectrum is 12 hours old ,and the current operating load is unavailable.
A conventional chatbot is optimized to answer the question of shutdown. But a trustworthy maintenance system should recognize a different problem: answer confidence and evidence sufficiency are not the same thing. The relevant question is not “Can the model answer?” but “Does the evidence available now justify the decision being requested?”
The correct response may be to withhold the continued-operation recommendation, explain why it cannot yet be justified, and request a current spectrum and operating load. Once those observations arrive, the investigation can resume.
I call this engineering-aware abstention: the ability to withhold a recommendation because evidence is insufficient while identifying what additional observation or measurement is required. This is stronger than simply saying “I am uncertain.” A competent industrial assistant should know what is missing, why it matters, and what should be obtained next.
In maintenance, intelligence is not only knowing an answer. It is knowing when the evidence is insufficient to answer.
From retrieval to investigation. The next step is agency. The distinction is simple but fundamental: RAG gives the system access to relevant knowledge. Agency allows it to use that knowledge in pursuit of an objective.
To solve the riddle of P-204, this means participating in an evidence-seeking investigation. The abnormal vibration is observed; relevant history and procedures are retrieved; plausible hypotheses are considered; missing or discriminating evidence is identified; a measurement or inspection is requested; the result updates the hypotheses; the system then recommends, abstains, or escalates according to the strength of the case.
The sequence is essentially observe • retrieve • hypothesize • identify missing evidence • obtain evidence • update • recommend / abstain / escalate. What matters is that the AI participates in the process by which the engineering answer is reached. It does not merely produce a better paragraph at the end.
The LLM does not replace vibration analysis, reliability models, CMMS/EAM, FMEA, RCM, or physical inspection. Specialized systems still produce engineering evidence; the role of
LLM/RAG/agentic systems is to connect, contextualize, retrieve, interpret, and orchestrate their outputs around the problem being investigated.
Authority is part of intelligence. A system capable of investigating a problem is not necessarily allowed to act on its conclusion. Capability does not imply permission. Depending on criticality, consequence, uncertainty, evidence quality, and organizational rules, the system may inform • recommend • prepare • request approval • execute within defined authority • escalate. Authority belongs inside the Maintenance Understanding Stack. A technically defensible recommendation does not automatically confer the right to execute it. Nor should human-in-the-loop be treated as an inconvenience that disappears when AI improves. For many industrial decisions, human-governed agency may be the correct final architecture. Humans retain physical verification, judgement, consequence assessment, and accountability; machines assume more of the burden of finding and connecting knowledge.

A model of intelligent machines that is perhaps closer to what maintenance actually needs is KITT, a car with AI in the TV series Knight Rider. KITT was extraordinarily capable: it monitored the vehicle, interpreted situations, warned protagonist Michael Knight, proposed actions, and intervened when necessary. Yet the partnership mattered as much as the intelligence. KITT was not valuable because it eliminated the human from the loop, but because machine capability and human judgement complemented each other. Industrial maintenance may ultimately require something similar: not autonomous machines replacing engineers, but increasingly capable technical partners operating within clearly defined boundaries of evidence and authority.
When the technical memory of the organization becomes active. Today, the engineer often acts as the mediator between maintenance information systems, synthesizing their input. To understand a difficult event, the engineer moves among condition monitoring, historian, CMMS/EAM, inspection records, manuals, and previous interventions, mentally reconstructing a coherent picture from systems that were never designed to reason together.
The deeper opportunity is to create a cognitive integration layer over the technical memory of the plant.
The manuals, measurements, work orders, failure histories, procedures and engineering analyses already exist. Traditionally, that technical memory is passive: it waits until a human knows where to look and how to reconstruct its relevance. An active technical memory can participate in determining what knowledge applies to P-204, what evidence is missing, what should be investigated next, and what the accumulated information actually justifies.
The real disruption of generative AI in maintenance may not be that machines become better at answering questions. It may be that the accumulated technical memory becomes an active participant in maintenance decisions.
Prediction can tell us that something may fail. Specialized diagnostics provide evidence about what is happening. Retrieval connects that evidence with technical memory. Context determines what applies to this asset now. Investigation seeks what is still missing. Authority determines what legitimately follows. Humans retain judgement and assume responsibility for consequences.
Ultimately, industrial maintenance needs something less theatrical and more useful than HAL 9000. It needs a system that knows which asset to focus on, connects events with defensible evidence, recognizes when that evidence is insufficient, understands the limits of its authority, and knows when a human must decide.
At that point, machines will stop merely talking about maintenance and begin, in a meaningful industrial sense, to understand it.
Text: Prof. Diego Galar Figures: Prof. Diego Galar Photo: shutterstock



