Lekker Code

Case study · Lekker Code

Field service in wind energy

From network drive to reliable answer: AI assistant for wind turbine service technicians

Wind farm operator · Architecture assessment and target architecture

Starting point

Service technicians work in the tower or the substation, often alone and under time pressure. They need to know what a fault code means on this turbine type and which version of an instruction applies today. The documentation lives on a network drive that has grown over roughly two decades: scanned wiring diagrams without a text layer, HTML manuals with tens of thousands of pages, status lists and photos, in total several hundred thousand files from more than thirty manufacturer folders. A RAG assistant with hybrid search, a reranker and page-level citations was already running on the pilot data. The open question was whether the approach would carry the real document base.

Analysis: five break points

We surveyed the document base and reviewed the ingestion code. The existing path sent every page through a large language model before anything was filtered. Extrapolated, a full run meant several million model calls, costs in the mid four-figure euro range and a runtime of days to weeks. Three of the five break points cost money, two put safety in the field at risk.

  • Format: The pipeline read PDFs only. For the largest manufacturer, though, the actual documentation is an HTML manual system, and what got indexed instead were cryptically named PDF attachments.
  • Text yield: Text recognition was switched off. A fifth to a third of all pages contain no readable text and still went through the model. A single wiring diagram with 258 text-free pages produced several hundred model calls and as many empty index entries.
  • Duplicates and metadata: Duplicates were not detected, about half of the files for one manufacturer. Turbine model, controller type and revision were left for the model to guess, although they are almost always in the file path.
  • Discarded documents: Anything that did not fit one of four categories was deleted. That hit safety data sheets, HSE warnings and status code sheets of all things.
  • Revisions: Competing revisions existed side by side, and the pipeline had no notion of supersession. A withdrawn version could reach the technician as a valid answer.

Solution

The target architecture does not replace the platform, it changes the order. A deterministic stage that works without a model goes in front of the language model. It detects duplicates, reads metadata from path and file name, resolves revisions, switches on text recognition only where needed and routes by document type. Documents it cannot classify go into a review queue rather than the bin, and status codes are looked up exactly in a table. SharePoint replaces the network drive as the source, which means incremental indexing and metadata maintained by the business. We ruled out a generic Copilot solution. It loses filtering by model, controller and revision and moves confidential manufacturer documentation out of sovereign EU operation.

Result

The client has a decision paper with a prioritised plan. The first steps do not depend on the document source and can start right away: a sample run that turns the forecast into a measurement, status code tables as the cheapest safety gain, the deterministic stage and the review queue. A smaller overlap window for text fragments alone cuts model costs by up to two thirds. The assistant stays advisory. It shows the cited passage in the manual, and the person on site makes the decision.

Technologies

  • RAG
  • Hybrid search
  • Reranker
  • OCR
  • SharePoint
  • EU operation
Similar situation? Start a conversation

Source: published project account by Lekker Code. Metrics apply only to this case.