Skip to content
24 September 2026

Comparing AI assistants for engineering

Four categories of tools are on the table, and they solve different problems. Which criteria actually decide the question in engineering.

Quick answer

Generic chat assistants, search inside the Microsoft stack, search in a DMS or PLM, and self-built RAG systems each solve a different problem. In engineering the language model is rarely what decides. What decides is whether an answer traces back to file, revision and page, whether the documents stay inside your own legal perimeter, and whether the system reads drawings.

Companies in mechanical engineering looking for an AI tool for their own document base rarely compare product against product. Usually four quite different categories sit side by side, promising the same thing at first glance and delivering something entirely different in daily work. This article sorts the categories, names the criteria the decision actually turns on, and describes a test you can run yourself in about two weeks. The direct, product-by-product comparisons live on the KoAssist comparison page.

What is an AI assistant for engineering usually compared against?

Generic chat assistants. They answer from their training knowledge and from whatever gets uploaded into the conversation. Strong at phrasing, translation and explaining general matters. The break comes as soon as the answer has to come out of your own material: what is not in the chat does not exist for the system, and a statement about an internal guideline cannot be backed up.

Search inside the Microsoft stack. Deeply integrated into the Office world and, where content genuinely lives in SharePoint and Teams, close to where people work. For engineering document bases the fit is different: a large part of the relevant knowledge sits in PDFs, drawings and grown file shares rather than Office files, and how precise the citation gets varies considerably between tools.

Search in a DMS or PLM. The existing system knows metadata, revision states and structures, and is hard to beat there. The question "which bill of materials belongs to which assembly" belongs in that system. The question "what does our design guideline require for this weld" rarely does, because full-text search across document boundaries usually returns no more than a list of hits.

Self-built RAG systems. With off-the-shelf building blocks a prototype comes together quickly. The effort then shifts from building to operating: document preprocessing, a permissions model, handling drawings, keeping models and dependencies current. With a team for it and a narrow use case this works well. Run on the side, it becomes a maintenance topic within a year.

Which criteria decide in engineering?

Not the language model. Models are similar across vendors and keep getting better. The difference comes from everything around them:

Criterion The question it answers
Verifiability Does the answer name file, revision and page, and does one click take me to the passage?
Handling of what it does not know Does the system say "not found", or does it produce something plausible?
Image understanding Are drawings, calculation sheets and scans read, or skipped?
Legal perimeter Where are the documents processed, who are the subprocessors, is training on your data contractually excluded?
Permissions model Does the assistant see exactly what the person asking is allowed to see in the file system?
Effort to the first answer Weeks or months, and who carries it?
Operational ownership Who maintains index, models and dependencies once the project becomes routine?

The first two rows are not comfort features in engineering. A statement about a standard without a file and page reference is not merely inconvenient, it is unusable, because it cannot be checked. The reasoning is laid out in the article on citations in standards search. The third criterion is routinely underestimated: a substantial share of engineering knowledge sits in images, and a system that skips them answers half the question. See the article on multimodal AI in engineering.

When is another category the better choice?

These cases exist, and naming them is more honest than leaving them out:

  • The question does not depend on your documents. For research, phrasing and translation a generic assistant is the faster tool, and a specialised system adds nothing.
  • Your knowledge lives entirely in Microsoft 365 and consists mostly of Office documents. Then integration into the existing environment is an argument that outweighs domain focus.
  • You are looking for structure, not content. Revision states, part numbers, bills of materials: your PLM answers those better, and no assistant changes that.
  • You have a development team and a narrow use case. Then building it yourself can be cheaper, provided operations are part of the plan.

And the boundary in the other direction: KoAssist does not write standards, does not assess conformity and does not take over responsibility for the judgement. It shortens the path to the passage. The engineering decision stays with the engineer, and that is the right division of labour.

How do you test this in two weeks?

A comparison assembled from datasheets rarely produces a decision that holds. The following process costs a few person-days and yields comparable results:

  1. Collect twenty real questions. From daily work, not invented. Five of them should be questions whose answer is demonstrably not in the set.
  2. Provide a bounded test set. Frequently used standards and guidelines, plus one closed project with drawings and correspondence. The typical mess belongs in it, not a tidied-up special case.
  3. Score every answer on three axes: is it correct, does the citation lead to the right page, and how long would the manual search have taken.
  4. Evaluate the five trap questions separately. This is where it shows whether a system stays silent or invents. This point separates candidates more sharply than any feature list.
  5. Put the operational questions in writing: hosting, subprocessors, permissions model, exclusion from training. A checklist for that sits under GDPR-compliant AI, and how KoAssist implements it under security and EU hosting.

If you are weighing operating models, the trade-off is set out in the article on on-premise or EU cloud. How the question plays out with the most widely used generic tool is covered in the GDPR check on ChatGPT for internal documents.

Conclusion: the comparison is decided at the citation

The four categories solve different problems, and none of them is right for everything. For questions whose answer sits in standards, specifications and project documents, and which have to stand up to checking, the field narrows quickly: what remains is whatever traces every statement back to file and page, reads drawings, admits what it does not know, and keeps the documents inside a legal perimeter you can defend to your data protection officer. That is what KoAssist is built for.

The product-by-product comparisons are on the KoAssist comparison page. If you would like to run the process above against your own material, book a demo.

For an overview of all the use cases, see our topic hub AI in Engineering Design.

The feature at the centre of all this is the knowledge assistant.

FAQ

What is the difference between an AI assistant and full-text search in a DMS?

Full-text search returns documents, an assistant answers the question. In practice the difference is the work in between: with full-text search you have to know the right term and read the hits yourself. An assistant takes the question in your own wording, searches across documents and points at the passage. The price is that you have to verify the answer, and that is exactly what the citation is for.

Is a generic assistant not enough if our data is excluded from training?

Exclusion from training is an important commitment, but it answers only one of several questions. Still open: where the documents are processed and stored, which subprocessors are involved, and whether the assistant reaches your actual document base or only what someone uploads into a chat. For a dependable answer out of standards and project files, one upload per question is not enough.

How many documents does a meaningful test need?

Fewer than most people assume. A bounded set of a few hundred documents that reflects everyday work tells you more than a full import: the standards and guidelines you reach for often, plus one closed project with its drawings and correspondence. What matters is not volume but that the typical formats and the typical mess are represented.

How do you tell whether a vendor only claims to cite sources?

Three probes. First: does the citation point at the specific page or only at the document. Second: when you open the passage, does it say what the answer claimed. Third: what happens with a question whose answer is not in the set. A system that produces something plausible instead of saying "not found" is not usable in engineering.

How much internal time does such a comparison cost?

For the process described here, plan on roughly two weeks of elapsed time and a few person-days: half a day for the question list, a day to provide the test set, then distributed review across the team. The largest item is not the technology but agreeing on the twenty questions you will measure against.

Yanik Yeganehfar, co-founder of KoAssist

Find out with Yanik whether KoAssist fits your engineering team.

Book a free conversation