Skip to content
07 July 2026

ChatGPT for internal documents: the 2026 GDPR check

Uploading internal documents to the public ChatGPT? Risky. The 2026 GDPR check: what is allowed, where the pitfalls are, and which alternative cites its sources.

Quick answer

The public ChatGPT is not readily suitable for confidential company documents. Inputs may be processed outside the EU and, depending on the plan, used to improve the model. The use only becomes GDPR-compliant with EU hosting, a data processing agreement, a contractual exclusion of training use, and traceable answers with source citations. Dedicated B2B systems provide this, the standard chatbot does not.

ChatGPT for internal documents: the 2026 GDPR check

Can you use ChatGPT with internal company documents?

The impulse is understandable. ChatGPT answers questions in seconds, and pasting in your own manual or specification is the obvious next step. Technically it works. The question is not whether it can be done, but whether it may be.

As soon as confidential engineering data, customer information or personal data enters the picture, a convenience question turns into a compliance question. And that one is not answered by feature lists but by data location, contractual position and traceability.

This article works through it in order: what the GDPR actually requires here, which risks genuinely arise, what changes between the ChatGPT plans, when the use is defensible and when it is not, and what the alternative has to deliver.

What the GDPR actually requires here

The GDPR does not prohibit the use of AI. It sets conditions, and those come down to four questions.

Is there a legal basis? As soon as personal data is processed, that processing needs a basis, for example legitimate interest or performance of a contract. Technical documents look harmless at first glance but carry personal data more often than expected: names of inspectors in test reports, contacts in project files, signatures on approvals.

Are the roles clear? Whoever deploys a tool is the controller. The vendor is the processor. That relationship needs an agreement under Art. 28 GDPR. Without it there is no basis for a service provider to process personal data on your behalf.

Where does the data go? If processing leaves the EU, the requirements for third-country transfers apply. There are ways to handle that, but they have to be chosen and documented deliberately rather than accepted in silence.

Does the processing stay tied to its purpose? If inputs are used to improve models, that is a further processing for a different purpose. This is precisely the difference between a tool that answers your question and a tool that learns from it.

These points are laid out in full on the page about GDPR-compliant AI tools, including a checklist you can take into a vendor call.

What risks arise when uploading internal documents?

Data location. In the free version it is often not transparent where inputs are processed. A third-country transfer triggers additional GDPR requirements. On top of that, a European data center alone does not settle the question when the operator belongs to a group headquartered in the US and therefore falls under the US Cloud Act. What the location actually changes, and what it does not, is covered in the article on EU hosting and AI.

Training use. Depending on plan and setting, inputs can be used to improve the models. For internal documents that is a non-starter and has to be contractually excluded. What matters is the legal form of the assurance: a setting in an account can be changed by an update or a change of plan, a contract clause cannot.

No data processing agreement. Without a DPA there is no legal basis for a service provider to process personal data. For a privately created account that agreement usually does not exist at all, because the account does not belong to the company.

Trade secrets, not only data protection. An engineering drawing often contains no personal data and is still the most valuable thing in the building. Trade secret protection depends on whether reasonable confidentiality measures were taken. An uncontrolled upload into a private chatbot account works directly against that evidence.

No traceability. A general chatbot phrases answers from its training knowledge and does not cite them. For technical content, an answer without a source cannot be verified and therefore cannot be relied on. Why that matters most in technical documentation is shown in the article on AI with source citations.

What changes between the ChatGPT plans

The most common confusion in this debate: ChatGPT is not one product but a product family with markedly different terms.

In the consumer plans, control over training use sits in an account setting. The account belongs to the person, not the company. There is no company-wide contract, no central administration of access, and no way to enforce data handling organisationally.

In the business plans the picture changes: there are contractual assurances, administrative controls and usually an exclusion of training use. That is a real difference, and the reason why the blanket statement "ChatGPT is forbidden" is just as wrong as "ChatGPT is approved".

What this does not automatically resolve are the questions of processing location and of which group owns the provider. Both remain to be assessed regardless of plan.

Because plan terms change, what belongs here is a method rather than a snapshot: ask to see the currently applicable terms and the data processing agreement instead of relying on a summary found online. That applies to every vendor, including us.

The application and the API are two different things

A distinction that causes recurring misunderstandings: the chat application in the browser and the programming interface of the same vendor are governed by different terms. When your IT department says "we use the model through the API, that is handled", it says nothing about what happens when someone opens the web interface with a private account in parallel.

Both routes have to be assessed and governed separately. In practice that means an approval for an integrated system is not an approval for the chatbot in the browser, and the other way round.

Four misconceptions that show up in every discussion

"We anonymise the documents first." It sounds tidy and fails in practice on effort. Someone who needs an answer in two minutes does not redact 40 pages. And with technical material the problem is often not the name but the drawing itself. Anonymisation helps with data protection and not at all with trade secrets.

"We only upload an excerpt." The excerpt is frequently the part that counts: the tolerance, the calculation basis, the deviation from the standard. Volume does not determine how worthy of protection something is.

"We are too small to be interesting." The GDPR sets no lower threshold for the applicability of its principles, and trade secret protection does not depend on company size but on whether reasonable measures were taken. Smaller suppliers also frequently work with documents belonging to their large customers, whose non-disclosure agreements travel with them.

"The model does not remember it anyway." Whether a model learns from an input is a question of contractual terms, not intuition. And even where no training happens, the data was transmitted and processed. The data protection question arises at transmission, not first at training.

Decision tree: when is ChatGPT defensible, and when is it not?

The following sequence answers the question for one specific use case. Work through it top to bottom and stop at the first no.

Step 1: Does the document contain personal data or trade secrets? No, it is publicly available content or general help with phrasing. → Uncritical. Reading on is still worthwhile for the next case. Yes. → Go to step 2.

Step 2: Are you using a company account under a company contract? No, it is a private or self-created account. → Not defensible. The contractual basis is missing entirely, regardless of any setting. Yes. → Go to step 3.

Step 3: Is there a data processing agreement, and does it exclude training use? No, or unclear. → Not defensible until that is settled. Settling it takes days, not months. Yes. → Go to step 4.

Step 4: Is the processing location clear and the third-country question assessed? No. → Park it and assess with your data protection officer. The assessment lands differently depending on the data category. Yes. → Go to step 5.

Step 5: Does the answer have to be verifiable? Yes, it concerns standards, tolerances, test steps or figures that feed into a document or a decision. → The general chatbot is the wrong tool. Not for data protection reasons, but because an answer without a source has no value here. No, it concerns phrasing, summarising or structure. → Defensible within the clarified frame.

The sequence exposes a point that often gets lost in privacy discussions: even when every legal question is answered cleanly, a technical problem remains for engineering content. That is step 5, and in practice it decides more often than the steps before it.

Public ChatGPT vs. dedicated AI for internal documents

Criterion Public ChatGPT Dedicated AI for internal documents
Data location often US or third country EU data center
Training use of inputs possible depending on plan contractually excluded
Data processing agreement limited standard
Source citation no yes, down to file and page
Access control and audit log limited role-based and logged
Knowledge base general training knowledge your documents only
Behaviour at knowledge gaps keeps phrasing plausibly points to the missing source
Deletion on request depends on plan with a deadline and a record

The last row gets overlooked and is one of the most effective questions in a vendor call: what exactly is deleted, within what deadline, and do you get proof? Document copies, index, embeddings and caches are four different things.

The real exposure is shadow IT

In practice, what happens is rarely decided by the policy but by the availability of an alternative.

An engineer faces a question whose answer sits in a 300-page standard or in a project folder from 2019. There are two options: spend half an hour searching, or paste the passage into a chat window. If there is no approved third route, the unapproved one appears by itself, and invisibly.

That leads to an uncomfortable conclusion: a ban without an alternative raises the risk rather than lowering it. It only moves the usage to where nobody can see it. A properly vetted tool is therefore also a privacy instrument. It removes the reason for the detour.

In practical terms: policy and tool belong together at rollout. The policy alone produces shadow IT, the tool alone produces sprawl.

What should you look for in a GDPR-compliant solution?

Is the data processed in the EU, and is the operator free of a US parent company? That determines whether the US Cloud Act applies.

Is training use contractually excluded rather than merely toggled off in the settings?

Does the system back every answer with a source? Without proof, every statement remains an act of faith.

Is there role-based access control and an audit log? Not everyone in the company should be able to query every document.

How long does a full deletion take, and what does it cover? A deletion promise that only covers the document copies is incomplete.

How these points are implemented at KoAssist, including subprocessors and the data flow, is set out on the page about security and EU hosting. A vendor-neutral checklist for your own selection process is available under GDPR-compliant AI tools. And how KoAssist sits next to ChatGPT and other tools is shown in the comparison.

What you can settle in the next two weeks

The discussion often stalls because it gets framed as a matter of principle. It can be made operational in a few steps.

Establish the current state first. Ask the team openly, and without threatening sanctions, who currently uses which tool for which task. The answer is rarely "nobody", and it is the basis for everything that follows. Starting with a reprimand produces no reliable information.

Then sort the tasks, not the tools. Separate tasks with no material worth protecting, such as phrasing help and translation of public texts, from tasks on internal documents. The first group usually needs only a clear rule. The second group is the actual need, and it tends to be smaller and more concrete than feared.

Next, check the contractual position. For the plan already in use: is there a data processing agreement, is training use excluded, and where does processing happen? A vendor answers those three questions within a few days or not at all. Either outcome is usable information.

Finally, decide whether a second tool is needed. If the second group of tasks dominates and answers have to be verifiable, the path leads to a dedicated system. If not, a governed use of the existing one is enough.

This order has a practical advantage: it produces a basis for a decision within two weeks instead of extending a debate while uploads continue in the background.

What does this mean in practice?

The public ChatGPT is an excellent tool for general tasks. For confidential internal documents it is the wrong tool, not because it is too weak but because it was built for a different purpose.

What technical teams need is not an all-knowing chatbot but a system that answers solely from their own files, backs every statement, and processes the data within a defensible legal frame. The decisive difference is not the intelligence of the model but the origin of the answer and control over the data.

If you have to decide this for your company, the decision tree above works as a basis for the conversation with IT and privacy. It answers in ten minutes what otherwise takes three meetings.

This article does not replace legal advice. It is meant to give you the questions you take into the internal discussion.

To see what this looks like with your own documents, the fastest route is direct: book a demo.

For the wider context, see our topic hub AI in engineering, which places data protection, source citations and further use cases in an overview.

The alternative inside your own document base is the knowledge assistant: it works on the documents your company already holds.

FAQ

May I upload internal documents to the public ChatGPT?

For confidential or personal content this is not advisable as long as there is no data processing agreement and the data location and training use are unresolved. Many companies therefore prohibit uploading internal documents to public chatbots by policy.

Is it safe to upload internal documents to ChatGPT?

It depends on three things: the plan you are on, where the processing happens, and what your contract says. Without a data processing agreement, without a clarified processing location and without a contractual exclusion of training use, uploading confidential material is not defensible. With those three settled it becomes a case-by-case assessment rather than a blanket approval, and the decision tree further down walks through it.

Is ChatGPT Enterprise GDPR-compliant?

Enterprise and Team plans offer more control than the free version, such as excluding training use and contractual assurances. Whether the use is fully GDPR-compliant depends on the specific data location, the data processing agreement and the use case, and should be assessed case by case.

Are my inputs used to train the AI?

In free consumer versions this is possible depending on the setting. In B2B plans, training use can usually be excluded. That assurance should be fixed contractually, not only enabled in the settings, because a setting can be reset by a product update or a change of plan.

What happens to documents we have already uploaded?

First establish which plan and which setting applied at the time of the upload, and whether the content contained personal data or trade secrets. Delete the affected conversations and document what you did. Where personal data is involved, the case belongs with your data protection officer, who assesses whether a reportable breach occurred.

Is an internal policy enough protection?

A policy is necessary but ineffective on its own when there is no approved route. Someone who needs an answer and has no sanctioned tool will find a workaround. The policy only becomes effective in combination with a provided alternative that actually solves the underlying task.

Can we at least use ChatGPT to translate technical texts?

That depends entirely on the content, not on the task. In data protection terms a translation is no lighter form of processing than a summary, because the full text is transmitted either way. For publicly available texts this is uncritical. For an internal operating manual or a specification, the same questions about contract, processing location and training exclusion apply as for any other use.

What is the alternative to ChatGPT for internal documents?

Dedicated B2B systems process your documents in an EU data center, exclude training use contractually, and back every answer with a source. They answer solely from your files instead of general model knowledge. More on this in the article on AI with source citations.

Yanik Yeganehfar, co-founder of KoAssist

Find out with Yanik whether KoAssist fits your engineering team.

Book a free conversation