In-House AI Chatbots Can Be Built on Your Own Server — Setup Standards That Keep Company Data From Going Out
An in-house AI chatbot can be built on your own server. The line that divides setups is where your questions and company documents go. This article compares three setups and walks through, in order, where the data sits and the legal checkpoints.
| Setup | Where the data stays | What to check first |
|---|---|---|
| ① External AI API connection | Questions and documents sent to external servers | Scope of transmitted items, storage and retention conditions |
| ② Own server (on-premise) direct operation | Inside the internal network | Model specs, number of concurrent users |
| ③ Hybrid (sensitive data only in-house) | Sensitive documents in-house, general queries only external | Boundary classification criteria |
① When using an external AI API — what goes out
OpenAI states in its official documentation (Data controls in the OpenAI platform) that, since March 1, 2023, data sent via the API is not used for model training or improvement unless the user explicitly agrees. However, abuse-monitoring logs are generated by default and retained for up to 30 days (with exceptions where the law requires longer retention, etc.), and these logs may contain customer content such as questions and answers. Anthropic also notes in its Privacy Center that inputs and outputs of its commercial products (API, Claude for Work, etc.) are not used for training by default, but with the exception of cases where the user explicitly sends feedback.
“Not used for training” and “not stored” are different questions. Even with a training-exclusion policy, the transmission, monitoring, and log-retention segments must be checked separately.
② When personal information is involved — what to check in the law
If questions may contain employee or customer personal information, review the relevant provisions of the Personal Information Protection Act. Article 26(1) specifies what a delegation document must include: “When a personal information controller delegates the processing of personal information to a third party, it shall be done by a document containing the following items.” The items are: 1. matters concerning the prohibition of processing personal information for purposes other than performing the delegated task, 2. matters concerning the technical and managerial protective measures for personal information, 3. other matters prescribed by Presidential Decree for the safe management of personal information. Paragraph 4 of the same article imposes on the delegating party the obligation to educate the delegatee and supervise that processing is carried out safely.
If it is a structure that entrusts processing to an external service, whether these provisions apply becomes a subject of review. Also, if the chatbot server is located abroad, check Article 28-8(1): “A personal information controller shall not provide (including where it is queried), entrust the processing of, or store personal information abroad (referred to in this Section as ‘transfer’).” However, there are, among others, paths that fall under the exceptions: the data subject’s separate consent, special provisions in laws or treaties, entrustment of processing or storage for contract performance with the processing policy disclosed or notified, certification under a public notice of the Personal Information Protection Commission(개인정보보호위원회), and countries recognized by the Commission as providing an equivalent level of protection.
③ When keeping it on your own server — what to decide
The specifications of the model you operate are determined by the size of the model you intend to use and the number of concurrent users. Rather than fixing a specific GPU count first, weigh the response-speed target, the number of concurrent users, and the model’s parameter size together.
Departmental access rights are also designed at this stage. If HR, finance, and sales documents are placed in the same index, a given employee may receive answers beyond their viewing scope.
Using a method that attaches sources to answers (RAG: Retrieval-Augmented Generation, a setup that indexes documents, finds passages relevant to the query, and appends the source to the answer) lets users confirm which document from which department the information came from.
When you build it yourself, there are points where you actually get stuck. Reading methods differ by document format — spreadsheets, PDF tables, scanned images — and departmental viewing scope must be enforced consistently all the way to the answer stage. Reviewing these two points before finalizing the setup reduces the need to rebuild after deployment. → Request a setup review
Verification order
- Classify internal materials by whether they contain personal information and by sensitivity.
- Separate materials that can be sent out from those to keep in-house.
- Choose one of the three setups above, and for external transmission verify the storage and log conditions in the vendor’s documentation.
- If personal information is involved, review whether the Personal Information Protection Act applies with a legal professional.
To sum up
Most companies can start with a setup that keeps an AI chatbot based on internal documents on their own server. If you fix the data location and permission boundaries first, later model swaps or expansions can proceed within the same structure.
An in-house AI chatbot setup begins with data-location criteria and per-department permission design.
If you share the formats of your documents and the departments that use them, we will propose a suitable setup and scope of work. → Request a setup review · View our work
References
- Personal Information Protection Act Article 26(1) and (4), Article 28-8(1) [Verified — National Law Information Center, effective Sep 11, 2026 text]
- OpenAI, “Data controls in the OpenAI platform” [Verified — official documentation]
- Anthropic Privacy Center, “Is my data used for model training?” [Verified — official documentation]
Sources for this article
Laws were cross-checked against the original text on the National Law Information Center(국가법령정보센터) — Personal Information Protection Act(개인정보 보호법) Article 26(1) and (4), Article 28-8(1) (effective Sep 11, 2026 text, verified 2026-09-27). External AI API data policies were verified against the official documentation text of OpenAI's "Data controls in the OpenAI platform" and Anthropic Privacy Center's "Is my data used for model training?" (2026-09-27). Search terms were confirmed for demand via Google autocomplete.
Related services
Work on this topic
- The forgery verdict comes at the very moment identity verification is in progress National R&D · 2024 · 2023.05 – 2024.07 (15 months)
- Custom Self-Hosted AI Work Assistant with No Data Leakage Own product · 2026 · Built September 2026 · In operation
- A document that arrives by email becomes a row in Excel, with no human hand In-house product · 2026 · Deployed 2026.01 · pilot service in operation
- An analysis tool that shows thousands of research comments split by topic and sentiment Research · Academia · 2025 · 2 weeks