All articles

AI Solutions

How to Prepare a Knowledge Base for an AI Chatbot

Prepare reliable chatbot knowledge with approved sources, clear ownership, permissions, policy updates, evaluation questions, and human handoffs.

MyPocket · 6 min read · Updated

Business team reviewing AI tools and data on a laptop in a modern office

An assistant can't give a dependable explanation of a policy your team hasn't agreed on. Connecting more files doesn't solve that problem. It may simply give the system more contradictory material to choose from. Knowledge preparation starts with deciding what is current, approved, and appropriate for the intended user.

A chatbot knowledge base is not just a folder of documents. It's an operating process for maintaining information, controlling access, evaluating answers, and handling questions the source material cannot support. That process matters before you spend time adjusting the assistant's tone.

Define the questions the assistant should answer

Choose a bounded audience and task. A public service assistant may need opening hours, service descriptions, preparation instructions, and an enquiry route. An internal staff assistant may need procedures and role-specific documents. Those should not automatically share the same sources or permissions.

Write examples of intended questions and examples outside the scope. Identify questions requiring live data or a human decision. “What areas do you serve?” can be answered from maintained information. “Can you guarantee an appointment tomorrow?” may require availability and authorization the knowledge assistant doesn't have. Our chatbot versus agent guide explains that distinction.

Inventory sources before importing them

List the website pages, documents, approved FAQs, and other material you plan to use. Record the owner, intended audience, revision date, and access level for each source. Remove obsolete versions or label them clearly enough that they won't be treated as current policy.

Check your right to use the material. Publicly accessible text isn't automatically appropriate to copy into a commercial service, and a customer conversation can contain personal information. Don't connect an entire shared drive merely because it is convenient. Select what the task requires and exclude information the users should not receive.

Resolve contradictions in the underlying content

Suppose one page says appointments can be cancelled up to the previous day and another says cancellation closes several days earlier. A more persuasive prompt can't decide which policy the business intends. The content owner must choose the authoritative rule and update the conflicting sources.

Look for inconsistent service names, expired promotions, informal promises, and instructions written for staff rather than customers. Separate examples from binding policies. If a price depends on assessment, state that directly instead of supplying a sample amount that could be mistaken for a quote. Better source material reduces avoidable ambiguity before any model is involved.

Structure information around a real question

Use clear headings and complete explanations. A useful entry may describe the question, the approved answer, important exceptions, and the next step. Include relevant context in the same section rather than relying on a sentence elsewhere in a long document to qualify the answer.

Avoid turning everything into tiny fragments that lose their meaning when retrieved. An isolated sentence about refunds can be misleading without the conditions around it. Also avoid repetitive variations of the same answer solely to match more keywords. The aim is a source a person can understand and maintain, not a large amount of text for its own sake.

Separate public and restricted information

Enforce permissions in the retrieval and tool layer. If an employee cannot access a document directly, an assistant should not expose it through a summary. Hiding the document link while quoting its private contents is not an effective access control.

Decide whether different users need different knowledge collections. Keep credentials, private financial details, and sensitive records out of general information sources. Treat user messages and retrieved content as information that cannot independently grant authority. A prompt asking the model to be careful is not a substitute for controlling what it can retrieve.

Distinguish maintained facts from live data

A knowledge base can explain how a booking works, but it may not know current availability. It can explain a product's general features, but stock or account status may require an authorized connection. Don't pretend a periodically updated document is a live system.

If live access is needed, define the source of truth and allowed actions. Decide how the assistant responds when the connected service is unavailable. It should acknowledge the limitation and offer a practical alternative rather than turn stale information into a confident commitment. The business AI use case guide discusses these workflow boundaries.

Prepare evaluation questions before launch

Create examples covering common questions, ambiguous wording, incomplete requests, and topics the assistant should decline or hand off. Include questions whose answer is genuinely absent. The system should not be rewarded for inventing a plausible response simply to avoid saying it cannot help.

For each example, record the approved answer or expected behavior and the relevant source. Keep some evaluation examples separate from those used to adjust the configuration. Check whether answers are correct, complete enough, and appropriately qualified. A friendly tone is valuable only when it accompanies reliable information.

Plan missing-answer behavior and human handoffs

Decide how a user reaches a person and what context staff receive. Explain expected response times without implying constant monitoring. Ask only for the information needed for the handoff, and avoid copying an entire sensitive conversation into another system when a minimal summary is sufficient.

Name the person responsible for reviewing recurring unanswered questions. Some should lead to a new approved knowledge entry. Others belong outside the assistant's scope. Treat missing answers as a useful signal about the content and service process rather than automatically expanding the system's authority.

Maintain the knowledge after launch

Assign a review routine and a process for urgent policy changes. Track which source changed, who approved it, and when the assistant's knowledge was refreshed. Re-run relevant evaluation questions after important updates. A correct answer last month may be wrong after the business changes its service.

Review provider terms, data retention, and connected permissions as the implementation evolves. The NIST AI risk-management resources can help frame a broader risk review. They don't replace your own security and operating controls or certify the assistant's output.

A knowledge preparation checklist

  • Define the audience and intended questions.
  • Inventory sources with owners, dates, and access levels.
  • Remove obsolete versions and resolve contradictions.
  • Keep approved answers, exceptions, and next steps together.
  • Exclude unnecessary private information.
  • Separate static knowledge from live system access.
  • Prepare ordinary, ambiguous, missing-answer, and restricted-data evaluations.
  • Provide a human handoff and a correction process.
  • Assign ongoing review and urgent-update responsibilities.

Common knowledge-base questions

Should we upload every company document?

No. Select information appropriate to the task and user. Excess material can create contradictions, privacy exposure, and maintenance work without making the assistant more useful.

Does better knowledge prevent every wrong answer?

No. It can reduce avoidable errors, but retrieval, interpretation, permissions, and output generation still need evaluation. Keep human escalation available and monitor actual results.

Can the assistant learn policies from customer conversations?

Not without a controlled review process. A user's statement is not automatically an approved fact. New knowledge should be checked, authorized, and maintained like other source information.

Build the information foundation first

Explore MyPocket AI solutions or describe your audience and approved sources. A smaller, maintained knowledge base with clear boundaries can be more valuable than a large collection nobody owns.