Back to blog
·6 min read

How to optimize knowledge base chunking and token limits for precise AI answers

How to optimize knowledge base chunking and token limits for precise AI answers

When an AI chatbot gives vague, incomplete, or truncated answers, the language model itself is rarely to blame. Usually the breakdown happens in the Retrieval-Augmented Generation (RAG) step, where your knowledge base is split into smaller passages called chunks. If a chunk is cut off in the middle of an explanation, or the token budget fills up with irrelevant text, the chatbot loses the context it needs to answer fully. In this guide we look at how chunk size, overlap, and document structure affect answer quality – and what you can do about it in your own content.

Why chatbots truncate answers and lose critical context

A RAG-based chatbot retrieves relevant passages from your source documents and inserts them into the language model prompt before the answer is written. The problem arises when content is split mechanically, without regard for where a line of reasoning starts and ends. If a policy condition begins in chunk A and its exception sits in chunk B, retrieval may only return chunk A. The customer then gets an answer that is technically taken from your documentation, yet still misleading because a key qualifier is missing.

Example

Chunk A: "You can return items within 30 days for a full refund."

Chunk B: "This does not apply to customised products or sale items."

A customer asks whether they can return a customised jacket. Retrieval matches chunk A, and the chatbot says "yes" – because the exception never made it into the context.

Token limits make this worse. Every answer is generated within a fixed token budget that has to fit the system instructions, conversation history, the user's question, the retrieved chunks, and the answer itself. If retrieved chunks are bloated with filler, there is less room for what actually matters. Read more about how RAG grounds answers in your sources.

When the budget is poorly used, you typically see two failure modes:

  • Truncated answers: The model hits its maximum output length mid-sentence or before reaching the conclusion.
  • Context blindness: Passages containing important exceptions or conditions rank too low or do not fit, and never reach the context.

Balancing chunk size, overlap, and token budget

Optimising a knowledge base is about balancing precision against context. Split documents into chunks that are too small (say 50–100 tokens) and each chunk loses its broader meaning. The model sees an isolated sentence like "This does not apply on weekends" with no idea which service, warranty, or pricing rule it refers to. Chunks that are too large (say 1,500 tokens) pull in so much unrelated text that search precision drops.

The fix is a sensible chunk size combined with overlap between consecutive chunks. An overlap of 10–20 percent keeps sentences that fall on a boundary attached to their context. If chunk A ends with a condition, the start of chunk B repeats it before describing the outcome.

The goal is for every chunk to make sense on its own. This is also the core idea behind writing precise chatbot content: every part of your knowledge base should be unambiguous, even when read in isolation.

Comparing common chunking strategies

Your chunking strategy determines how well the chatbot handles everything from simple FAQs to long service agreements. The table shows three common approaches and how they compare on context retention and truncation risk.

Chunking strategyTypical sizeContext retentionTruncation riskBest suited for
Fixed character length200–500 charactersLow (cuts sentences arbitrarily)HighUnstructured text without paragraphs
Structure-aware (paragraphs and sentences)300–800 tokensHigh (keeps reasoning intact)LowDocumentation, manuals, and web pages
Hierarchical with metadata500–1,200 tokens + metadataVery high (links parent and child sections)Very lowLong contracts, technical manuals, and price lists

Fixed character length is the most common source of problems: when a sentence is split mid-word or right before a qualifier, the risk of wrong or truncated answers rises. Hierarchical chunking performs best but takes more work at ingestion. For most businesses, structure-aware chunking is the best compromise.

Chatly applies structure-aware chunking automatically. Content from websites, PDFs, and other sources is split along paragraphs where possible, and along sentence boundaries when paragraph structure is unclear. Chunks also overlap, so sentences on a boundary keep their context. You do not need to build your own RAG pipeline – but the better structured your content is, the better the results.

How to structure content for better chunks

Even the best chunking algorithm struggles with long, unstructured blocks of text. Follow these principles when preparing your content:

  • Use short paragraphs with one topic each: Paragraph breaks are the most important split points. A paragraph about one thing usually ends up intact in a single chunk.
  • Repeat the subject in the text, not just the heading: Write "The return policy for customised products is…" rather than "It is…". The chatbot then knows what the chunk is about even if the heading lands in a different chunk.
  • Keep rules and exceptions together: Do not bury important terms or pricing caveats in a separate paragraph further down. Put the rule and its exception right next to each other, ideally as a bullet list.
  • Split up large tables: Tables spanning many pages are often split incorrectly. Use several smaller, focused tables, each with its own column headers.
  • Avoid distant cross-references: If a paragraph refers to "the terms in section 2.1 on page 40", restate the key point inline so the chunk stands on its own.

These principles go hand in hand with how you structure training data for your AI chatbot. When sources are clean, every token in the context window goes to useful information instead of filler.

Monitor answer quality in live conversations

Optimising chunking is not a one-off job. After launch, regularly review conversations where customers ask follow-ups like "What do you mean by that?" or "Can you finish that sentence?". These patterns suggest the answer was cut off, or that the retrieved chunks were missing key information.

Connecting conversation analytics with content audits helps you quickly find the documents causing problems. Set up a recurring routine where your team reviews escalations and incomplete answers, and track key chatbot metrics such as resolution rate and the share of unanswered questions.

How Chatly helps you deliver precise answers

Chatly brings conversation logs, answer quality, and analytics together in one dashboard. You can update sources, add new content, and see how the changes affect real customer questions – without involving developers.

Ready for a chatbot that answers completely and precisely? Try Chatly for free or get in touch with our team to see how a well-structured knowledge base can improve your customer support.