No universally correct chunk size exists, but a strong general baseline is roughly 256 tokens with 32-token overlap. That baseline should be tested against representative support content, not applied blindly.
A customer asks, “Does my plan include SSO?” The support agent retrieves a relevant passage, answers confidently, and omits the exception buried in a pricing table or the eligibility condition stated under a separate heading. The problem isn't always the language model. Often, the evidence was split badly before retrieval ever began.
That's why chunking of information deserves treatment as a retrieval design decision, not a routine preprocessing task. The useful question isn't whether every document was indexed. It's whether the retrieved evidence lets an agent give a complete, grounded answer that resolves the customer's issue.
Table of Contents
- When Bad Chunking Breaks a Support Answer
- What Chunking of Information Actually Means
- How Chunk Size, Overlap, and Retrieval Work Together
- Matching Chunk Policy to Support Intent
- Evaluating Chunking Against Real Support Outcomes
- Designing Coherent Chunks for Human-Centered Answers
When Bad Chunking Breaks a Support Answer
A customer follows every troubleshooting step in a retrieved answer, then reaches the point where the article says to stop and contact support. That warning sits below the final step, separated by a chunk boundary. The system returns the procedure without the escalation condition, so the customer keeps trying actions that should have ended earlier.
Clean documents, capable embeddings, and a well-configured vector database cannot prevent that failure by themselves. Citation presence isn't the same as citation coverage. The answer may cite the correct article and give technically accurate instructions, yet leave the customer without the condition they need to stop and escalate.

The hidden failure inside a confident answer
Poor chunking creates recurring failure modes:
- Boundary gaps: A sentence, table row, exception, or cross-reference is divided between chunks.
- Topic dilution: A large chunk combines eligibility rules, marketing copy, and unrelated product details, weakening its retrieval signal.
- Procedural fragmentation: Steps are separated from prerequisites, expected results, or escalation criteria.
- Metadata loss: The text survives extraction, but its section, page, product, or version context disappears.
More context does not automatically improve an answer. An oversized chunk can bury the relevant sentence among unrelated material, lowering its retrieval rank or making it easier for the model to overlook after retrieval. Research on chunk evaluation describes this trade-off: oversized chunks can reduce retrieval precision, while fixed token splitting can ignore headings, lists, tables, and procedural dependencies (HiChunk evaluation research).
Practical rule: Judge a chunk by the answer it enables, not by how neatly it fills an index.
Treat the pipeline as one system
Chunking interacts with document parsing, segmentation, embedding, retrieval, and answer generation. A parser that turns a table into scrambled text can defeat a sensible segmentation policy. An embedding model may handle prose well but struggle with table-derived questions. A retriever may rank a concise fact correctly while missing the second chunk needed to explain a multi-step process.
A controlled benchmark covering 800 questions, four structurally different government-regulation documents, three parsers, three chunking strategies, five dense embedding models, and BM25 found that no retriever family or parser-chunker combination consistently dominated (benchmark findings). Test complete configurations on the content your support team serves rather than selecting a chunk size in isolation. An AgentStack-style ingestion pipeline is a practical place to compare parsing and chunking variants before they reach production.
Classify failures by support intent. For fact lookups, check whether the retrieved chunk contains the full rule, qualifier, and source context. For procedures, check whether it preserves prerequisites, ordered steps, expected results, and escalation conditions. Then measure answer-level outcomes such as resolution and groundedness, not retrieval scores alone.
Log questions that produce incomplete, unsupported, or escalated answers. Inspect the missing evidence span: was it absent from the index, split across a boundary, misread from a table, filtered by metadata, ranked too low, or ignored during generation? That diagnosis shows whether chunking needs attention or another layer is responsible.
What Chunking of Information Actually Means
Chunking means grouping lower-level information into meaningful units. In a support system, that unit might be a product feature explanation, an eligibility rule, a troubleshooting stage, or a policy exception. It shouldn't be defined only by a character count.
Psychologist George A. Miller's 1956 paper, “The Magical Number Seven, Plus or Minus Two,” made chunking a foundational idea in cognitive psychology. Miller argued that short-term memory could hold approximately 7 ± 2 meaningful units, while emphasizing that a chunk isn't necessarily one character or word. People can group several elements into one familiar unit, such as a recognizable telephone-number pattern (Cowan review of Miller's work).
The practical lesson survives even though later research challenged the breadth of Miller's estimate: organization changes how information is handled. A support answer that groups eligibility, steps, exceptions, and escalation conditions is easier to understand than one that presents the same facts as disconnected fragments.

A chunk is more than a text slice
A useful retrieval chunk usually preserves several kinds of context:
- Structural context: Keep the heading or section label attached to its content.
- Logical context: Keep a question with its answer, a prerequisite with its procedure, and an exception with the rule it qualifies.
- Format context: Preserve tables, lists, code blocks, and labels instead of flattening them into ambiguous prose.
- Source context: Retain product, version, page, document, and section metadata.
This makes document parsing part of chunk design. Before choosing a splitter, inspect how source material becomes machine-readable. A practical overview such as the Writingmate document extraction guide is useful when comparing extraction approaches for PDFs and other formats, because a parser can preserve or destroy the structure your chunker needs.
Why smaller isn't automatically better
Human working-memory research offers a related technical rationale. A major review describes broad attentional capacity as approximately four independent chunks, with other evidence placing effective working-memory capacity around three to five meaningful items when rehearsal and long-term-memory support are controlled (working-memory review).
That doesn't mean every retrieval chunk should contain only a few items. It means the contents should form a coherent unit. A short fragment that omits the condition needed to interpret it may be less useful than a longer section that preserves the complete rule. Conversely, a large block containing unrelated topics forces both retriever and model to separate signal from noise.
The right mental model is therefore semantic packaging. Token limits matter, but headings, dependencies, tables, and expected answer shape matter too.
How Chunk Size, Overlap, and Retrieval Work Together
A support answer can fail even when the right document is indexed. A boundary may separate a requirement from its exception, or a table parser may detach values from their headers. Chunk size is one part of the configuration alongside parsing, segmentation, embeddings, and retrieval. A controlled RAG benchmark found that no single combination consistently performed best, so changing one setting without testing the others can produce misleading conclusions.
The benchmark reported a useful baseline of 256-token chunks with 32-token overlap. Top-five evidence recall was approximately 0.90, compared with approximately 0.87 for the same chunk size without overlap (benchmark configuration and results). Treat those figures as a starting point, not a universal rule. Overlap can preserve evidence that crosses a boundary, while still adding duplicate passages and retrieval noise.
| Configuration aspect | Effect on retrieval | Support impact |
|---|---|---|
| Smaller semantic units | Improve precision for narrow factual questions, but may omit surrounding conditions | Useful for feature, eligibility, and definition lookups |
| Larger coherent sections | Preserve relationships across steps and exceptions, but may dilute relevance | Useful for troubleshooting and policy explanations |
| Modest overlap | Reduces the chance that boundary-spanning evidence disappears | Can preserve qualifications and adjacent procedural context |
| Heavy overlap | Adds duplicates and context redundancy | May increase index size and cause repetitive retrieval |
| Structure-aware parsing | Preserves headings, lists, tables, and document relationships | Particularly important for policies, FAQs, and table-heavy content |
| Blind token splitting | Produces predictable units but can sever meaning | Suitable as a baseline, not a default for every format |
Overlap solves one problem and creates another
Overlap helps when a condition begins near the end of one chunk and finishes at the start of the next. A modest shared span gives retrieval two chances to surface the complete rule. Excessive overlap copies the same material across many records, enlarges the index, and can return near-identical passages instead of complementary evidence.
For an AgentStack-style ingestion workflow, start with structure-aware parsing, approximately 256-token semantic units, and 32-token overlap. The AgentStack RAG chunking strategies guide provides a practical reference for configuring ingestion around document structure and retrieval behavior. Validate that baseline separately against prose, tables, FAQs, and procedures. A policy that helps fact lookup may still produce incomplete troubleshooting answers.
Chunk policies also affect how reliably support content is surfaced in generated answers and other systems. Guidance on optimizing for AI visibility can provide broader context, but retrieval testing should remain tied to evidence quality, answer groundedness, and resolution. A chunk that ranks well but omits a prerequisite can make an answer appear relevant while leaving the customer unable to complete the task.
Parse structure before tuning tokens
Tables require separate handling. A row may depend on column headers, units, plan names, or footnotes. Flattening those relationships into a value sequence can create plausible embeddings that support an incorrect answer. Preserve the table's semantic relationships, or store a representation that keeps headers and qualifying notes with each answerable unit.
Test the parser, chunker, embedding model, and retriever as one pipeline. Measure whether the retrieved evidence contains the required condition, whether the final answer is grounded in that evidence, and whether the customer's issue is resolved. If only chunk size changes, a failure caused by table extraction or ranking may be blamed on segmentation.
Matching Chunk Policy to Support Intent
A customer asking, “Does the Pro plan include audit logs?” needs a narrow, quotable fact. A customer asking, “Why did my webhook stop working after the migration?” needs a sequence of prerequisites, diagnostic checks, expected outcomes, and escalation conditions. These questions should not retrieve identical evidence.
Recent retrieval experiments found that 64 to 128 tokens performed best for concise, fact-based questions, while 512 to 1,024 tokens performed better for descriptive or technical answers (intent-sensitive chunking research). The same study reported that increasing chunks to 512 tokens reduced recall by roughly 10% to 15% in some settings, while NewsQA reached a 55.9% recall@1 peak at 512 tokens. These results show why chunk size is an intent-specific choice, not a universal setting.

Build policies around question shape
Use intent categories to define the evidence a retriever should prefer:
- Fact lookup: Keep the feature, plan, definition, or eligibility statement compact and tied to its heading.
- Procedure: Keep prerequisites, ordered steps, expected results, and failure paths together where possible.
- Troubleshooting: Preserve the diagnostic branch, symptoms, checks, and escalation condition as one coherent unit.
- Policy explanation: Include the rule, exception, scope, and effective context rather than an isolated sentence.
- Comparison: Keep the dimensions being compared aligned, especially when the source uses a table.
A single fixed size can work across mixed content, but it often hides opposing failure modes. Small chunks can improve fact matching while excluding the conditions that make a procedure safe. Large chunks can preserve a policy's logic while reducing precision for a simple feature lookup.
The useful question isn't “What chunk size should we use?” It's “What evidence does this support intent require?”
Apply the same intent-aware thinking to ingestion. AgentStack's document ingestion pipeline offers a practical reference for handling websites, documents, and structured knowledge before retrieval. Preserve headings, lists, and relationships during ingestion, then test whether each intent receives the evidence its answer requires.
Retrieval balances precision against coverage. Narrow passages can rank strongly for a specific fact but omit a prerequisite or exception. Wider passages provide more surrounding context, but introduce unrelated text that can dilute ranking and increase the chance of unsupported synthesis. The right policy depends on which failure is more costly for the support workflow.
Measure the answer, not just the hit
Retrieval recall is useful, but it is not the finish line. Compare policies by groundedness, citation coverage, successful resolution, and escalation behavior. A chunk that retrieves a relevant sentence while omitting an exception may look healthy in a retrieval report and still produce an unsafe answer.
Keep separate test sets for each intent. A compact fact set should contain precise eligibility and feature questions. A procedure set should contain multi-step tasks. A policy set should include exceptions and scope. One configuration may perform well on fact lookup and poorly on troubleshooting, so route or adapt policies instead of declaring a universal winner.
Evaluating Chunking Against Real Support Outcomes
A customer asks whether a feature is available, retrieves a paragraph that mentions it, and receives an answer that misses the plan restriction below the heading. The retrieval result looks relevant, yet the support outcome is wrong. Chunking needs an evaluation loop that exposes this kind of failure, not one that only ranks configurations.
Start with real support questions, especially those that caused an escalation, correction, low-confidence response, or repeated customer follow-up.

1. Log unanswered queries
Capture the customer question, retrieved passages, generated answer, citations, final disposition, and the evidence a human agent used to correct or complete the response. Record whether the answer resolved the issue and whether escalation was required, rather than recording only whether the system returned text.
A missing answer can originate in several places:
- The source document may not contain the information.
- The parser may have lost a table, heading, or footnote.
- A chunk boundary may have separated dependent evidence.
- Metadata may have filtered out the correct passage.
- Retrieval may have ranked the evidence too low.
- Generation may have ignored or misread retrieved content.
This classification stops teams from answering every failure by adding more context.
2. Audit the evidence boundary
For each failed answer, locate the smallest span that would support a complete response. Inspect the surrounding chunks and check whether the heading, prerequisite, exception, table header, or escalation rule remained attached.
A useful diagnostic artifact is a boundary report showing each chunk with its source section, neighboring chunks, token length, overlap, and retrieved rank. Engineers and support leads can use it to see fragmentation that ordinary retrieval logs hide.
3. Compare complete configurations
Run structural, hierarchical, and adaptive policies against the same documents and question sets. Keep the parser, embedding model, retriever, answer prompt, and evaluation criteria controlled where possible. In a second experiment, change one pipeline choice at a time to isolate its effect.
Benchmark design also matters. Standard RAG tests can under-assess segmentation when they lack multi-level boundary annotations and evidence-dense questions. A modular benchmark covers text, tables, and knowledge graphs, with varied query and chunk relationships. Its HiChunk benchmark description supports testing cross-format and multi-step questions, not only clean paragraph lookups.
4. Score business-level outcomes
Use a scorecard that connects retrieval behavior to support work:
- Evidence recall: Did retrieval include the passages needed for the answer?
- Groundedness: Did the response stay within the retrieved evidence?
- Citation coverage: Did citations support every material claim, including exceptions?
- Resolution: Did the customer receive an actionable, complete response?
- Escalation quality: Did the system escalate only when evidence was missing or human review was warranted?
- Maintenance signal: Which unanswered questions reveal a documentation gap?
For mixed-format corpora, AgentStack's unstructured data processing guide provides a practical ingestion reference. Extraction quality must remain consistent across policies, or a chunking comparison may be measuring parser performance.
5. Refine the policy deliberately
If the missing span crosses a heading or procedure boundary, improve structural parsing or segmentation. If the evidence exists but ranks poorly, investigate embeddings, metadata, hybrid retrieval, or reranking. If the answer cites the correct chunks but omits a condition, improve answer structure and grounding controls instead of just increasing chunk size.
Keep a small, permanent regression set of difficult questions, separated by support intent. Fact lookups should test precise answers, procedures should test ordered actions, and policy questions should test exceptions and scope. A policy is ready for wider use when it performs reliably on the question types that affect resolution, groundedness, and escalation, not when its chunks look uniform.
Designing Coherent Chunks for Human-Centered Answers
A support article can retrieve the right rule and still produce a poor answer. The failure often begins when eligibility, actions, exceptions, and escalation evidence are split across unrelated chunks. Retrieval should preserve the relationships an agent or customer needs to act.
Chunk boundaries should follow the answer's information architecture. For fact lookups, keep the definition, qualifying conditions, and authoritative value together. For procedures, keep the ordered steps with prerequisites, expected outcomes, and failure handling. Tables need row and column context, while FAQs should retain each question with its answer and scope. These policies reduce ambiguity without forcing every document into one format.
Chunking also shapes how people process retrieved evidence. Working-memory research describes chunking as recoding related elements into a familiar unit, with benefits depending on what the unit contains (working-memory research). A boundary that joins unrelated sentences may save retrieval space while weakening interpretation.
Use a decision matrix rather than one universal policy:
| Support intent | Preferred source unit | Failure signal |
|---|---|---|
| Fact lookup | Definition, value, conditions, and scope | Correct topic, missing qualifier |
| Procedure | Ordered steps with prerequisites and exceptions | Steps retrieved out of order |
| Policy or eligibility | Rule beside exclusions and effective scope | Overconfident answer at the boundary |
| Troubleshooting | Symptom, diagnostic action, result, and escalation path | Repeated or premature escalation |
For AgentStack-style ingestion, apply structure-aware parsing before comparing chunk policies across prose, tables, FAQs, and procedures. Record the policy used for each document type, then evaluate answer resolution, groundedness, citation coverage, and escalation quality. Inspect failed answers at the answer level: a missing condition points to boundary or retrieval design, while an unsupported conclusion points to generation controls.
AgentStack provides website and document ingestion, automatic chunking and indexing, multi-model support, omnichannel delivery, and analytics for unanswered questions and resolution outcomes. Use those signals to revise the matrix as the knowledge base changes.
