Learn / Basics

What is chunking, and which strategy should you use?

Updated 3 October 2026 · 3 min read

Short answer

Chunking is splitting a document into passages before they are indexed for search. The goal is passages small enough to match a question precisely but complete enough to answer it, which is why splitting along the document's own structure usually beats cutting every fixed number of characters.

With troveGEN

troveGEN chooses a chunking strategy from each document's structure until you set one yourself, and a built-in simulator shows exactly how a sample will be split.

See what troveGEN provides ↓

Why chunking matters

A search system retrieves passages, not whole documents. If a passage is too large, the relevant sentence is diluted by everything around it and ranks poorly. If it is too small, it loses the context that makes it meaningful, such as which product, clause or patient it refers to.

Published comparisons of retrieval setups consistently find that the chunking choice can influence quality as much as the choice of embedding model.

The common strategies

  • Fixed-size: cut every N characters or tokens, often with overlap. Simple and fast, but it slices through sentences, tables and lists.
  • Recursive: try to split on paragraph breaks first, then sentences, then words, only cutting harder when a piece is still too large. A good general default.
  • Structural: follow the headings, sections and tables of the document, and carry the heading path with each passage so context is not lost.
  • Semantic: split where the meaning changes, found by comparing embeddings of neighbouring sentences. Better boundaries at a higher compute cost.
  • Parent and child: index small passages for precise matching but return the larger section around them for context.

Things that should never be split

Some content is only meaningful whole. A table row separated from its column headings is a string of numbers. A numbered clause cut away from its sub-points changes what it says. A list of medications or a policy with its exceptions answers a question only when complete. A good strategy recognises these shapes and keeps them together.

Overlap, size and metadata

Overlap repeats a little text at each boundary so a sentence on the edge is not lost, at the cost of some duplication. Size is a trade-off: a few hundred tokens is a common starting point, but the right number depends on the documents and the questions. Attaching metadata such as the page, the heading path and the document type to every passage makes later filtering and citation far easier.

Choosing in practice

Start with the document type. Highly structured documents with headings and tables reward structural chunking. Flowing prose is well served by recursive chunking. Then test: build a small set of real questions and measure whether the right passage is retrieved, rather than choosing by instinct.

Key takeaways

  • Chunk boundaries decide whether an answer can be found and read in one piece.
  • Splitting along headings, sections and tables beats blind fixed-size cutting for structured documents.
  • Choose a strategy by testing it on your own questions.

How troveGEN helps with chunking

troveGEN chooses a strategy for each document automatically from its structure, until you set one yourself. Passages carry their heading path, page and type, tables are handled as tables, and a built-in simulator shows you exactly how a sample would be split before you ingest it.

What troveGEN provides

  • Automatic strategy selection from headings, tables and layout
  • Fixed-size, recursive, structure-based and semantic options when you want control
  • Heading path, page number and type attached to every passage
  • A simulator to preview chunk boundaries before ingesting
  • Neighbour, section and parent expansion to return the context around a match

Preview your own documents Start free — 500 pages

Frequently asked questions

What chunk size is best?

There is no universal number. A few hundred tokens is a common starting point. Test with real questions: if answers are cut off, go larger or split along structure; if results are vague, go smaller.

Does overlap help?

A little overlap protects boundary sentences, but too much creates near-duplicate results and extra cost. Structural splitting often needs less overlap because boundaries fall at natural breaks.

Is semantic chunking worth it?

It can give cleaner boundaries for unstructured prose, but it costs more to run and does not help with tables or numbered clauses. Many teams get more from structure-aware splitting plus testing.

How does troveGEN help with chunking?

troveGEN chooses a chunking strategy from each document's structure until you set one yourself, and a built-in simulator shows exactly how a sample will be split. It provides: Automatic strategy selection from headings, tables and layout; Fixed-size, recursive, structure-based and semantic options when you want control; Heading path, page number and type attached to every passage; A simulator to preview chunk boundaries before ingesting; Neighbour, section and parent expansion to return the context around a match.

Keep reading