Chunking for rag under worst-case document conditions: end-to-end qa evaluation of chunking strategies across squad and hotpotqa with dense and sparse retrievers
Loading...
Authors
ORCID
Other Contributors
Issue Date
Type
Electronic thesis
Thesis
Thesis
Language
en_US
Keywords
Degree
MS
Alternative Title
Abstract
Retrieval-Augmented Generation (RAG) systems depend critically on how documents are segmented into chunks before embedding and retrieval, yet the interaction between chunking strategy and embedding model is not well understood, especially on noisy real-world corpora. We examine five chunking strategies (fixed-length, fixed-length respecting sentence boundaries, recursive, semantic boundary-based, and late-token) paired with three embedding models across a possible “worst-case”, unorganized dataset consisting of stitched-together unrelated paragraphs using end to end metrics such as answer-EM and F1 as well as recall@k chunks under a fixed context budget. While preliminary experiments suggested semantic and late-token chunking would be strongest, results on our dataset show recursive and semantic chunking consistently outperform alternatives, with recursive chunking yielding the largest gains on the noisiest documents and most complex questions. We attribute this to dataset heterogeneity and noise, as traditional fixed-length methods struggle with this when the chunk lies across paragraph boundaries, where recursive segmentation better preserves local coherence and avoids spurious boundaries that degrade embedding quality. These findings suggest chunking choice should be conditioned on corpus noise and complexity: recursive for robustness, semantic for less complex questions, and late-token approaches may require explicit noise/boilerplate mitigation.
Description
May2026
School of Science
School of Science
Full Citation
Publisher
Rensselaer Polytechnic Institute, Troy, NY
