The measurement that only 5% of fixed-size boundaries landed on a sentence terminator is the kind of number more chunking posts should report, because it makes the heading-split failure concrete instead of hand-waved. One thing I have seen bite people with SemanticChunker is that the similarity threshold is corpus-sensitive, so a value tuned on prose quietly over-splits on code-heavy or tabular docs where sentence embeddings are noisy. Did you measure retrieval quality downstream of the three splitters, or just the chunk boundary statistics, since the boundary that reads cleanest to a human is not always the one that gives the best recall?