feat(vector): page-aware PDF chunking for predictable per-page retrieval

Add PageAwareChunker, which splits paginated documents (PDFs) on page
boundaries first and only character-splits pages larger than chunk_size.
No chunk spans a page boundary, so page_number is always exact and stored
excerpts never lead with a neighbouring page's text. When chunk_size is at
least the largest page, this yields exactly one chunk per page: a
predictable vector count (== page count), a flat per-page embedding cost,
and zero cross-page overlap duplication.

Gated by DOCUMENT_CHUNK_PAGE_AWARE (default true). When false, the legacy
char-based DocumentChunker + post-hoc assign_page_numbers path runs
unchanged. Only doc_type="file" with page_boundaries (PDFs) takes the
page-aware path; notes/deck/news are unaffected.

Measured on a 15-page record (query "leadership award louis", target =
top-half of page 15): char-based degraded the target to dense-rank 10 at
cs=2048 (OCR) and mislabeled its page; page-aware restored rank 1 across
every fusion/modality and chunk size, with correct page labels and clean
snippets.

BREAKING CHANGE: PDFs are re-chunked page-aware by default. Existing
deployments will re-index PDF content on the next vector sync (different
chunk counts and page_number labels). Set DOCUMENT_CHUNK_PAGE_AWARE=false
to retain the previous char-based behaviour.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-06 13:42:25 +02:00
co-authored by Claude Opus 4.8
parent 40e56aeaea
commit 2f2a7f9659
7 changed files with 347 additions and 10 deletions
+1
View File
@@ -743,6 +743,7 @@ equivalent.** Operators who need a runtime toggle should open an issue.
| `SIMPLE_EMBEDDING_DIMENSION` | ⚠️ Optional | `384` | Dimension for the fallback Simple provider |
| `DOCUMENT_CHUNK_SIZE` | ⚠️ Optional | `512` | Words per chunk for document embedding |
| `DOCUMENT_CHUNK_OVERLAP` | ⚠️ Optional | `50` | Overlapping words between chunks (must be < chunk size) |
| `DOCUMENT_CHUNK_PAGE_AWARE` | ⚠️ Optional | `true` | Split PDFs on page boundaries first (one chunk per page; oversized pages split within the page). Exact page numbers, clean snippets, and a predictable ~1 chunk/page when chunk size ≥ the largest page. Set `false` for the legacy char-based path. |
**Deprecated variables (still functional):**
- `VECTOR_SYNC_ENABLED` - Use `ENABLE_SEMANTIC_SEARCH` instead (will be removed in v1.0.0)
+4
View File
@@ -206,6 +206,10 @@ NEXTCLOUD_PASSWORD=
# Configure how documents are split before embedding
#DOCUMENT_CHUNK_SIZE=512
#DOCUMENT_CHUNK_OVERLAP=50
# Page-aware chunking for PDFs: split on page boundaries first so no chunk spans
# a page (exact page numbers, clean snippets, ~1 chunk/page when chunk size >=
# the largest page). Set false to use the legacy char-based path. Default: true
#DOCUMENT_CHUNK_PAGE_AWARE=true
# ===== SEMANTIC SEARCH TUNING =====
# Advanced parameters for vector sync background operations
+12
View File
@@ -132,6 +132,10 @@ _DEFAULTS: dict[str, Any] = {
# Document chunking
"document_chunk_size": 2048,
"document_chunk_overlap": 200,
# Page-aware chunking for paginated docs (PDFs): split on page boundaries
# first so no chunk spans a page (exact page_number, clean snippets, and
# predictable ~1 chunk/page when chunk_size >= the largest page).
"document_chunk_page_aware": True,
# PDF parse isolation (OOM guard)
"document_pdf_graphics_limit": 1000,
"document_parse_timeout_seconds": 120.0,
@@ -736,6 +740,13 @@ class Settings:
# Document chunking settings (for vector embeddings)
document_chunk_size: int = 2048 # Characters per chunk
document_chunk_overlap: int = 200 # Overlapping characters between chunks
# Page-aware chunking for paginated docs (PDFs). When True (default), PDF
# text is split on page boundaries first (one chunk per page; oversized
# pages are character-split within the page), giving exact page numbers,
# snippets that never lead with a neighbouring page, and a predictable
# ~1 chunk/page when document_chunk_size >= the largest page. When False,
# the legacy char-based path runs with post-hoc assign_page_numbers.
document_chunk_page_aware: bool = True
# PDF parse isolation (OOM guard). The parse runs in a subprocess so one
# pathological file fails that doc, not the pod.
@@ -1389,6 +1400,7 @@ def get_settings() -> Settings:
# Document chunking settings
"document_chunk_size": "DOCUMENT_CHUNK_SIZE",
"document_chunk_overlap": "DOCUMENT_CHUNK_OVERLAP",
"document_chunk_page_aware": "DOCUMENT_CHUNK_PAGE_AWARE",
"document_pdf_graphics_limit": "DOCUMENT_PDF_GRAPHICS_LIMIT",
"document_parse_timeout_seconds": "DOCUMENT_PARSE_TIMEOUT_SECONDS",
"document_parse_mem_limit_mb": "DOCUMENT_PARSE_MEM_LIMIT_MB",
@@ -2,6 +2,7 @@
import logging
from dataclasses import dataclass
from typing import Any
import anyio
from langchain_text_splitters import RecursiveCharacterTextSplitter
@@ -97,3 +98,150 @@ class DocumentChunker:
self.overlap,
)
return chunks
class PageAwareChunker:
"""Page-first chunker for paginated documents (PDFs).
Unlike :class:`DocumentChunker`, which splits the concatenated document text
on character boundaries and is therefore page-agnostic, this chunker splits
on the page boundaries FIRST and only falls back to character splitting for
pages larger than ``chunk_size``. As a result:
* No chunk ever spans a page boundary, so ``page_number`` is always exact
and the stored excerpt never leads with a neighbouring page's text (the
char-based path can bury a short page's content in the tail of a chunk
whose majority — and thus :func:`assign_page_numbers` label — is the
previous page).
* Chunks-per-page is ``ceil(page_chars / chunk_size)``. When ``chunk_size``
is at least the largest page, that is exactly one chunk per page, giving a
predictable vector count (== page count), a flat per-page embedding cost,
and zero cross-page overlap duplication.
Page numbers are assigned inline, so callers must NOT additionally run
``assign_page_numbers`` on the result.
"""
def __init__(self, chunk_size: int = 2048, overlap: int = 200):
"""
Initialize page-aware chunker.
Args:
chunk_size: Number of characters per chunk (default: 2048). Pages at
or below this size become a single chunk; larger pages are
character-split (with overlap) within the page only.
overlap: Overlapping characters between sub-chunks of an oversized
page (default: 200). Pages that fit in one chunk carry no
overlap.
"""
self.chunk_size = chunk_size
self.overlap = overlap
# Only used for pages that exceed chunk_size. Same hierarchical splitter
# as DocumentChunker so oversized pages keep semantic-boundary splitting.
self.splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size,
chunk_overlap=overlap,
add_start_index=True,
strip_whitespace=True,
)
async def chunk_text(
self, content: str, page_boundaries: list[dict[str, Any]]
) -> list[ChunkWithPosition]:
"""
Split ``content`` into per-page chunks using ``page_boundaries``.
Args:
content: Full document text. Offsets in ``page_boundaries`` must
index into this string (the extractor contract — see
``document_processors``).
page_boundaries: Ordered list of ``{"page", "start_offset",
"end_offset"}`` dicts. When empty, falls back to plain
character chunking (no page numbers), matching
:class:`DocumentChunker` for non-paginated input.
Returns:
List of chunks with character positions and ``page_number`` set.
"""
if not content:
return [ChunkWithPosition(text="", start_offset=0, end_offset=0)]
# No page info (e.g. a non-PDF that reached this path) — degrade to the
# char-based behaviour so the caller still gets sensible chunks.
if not page_boundaries:
docs = await anyio.to_thread.run_sync( # type: ignore[attr-defined]
self.splitter.create_documents,
[content],
)
return [
ChunkWithPosition(
text=doc.page_content,
start_offset=doc.metadata.get("start_index", 0),
end_offset=doc.metadata.get("start_index", 0)
+ len(doc.page_content),
)
for doc in docs
]
chunks = await anyio.to_thread.run_sync( # type: ignore[attr-defined]
self._chunk_by_page,
content,
page_boundaries,
)
logger.debug(
"Page-aware chunked document into %s chunks across %s pages "
"(chunk_size=%s, overlap=%s)",
len(chunks),
len(page_boundaries),
self.chunk_size,
self.overlap,
)
return chunks
def _chunk_by_page(
self, content: str, page_boundaries: list[dict[str, Any]]
) -> list[ChunkWithPosition]:
"""CPU-bound per-page splitting (runs in a worker thread)."""
chunks: list[ChunkWithPosition] = []
for boundary in page_boundaries:
page = boundary["page"]
start = boundary["start_offset"]
end = boundary["end_offset"]
page_text = content[start:end]
# Skip blank pages: embedding an empty/whitespace-only string wastes
# a provider call and a vector slot.
if not page_text.strip():
continue
if len(page_text) <= self.chunk_size:
stripped = page_text.strip()
# Tighten offsets to the stripped text so they stay meaningful
# even though the whole page is one chunk.
lead = len(page_text) - len(page_text.lstrip())
chunk_start = start + lead
chunks.append(
ChunkWithPosition(
text=stripped,
start_offset=chunk_start,
end_offset=chunk_start + len(stripped),
page_number=page,
)
)
continue
# Oversized page: split within the page only, keeping offsets
# absolute and the page number fixed.
for doc in self.splitter.create_documents([page_text]):
sub_start = start + doc.metadata.get("start_index", 0)
chunks.append(
ChunkWithPosition(
text=doc.page_content,
start_offset=sub_start,
end_offset=sub_start + len(doc.page_content),
page_number=page,
)
)
return chunks
+30 -8
View File
@@ -30,7 +30,10 @@ from nextcloud_mcp_server.observability.metrics import (
from nextcloud_mcp_server.observability.tracing import trace_operation
from nextcloud_mcp_server.search.pdf_highlighter import PDFHighlighter
from nextcloud_mcp_server.vector import payload_keys
from nextcloud_mcp_server.vector.document_chunker import DocumentChunker
from nextcloud_mcp_server.vector.document_chunker import (
DocumentChunker,
PageAwareChunker,
)
from nextcloud_mcp_server.vector.html_processor import html_to_markdown
from nextcloud_mcp_server.vector.placeholder import (
delete_placeholder_point,
@@ -682,27 +685,46 @@ async def _index_document(
logger.error("Failed to process file %s: %s", file_path, e)
raise
# Tokenize and chunk (using configured chunk size and overlap)
# Tokenize and chunk (using configured chunk size and overlap). Paginated
# files (PDFs with page_boundaries) use the page-aware chunker when enabled,
# which assigns page numbers inline; everything else uses the char-based
# chunker followed by post-hoc page assignment.
page_boundaries = file_metadata.get("page_boundaries")
use_page_aware = (
settings.document_chunk_page_aware
and doc_task.doc_type == "file"
and page_boundaries is not None
)
with trace_operation(
"vector_sync.chunk_text",
attributes={
"vector_sync.input_chars": len(content),
"vector_sync.chunk_size": settings.document_chunk_size,
"vector_sync.overlap": settings.document_chunk_overlap,
"vector_sync.page_aware": use_page_aware,
},
) as chunk_span:
chunker = DocumentChunker(
if use_page_aware:
page_boundaries_list = cast(list[dict[str, Any]], page_boundaries)
chunks = await PageAwareChunker(
chunk_size=settings.document_chunk_size,
overlap=settings.document_chunk_overlap,
)
chunks = await chunker.chunk_text(content)
).chunk_text(content, page_boundaries_list)
else:
chunks = await DocumentChunker(
chunk_size=settings.document_chunk_size,
overlap=settings.document_chunk_overlap,
).chunk_text(content)
record_document_chunks(doc_task.doc_type, len(chunks))
if chunk_span is not None:
chunk_span.set_attribute(_ATTR_CHUNK_COUNT, len(chunks))
# Assign page numbers to chunks if page boundaries are available (PDFs)
page_boundaries = file_metadata.get("page_boundaries")
if doc_task.doc_type == "file" and page_boundaries is not None:
# Assign page numbers for the char-based path (page-aware already sets them).
if (
not use_page_aware
and doc_task.doc_type == "file"
and page_boundaries is not None
):
# Type narrowing: page_boundaries is guaranteed to be list[dict] here
page_boundaries_list = cast(list[dict[str, Any]], page_boundaries)
with trace_operation(
+14
View File
@@ -167,6 +167,20 @@ class TestChunkConfigValidation:
assert settings.document_chunk_size == 2048
assert settings.document_chunk_overlap == 200
def test_page_aware_enabled_by_default(self):
"""Page-aware chunking is on by default."""
assert Settings().document_chunk_page_aware is True
@patch.dict(
os.environ,
{"DOCUMENT_CHUNK_PAGE_AWARE": "false"},
clear=True,
)
def test_page_aware_disabled_via_env(self):
"""DOCUMENT_CHUNK_PAGE_AWARE=false disables page-aware chunking."""
_reload_config()
assert get_settings().document_chunk_page_aware is False
def test_valid_chunk_settings(self):
"""Test valid chunk size and overlap configuration."""
settings = Settings(
+136
View File
@@ -3,9 +3,28 @@
from nextcloud_mcp_server.vector.document_chunker import (
ChunkWithPosition,
DocumentChunker,
PageAwareChunker,
)
def _make_doc(pages: list[str]) -> tuple[str, list[dict]]:
"""Build (full_text, page_boundaries) the way the PDF extractors do.
Page texts are concatenated with no separator and boundaries index exactly
into the result (the pypdfium2_fast contract).
"""
content = ""
boundaries: list[dict] = []
offset = 0
for i, text in enumerate(pages, start=1):
boundaries.append(
{"page": i, "start_offset": offset, "end_offset": offset + len(text)}
)
content += text
offset += len(text)
return content, boundaries
class TestDocumentChunkerPositions:
"""Test suite for DocumentChunker position tracking functionality."""
@@ -286,3 +305,120 @@ Fourth paragraph here."""
overlap_text = content[overlap_start:overlap_end]
assert overlap_text in chunks[i].text
assert overlap_text in chunks[i + 1].text
class TestPageAwareChunker:
"""Test suite for the page-aware chunker."""
async def test_one_chunk_per_page_when_chunk_size_exceeds_pages(self):
"""chunk_size >= largest page => exactly one chunk per page."""
pages = [
"Page one content.",
"Page two has rather more text than the first page does.",
"Third.",
]
content, boundaries = _make_doc(pages)
chunks = await PageAwareChunker(chunk_size=2048, overlap=200).chunk_text(
content, boundaries
)
# Predictable vector count == page count.
assert len(chunks) == len(pages)
for i, chunk in enumerate(chunks):
assert chunk.page_number == i + 1
# No leading/trailing whitespace in these pages -> text == page text.
assert chunk.text == pages[i]
# Offsets remain exact against the original document.
assert content[chunk.start_offset : chunk.end_offset] == chunk.text
async def test_no_chunk_spans_a_page_boundary(self):
"""Every chunk's character range stays within one page (the invariant)."""
pages = [
"Short.",
"word " * 200, # oversized page -> will be split within the page
"Another short page of text.",
" ", # blank page
"Final page content here.",
]
content, boundaries = _make_doc(pages)
chunks = await PageAwareChunker(chunk_size=200, overlap=20).chunk_text(
content, boundaries
)
for chunk in chunks:
assert chunk.page_number is not None
pb = boundaries[chunk.page_number - 1]
assert pb["start_offset"] <= chunk.start_offset
assert chunk.end_offset <= pb["end_offset"]
assert content[chunk.start_offset : chunk.end_offset] == chunk.text
async def test_oversized_page_splits_others_stay_single(self):
"""Only the oversized page yields multiple chunks; page numbers fixed."""
pages = ["small page one", "word " * 200, "small page three"]
content, boundaries = _make_doc(pages)
chunks = await PageAwareChunker(chunk_size=200, overlap=20).chunk_text(
content, boundaries
)
per_page = {1: 0, 2: 0, 3: 0}
for chunk in chunks:
per_page[chunk.page_number] += 1
assert per_page[1] == 1
assert per_page[2] > 1
assert per_page[3] == 1
async def test_blank_pages_skipped(self):
"""Whitespace-only pages produce no chunks (no wasted embeddings)."""
pages = ["Real content here.", " \n ", "More real content."]
content, boundaries = _make_doc(pages)
chunks = await PageAwareChunker(chunk_size=2048, overlap=200).chunk_text(
content, boundaries
)
assert len(chunks) == 2
assert {c.page_number for c in chunks} == {1, 3}
async def test_offsets_tightened_around_page_whitespace(self):
"""Leading/trailing page whitespace is stripped and offsets adjusted."""
pages = [" Leading and trailing. ", "Normal page."]
content, boundaries = _make_doc(pages)
chunks = await PageAwareChunker(chunk_size=2048, overlap=200).chunk_text(
content, boundaries
)
first = chunks[0]
assert first.text == "Leading and trailing."
assert content[first.start_offset : first.end_offset] == first.text
async def test_no_page_boundaries_falls_back_to_char_chunking(self):
"""Without page boundaries, behaves like the char-based chunker."""
content = "This is sentence one. " * 40
pa_chunks = await PageAwareChunker(chunk_size=100, overlap=20).chunk_text(
content, []
)
char_chunks = await DocumentChunker(chunk_size=100, overlap=20).chunk_text(
content
)
assert len(pa_chunks) > 1
# Same chunk boundaries as the char-based path, and no page numbers.
assert [(c.text, c.start_offset) for c in pa_chunks] == [
(c.text, c.start_offset) for c in char_chunks
]
assert all(c.page_number is None for c in pa_chunks)
async def test_empty_content_returns_single_empty_chunk(self):
"""Empty content returns one empty chunk regardless of boundaries."""
chunks = await PageAwareChunker().chunk_text(
"", [{"page": 1, "start_offset": 0, "end_offset": 0}]
)
assert len(chunks) == 1
assert chunks[0].text == ""
assert chunks[0].start_offset == 0
assert chunks[0].end_offset == 0