fix(document-processors): escalate glyph-corrupt PDFs to the structured tier

The fast (pypdfium2) extractor can leak raw glyph codes on subset fonts with a
broken /ToUnicode CMap. The result scores high on the existing text-quality
heuristic -- a uniform glyph/Caesar offset preserves whitespace and token
lengths -- yet is unsearchable. The structured (pymupdf) tier extracts the same
pages correctly.

Add a language-agnostic C0-control-character-ratio signal to the tier-0
classifier that detects this corruption and routes the document to a new
`structured` recommended_tier. Wire the fast->structured hop on the inline path
and generalise it so a low-quality-but-non-empty layer also tries structured
before OCR -- the inline and external ingest modes now follow the full
fast->structured->ocr ladder identically. A scanned / no-text-layer document
(total_chars == 0) still shortcuts straight to OCR, since a text extractor
cannot recover a pure raster.

New per-tenant tunable DOCUMENT_GLYPH_CORRUPTION_RATIO (default 0.02); escalation
metrics gain a `corrupt_glyphs` reason label.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-16 20:06:35 +02:00
co-authored by Claude Opus 4.8
parent 60df141442
commit cf7209cd85
8 changed files with 412 additions and 54 deletions
@@ -27,9 +27,10 @@ Two entry points:
Recommended tier:
* ``ocr`` -- scanned / no-usable-text-layer (route to tier 3, when enabled)
* ``structured`` -- a text layer that is present but glyph-corrupt (the fast
extractor leaked raw glyph codes; high C0-control-char ratio). A different
in-cluster extractor (the pymupdf ``structured`` tier) recovers it -- no OCR.
* ``fast`` -- a usable digital text layer (stay on tier 1)
``structured`` (tier 2 / docling) is a separate service, not produced here.
"""
import logging
@@ -56,9 +57,20 @@ MIN_TEXT_QUALITY = 0.5
OCR_PAGE_FRACTION = 0.5
# A page with fewer extracted chars than this has effectively no text layer.
MIN_PAGE_CHARS = 16
# Doc-level control-character ratio above which the text layer is treated as
# glyph-corrupt and routed to the ``structured`` (pymupdf) tier, which re-extracts
# such PDFs correctly. Kept in sync with the DOCUMENT_GLYPH_CORRUPTION_RATIO
# setting default (the registry passes the per-tenant value). See
# ``_control_char_ratio``.
GLYPH_CORRUPTION_RATIO = 0.02
_WORD_RE = re.compile(r"\S+")
# Whitespace control characters that legitimately appear in extracted text
# (tab / newline / carriage-return / form-feed / vertical-tab). Every OTHER C0
# control char is a corruption signal -- see ``_control_char_ratio``.
_TEXT_WHITESPACE_CONTROLS = frozenset("\t\n\r\f\v")
@dataclass
class PageSignals:
@@ -67,6 +79,7 @@ class PageSignals:
image_coverage: float # 0..1 of page area covered by images
text_quality: float # 0..1; low = mashed/space-less/garbage layer
needs_ocr: bool # scanned or unusable text layer
control_ratio: float = 0.0 # 0..1; high = corrupt/glyph-leak text layer
@dataclass
@@ -76,10 +89,11 @@ class DocClassification:
total_chars: int
mean_text_quality: float
ocr_page_fraction: float # fraction of sampled pages flagged needs_ocr
recommended_tier: str # "fast" | "ocr"
recommended_tier: str # "fast" | "structured" | "ocr"
mean_control_ratio: float = 0.0 # doc-level C0-control-char ratio (glyph-leak)
flags: set[str] = field(
default_factory=set
) # scanned | bad_text_layer | image_heavy
) # scanned | bad_text_layer | image_heavy | corrupt_glyphs
pages: list[PageSignals] = field(default_factory=list)
@@ -116,6 +130,25 @@ def _text_quality(text: str) -> float:
return round(ws_score * len_score * overlong_score * merge_score, 3)
def _control_char_ratio(text: str) -> float:
"""Fraction of C0 control characters (excluding whitespace controls) in ``text``.
Near 0 for clean text in ANY script; elevated when the extractor leaked raw
glyph codes instead of Unicode -- the broken-/ToUnicode failure mode where a
subset font's character codes are returned uniformly offset (e.g. "WKH" for
"THE"). This is the language-agnostic counterpart to :func:`_text_quality`: a
uniform glyph/Caesar offset preserves whitespace and token lengths (so every
``_text_quality`` factor scores it ~1.0), but it litters the text with C0
controls -- digits/punctuation map to bytes below 0x20 -- which clean prose
never contains. Unlike a dictionary or stop-word probe it makes no assumption
about the document's language.
"""
if not text:
return 0.0
bad = sum(1 for c in text if ord(c) < 0x20 and c not in _TEXT_WHITESPACE_CONTROLS)
return bad / len(text)
def _sample_indices(page_count: int) -> list[int]:
if page_count <= MAX_SAMPLED_PAGES:
return list(range(page_count))
@@ -143,6 +176,56 @@ def _page_image_coverage(page: Any) -> float:
return min(img_area / page_area, 1.0)
def _route_from_signals(
*,
total_chars: int,
ocr_frac: float,
mean_quality: float,
control_ratio: float,
image_heavy: bool,
page_fraction: float,
min_text_quality: float,
glyph_corruption_ratio: float,
) -> tuple[set[str], str]:
"""Shared flag-set + recommended-tier decision for both classifier paths.
Routing precedence (cheapest correct fix first):
1. scanned / no text layer (``ocr_frac >= fraction`` AND ``total_chars == 0``)
-> ``"ocr"``
2. glyph-corrupt text layer (``control_ratio > glyph_corruption_ratio``)
-> ``"structured"``. pypdfium2 leaked glyph codes; the pymupdf
``structured`` tier re-extracts these correctly, so no paid OCR is
needed. The registry re-classifies the structured output, so a doc that
is ALSO partly scanned can still escalate to OCR from there.
3. junk/mashed text layer (``ocr_frac >= fraction``) -> ``"ocr"``
4. otherwise -> ``"fast"``
Flags are diagnostic and independent of the verdict (e.g. ``image_heavy``
fires on ANY image-heavy page; the OCR route needs a page FRACTION).
"""
glyph_corrupt = total_chars > 0 and control_ratio > glyph_corruption_ratio
flags: set[str] = set()
if ocr_frac >= page_fraction and total_chars == 0:
flags.add("scanned")
elif ocr_frac >= page_fraction and mean_quality < min_text_quality:
flags.add("bad_text_layer")
if glyph_corrupt:
flags.add("corrupt_glyphs")
if image_heavy:
flags.add("image_heavy")
if ocr_frac >= page_fraction and total_chars == 0:
recommended = "ocr"
elif glyph_corrupt:
recommended = "structured"
elif ocr_frac >= page_fraction:
recommended = "ocr"
else:
recommended = "fast"
return flags, recommended
def classify_pdf(content: bytes) -> DocClassification:
"""Classify a PDF from its bytes.
@@ -171,7 +254,14 @@ def classify_pdf(content: bytes) -> DocClassification:
# raises the diagnostic image_heavy flag below.
needs_ocr = quality < MIN_TEXT_QUALITY or len(text.strip()) < MIN_PAGE_CHARS
pages.append(
PageSignals(n, len(text), round(coverage, 3), quality, needs_ocr)
PageSignals(
n,
len(text),
round(coverage, 3),
quality,
needs_ocr,
round(_control_char_ratio(text), 4),
)
)
sampled = len(pages)
@@ -180,25 +270,24 @@ def classify_pdf(content: bytes) -> DocClassification:
round(sum(p.text_quality for p in pages) / sampled, 3) if sampled else 0.0
)
ocr_frac = (sum(p.needs_ocr for p in pages) / sampled) if sampled else 0.0
# Char-weighted doc-level control-char ratio (p.control_ratio * char_count is
# the per-page bad-char count). The glyph-leak signal -- see _control_char_ratio.
control_ratio = (
sum(p.control_ratio * p.char_count for p in pages) / total_chars
if total_chars
else 0.0
)
# Flags are diagnostic signals, intentionally independent of the routing
# verdict: image_heavy fires if ANY page is image-heavy, while the OCR route
# needs a FRACTION of pages (OCR_PAGE_FRACTION). So a mostly-digital doc with
# one full-page photo is flagged image_heavy yet still routes "fast" -- the
# flag_total{image_heavy} count is expected to exceed classified{ocr}.
flags: set[str] = set()
if any(p.image_coverage >= IMAGE_HEAVY_THRESHOLD for p in pages):
flags.add("image_heavy")
if (
ocr_frac >= OCR_PAGE_FRACTION
and total_chars
and mean_quality < MIN_TEXT_QUALITY
):
flags.add("bad_text_layer")
if ocr_frac >= OCR_PAGE_FRACTION and total_chars == 0:
flags.add("scanned")
recommended = "ocr" if ocr_frac >= OCR_PAGE_FRACTION else "fast"
flags, recommended = _route_from_signals(
total_chars=total_chars,
ocr_frac=ocr_frac,
mean_quality=mean_quality,
control_ratio=control_ratio,
image_heavy=any(p.image_coverage >= IMAGE_HEAVY_THRESHOLD for p in pages),
page_fraction=OCR_PAGE_FRACTION,
min_text_quality=MIN_TEXT_QUALITY,
glyph_corruption_ratio=GLYPH_CORRUPTION_RATIO,
)
return DocClassification(
page_count=page_count,
@@ -207,6 +296,7 @@ def classify_pdf(content: bytes) -> DocClassification:
mean_text_quality=mean_quality,
ocr_page_fraction=round(ocr_frac, 3),
recommended_tier=recommended,
mean_control_ratio=round(control_ratio, 4),
flags=flags,
pages=pages,
)
@@ -239,6 +329,7 @@ def classify_from_text(
min_text_quality: float = MIN_TEXT_QUALITY,
min_page_chars: int = MIN_PAGE_CHARS,
page_fraction: float = OCR_PAGE_FRACTION,
glyph_corruption_ratio: float = GLYPH_CORRUPTION_RATIO,
image_coverage: list[float] | None = None,
) -> DocClassification:
"""Classify from text already extracted by tier-1 -- no PDF re-open by default.
@@ -246,9 +337,13 @@ def classify_from_text(
The hot-path classifier. A page is OCR-worthy when its text is near-empty
(``< min_page_chars``) or its text-quality is junk (``< min_text_quality`` --
the word-merging signal). The doc recommends ``ocr`` once
``ocr_frac >= page_fraction``. Thresholds are passed in by the registry from
per-tenant settings. ``image_coverage`` (when supplied) only feeds the
``image_heavy`` diagnostic flag -- it does NOT route (see module docstring).
``ocr_frac >= page_fraction``. A doc whose text layer is present but
glyph-corrupt (doc-level C0-control-char ratio ``> glyph_corruption_ratio``,
the broken-/ToUnicode failure mode) instead recommends ``structured`` -- the
pymupdf tier re-extracts it correctly, no OCR needed. Thresholds are passed in
by the registry from per-tenant settings. ``image_coverage`` (when supplied)
only feeds the ``image_heavy`` diagnostic flag -- it does NOT route (see
module docstring).
``page_boundaries`` are ``{page, start_offset, end_offset}`` indexing into
``full_text``; ``image_coverage[i]`` (if given) aligns with the i-th boundary.
@@ -293,7 +388,14 @@ def classify_from_text(
# ``cov`` still feeds the diagnostic ``image_heavy`` flag below.
needs_ocr = len(seg.strip()) < min_page_chars or quality < min_text_quality
pages.append(
PageSignals(b["page"], len(seg), round(cov, 3), quality, needs_ocr)
PageSignals(
b["page"],
len(seg),
round(cov, 3),
quality,
needs_ocr,
round(_control_char_ratio(seg), 4),
)
)
sampled = len(pages)
@@ -305,23 +407,21 @@ def classify_from_text(
# page_count guard also skips escalation; defaulting to 0.0 keeps the
# recorded classification metric accurate rather than a misleading "ocr").
ocr_frac = (sum(p.needs_ocr for p in pages) / sampled) if sampled else 0.0
# Doc-level (char-weighted) control-char ratio -- the glyph-leak signal that
# routes to the structured tier. Computed over full_text so it is robust to
# boundary edge cases.
control_ratio = _control_char_ratio(full_text)
# Flags gated on ocr_frac >= page_fraction (matching classify_pdf): a doc that
# routes "fast" must not carry a junk-layer flag just because a few isolated
# pages are bad -- otherwise the metric diverges from classify_pdf.
flags: set[str] = set()
if sampled and ocr_frac >= page_fraction:
if total_chars == 0:
# "scanned" (not "no_text_layer"): same name + meaning as classify_pdf
# so astrolabe_document_classifier_flag_total isn't split across two
# labels for the empty-text-layer case.
flags.add("scanned")
elif mean_quality < min_text_quality:
flags.add("bad_text_layer")
if any(p.image_coverage >= IMAGE_HEAVY_THRESHOLD for p in pages):
flags.add("image_heavy")
recommended = "ocr" if ocr_frac >= page_fraction else "fast"
flags, recommended = _route_from_signals(
total_chars=total_chars,
ocr_frac=ocr_frac,
mean_quality=mean_quality,
control_ratio=control_ratio,
image_heavy=any(p.image_coverage >= IMAGE_HEAVY_THRESHOLD for p in pages),
page_fraction=page_fraction,
min_text_quality=min_text_quality,
glyph_corruption_ratio=glyph_corruption_ratio,
)
return DocClassification(
page_count=len(page_boundaries),
@@ -330,6 +430,7 @@ def classify_from_text(
mean_text_quality=mean_quality,
ocr_page_fraction=round(ocr_frac, 3),
recommended_tier=recommended,
mean_control_ratio=round(control_ratio, 4),
flags=flags,
pages=pages,
)
@@ -49,7 +49,7 @@ class EscalationDecision:
kind: Literal["hop", "suppressed"]
to_tier: str
reason: Literal["empty_text", "low_confidence"]
reason: Literal["empty_text", "low_confidence", "corrupt_glyphs"]
def next_tier(current: str) -> str | None:
@@ -80,10 +80,11 @@ class EscalateError(Exception):
the junk text is never indexed, and it must never be swallowed by a broad
``except Exception`` on the indexing path.
``reason`` uses the existing escalation label vocabulary. This PR raises
``empty_text`` (scanned / no text layer) and ``low_confidence`` (junk text
layer); ``unsupported`` and ``forced`` are reserved for future callers and
not raised yet.
``reason`` uses the existing escalation label vocabulary: ``empty_text``
(scanned / no text layer), ``low_confidence`` (junk text layer), and
``corrupt_glyphs`` (a usable-looking layer whose extractor leaked raw glyph
codes -- the broken-/ToUnicode case -- recovered by a different in-cluster
extractor); ``unsupported`` and ``forced`` are reserved for future callers.
"""
def __init__(self, *, from_tier: str, to_tier: str, reason: str) -> None:
@@ -247,6 +247,63 @@ class ProcessorRegistry:
result, content, settings, record=True, filename=filename
)
# Escalate a poor fast extraction up the ladder (fast -> structured -> ocr),
# mirroring the external per-tier path so both modes behave identically. A
# glyph-corrupt layer (the extractor leaked raw glyph codes -- the
# broken-/ToUnicode case) OR a low-quality-but-non-empty layer first tries
# the structured (pymupdf) tier: free, in-cluster, and able to recover both.
# Only a scanned / no-text-layer doc (total_chars == 0) skips structured --
# a text extractor cannot conjure text from a pure raster -- and drops
# straight to OCR via the gate below. Structured is therefore NOT gated on
# document_ocr_enabled. Its output is re-classified (record=False -- the doc
# was already counted at the fast tier) so a doc that is ALSO partly scanned
# still reaches the OCR gate.
if (
classification is not None
and classification.page_count > 0
and (
classification.recommended_tier == "structured"
or (
classification.recommended_tier == "ocr"
and classification.total_chars > 0
)
)
):
structured = self._pdf_processor_for_tier("structured")
if structured is not None:
reason = (
"corrupt_glyphs"
if classification.recommended_tier == "structured"
else "low_confidence"
)
record_document_escalation("fast", "structured", reason)
logger.info(
"Escalating %s fast->structured (reason=%s)",
filename or "<bytes>",
reason,
)
structured_result = await self._run_processor(
structured,
content,
content_type,
filename,
options,
progress_callback,
escalated=True,
)
if structured_result.success:
result = structured_result
classification = self._classify_result(
result, content, settings, record=False, filename=filename
)
else:
logger.warning(
"structured escalation did not succeed for %s (%s); keeping "
"the tier-1 result",
filename or "<bytes>",
structured_result.metadata.get("parse_failed_reason", "error"),
)
# NOTE: the suppressed-escalation metric (document_escalation_suppressed_total,
# the "what-if OCR" signal; Deck #324) is intentionally NOT emitted on this
# inline/memory path -- it is instrumented only on the per-tier external
@@ -379,6 +436,7 @@ class ProcessorRegistry:
min_text_quality=settings.document_ocr_min_text_quality,
min_page_chars=settings.document_ocr_min_page_chars,
page_fraction=settings.document_ocr_page_fraction,
glyph_corruption_ratio=settings.document_glyph_corruption_ratio,
image_coverage=image_coverage,
)
except Exception:
@@ -520,6 +578,9 @@ class ProcessorRegistry:
- ``total_chars == 0`` (scanned / no text layer) -> target the ``ocr``
tier directly. Text-extractor tiers (``structured``) cannot conjure
text from a pure raster scan, so a structured hop would just be wasted.
- glyph-corrupt text layer (``recommended_tier == "structured"``) -> target
the ``structured`` tier; pymupdf re-extracts a broken-/ToUnicode layer
correctly, so OCR is never the target for this case.
- low-confidence but non-empty layer -> escalate to the next rung, so a
different in-cluster extractor can try before paying for OCR.
@@ -538,13 +599,23 @@ class ProcessorRegistry:
record=(current_tier == TIER_LADDER[0]),
filename=filename,
)
if classification is None or classification.recommended_tier != "ocr":
if classification is None or classification.recommended_tier not in (
"structured",
"ocr",
):
return None
# A zero-page (empty/corrupt) PDF gains nothing from any tier.
if classification.page_count <= 0:
return None
if classification.total_chars == 0:
minimum: str | None = "ocr"
minimum: str | None
if classification.recommended_tier == "structured":
# Glyph-corrupt text layer (the extractor leaked glyph codes): a
# different in-cluster extractor (the structured/pymupdf tier) recovers
# it -- never pay for OCR here. Target the structured rung specifically.
minimum = "structured"
reason = "corrupt_glyphs"
elif classification.total_chars == 0:
minimum = "ocr"
reason = "empty_text"
else:
minimum = None