docs(review): note image_heavy only fires when scan detection is on

Address PR #863 round 4: classify_from_text's docstring now states that the
image_heavy flag (and the image-coverage trigger) are only set when
image_coverage is supplied, so the flag reads zero for tenants with
DOCUMENT_OCR_DETECT_SCANNED=false -- self-documenting the metric semantics.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-05 05:16:25 +02:00
co-authored by Claude Opus 4.8
parent 820be135fb
commit 36209e3160
@@ -246,6 +246,10 @@ def classify_from_text(
``page_boundaries`` are ``{page, start_offset, end_offset}`` indexing into
``full_text``; ``image_coverage[i]`` (if given) aligns with the i-th boundary.
Note: the ``image_heavy`` flag (and the image-coverage trigger) are only set
when ``image_coverage`` is supplied, so for tenants with scan detection off
that flag is always zero -- the text-quality/empty signals still route.
"""
# image_coverage is expected to be one entry per page, capped at
# MAX_SAMPLED_PAGES (see image_coverage_per_page). Any other length means the