fix(document): timeout reason bucket + Sonar https hotspot (#892 round 2)

Round-2 review on PR #892:
- OcrProcessor.process now catches TimeoutError separately and returns
  parse_failed_reason="timeout" with a populated message ("OCR timed out after
  Ns"), instead of conflating timeouts with API errors under "error" and logging
  an empty suffix. Lets dashboards tell a too-low timeout from a failing
  provider. Test added.
- Add validator-rejection tests for DOCUMENT_OCR_TIMEOUT_SECONDS=0 (gte=1) and
  DOCUMENT_MAX_PDF_SIZE_MB=-1 (gte=0), matching the existing validator-test
  pattern.
- Comment the _Settings test fixture's max_pdf_size_mb=0.0 default.

SonarCloud: quality gate was failing on new_security_hotspots_reviewed (S5332
"use https") from an http:// URL in the new gateway-timeout test — switched to
https:// (mirrors commit 98c9d58e).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-11 06:12:21 +02:00
co-authored by Claude Opus 4.8
parent 64ea5c8631
commit 2f8875e736
4 changed files with 52 additions and 1 deletions
@@ -264,6 +264,22 @@ class OcrProcessor(DocumentProcessor):
text, boundaries = await backend.ocr(
content, content_type.split(";")[0].strip().lower()
)
except TimeoutError:
# anyio.fail_after / httpx read-timeout raise TimeoutError with an
# empty message; give it its own reason bucket and a useful log so a
# too-low DOCUMENT_OCR_TIMEOUT_SECONDS is distinguishable from a
# provider that's actually erroring.
timeout = settings.document_ocr_timeout_seconds
logger.warning(
"OCR timed out for %s after %.1fs", filename or "<bytes>", timeout
)
return ProcessingResult(
text="",
metadata={"parse_failed_reason": "timeout"},
processor=self.name,
success=False,
error=f"OCR timed out after {timeout:.1f}s",
)
except Exception as e:
logger.warning("OCR failed for %s: %s", filename or "<bytes>", e)
return ProcessingResult(