fix(ocr): round-2 review — lazy store lock, mode enum normalization, type hints

Round 2 review (PR #910):
- BLOCKING: BatchOcrJobStore._shared_lock is now lazy-init (anyio.Lock | None,
  created on first shared() call) instead of at class-definition time — matches
  the CLAUDE.md "no anyio primitives at import time" rule and OcrProcessor's
  pattern. The None-check->assign has no await between, so it's race-free.
- document_ocr_mode now normalizes via _enum_fields (case-insensitive, like
  document_ocr_provider) instead of a strict dynaconf is_in Validator, so
  DOCUMENT_OCR_MODE=Batch normalizes to "batch" rather than erroring. Tests for
  case-normalization + invalid-value rejection.
- TYPE_CHECKING-gated GatewayBatchOcrClient import so build_gateway_batch_client
  / _get_batch_client are typed `GatewayBatchOcrClient | None` instead of Any
  (runtime import stays lazy to avoid the import cycle).
- Rename ocr_options -> doc_identity_options (it's threaded to all tiers; only
  OCR reads it) + clarify the comment.
- Drop the redundant forward-ref quotes on _shared_instance.
- Add direct _batch_identity unit tests (partial/empty options branches).

Left as follow-up: reusing one httpx.AsyncClient across submit/poll (same
per-call pattern as the existing sync _GatewayOcrBackend; no clean aclose hook
on the cached client today).

1653 unit tests pass; ruff + ty green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-15 10:45:10 +02:00
co-authored by Claude Opus 4.8
parent 2b7dfc8535
commit 995e810d89
6 changed files with 75 additions and 16 deletions
+29
View File
@@ -225,6 +225,35 @@ async def test_mistral_backend_applies_timeout(mocker, monkeypatch):
# --- batch mode (Deck #332) --------------------------------------------------
@pytest.mark.parametrize(
"options",
[
None,
{},
{"doc_id": "d", "doc_type": "file"}, # missing user_id
{"user_id": "u", "doc_type": "file"}, # missing doc_id
{"user_id": "u", "doc_id": "d"}, # missing doc_type
{"user_id": "u", "doc_id": "d", "doc_type": ""}, # empty doc_type
],
)
def test_batch_identity_returns_none_without_full_identity(options):
assert ocr._batch_identity(options) is None
def test_batch_identity_extracts_tuple_and_defaults_etag():
assert ocr._batch_identity(
{"user_id": "u", "doc_id": "d", "doc_type": "file", "etag": "v1"}
) == ("u", "d", "file", "v1")
# etag may be absent/empty -> normalised to "".
assert ocr._batch_identity({"user_id": "u", "doc_id": "d", "doc_type": "file"}) == (
"u",
"d",
"file",
"",
)
from nextcloud_mcp_server.embedding.gateway_batch_client import ( # noqa: E402
BatchPollResult,
)