refactor(usage): extract indexing metering helper; address review round 2

Round-2 claude-review findings:

- 🟡 Base-class recursion invariant: documented on embed_with_usage /
  embed_batch_with_usage that a provider overriding embed()/embed_batch() to
  delegate to the *_with_usage variant MUST also override that variant, or the
  two recurse. (No recursion today; the shipped providers pair the overrides.)
- 🟡 Processor metering had no unit test: extracted the two-event recording
  into a module-level record_indexing_usage() helper and added
  tests/unit/test_processor_metering.py (value mapping, flag/zero-chunk no-ops,
  best-effort failure swallowed).
- 🟡 SonarQube hotspots (python:S5332) were 3 http:// URLs in the new test
  fixtures (mock hosts, never contacted) blocking the quality gate
  (new_security_hotspots_reviewed). Switched them to https:// so no hotspot is
  raised.
- 🟢 Zero-chunk guard: record_indexing_usage() no-ops when chunk_count == 0, so
  an empty document no longer writes zero-value billing rows.

Deferred (stated on the PR): Mistral x.index-or-0 sort key (pre-existing,
equivalent), CHANGELOG note for the Ollama /api/embed switch (CHANGELOG is
commitizen-generated from commit bodies, which document it), class-var
query_token_count (safe under the per-request instance pattern).

Deck #67.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-08 01:22:03 +02:00
co-authored by Claude Opus 4.8
parent a0bb5642cb
commit d15ce627ab
5 changed files with 196 additions and 53 deletions
+11
View File
@@ -75,6 +75,12 @@ class Provider(ABC):
usage from their embedding response override this. Used by the
usage-metering hooks (Deck #67) to bill ``embeddings_queries`` by
tokens rather than by operation count.
IMPORTANT (recursion invariant): this default calls ``self.embed``. A
provider that overrides ``embed()`` to delegate to ``embed_with_usage()``
(to avoid duplicating request logic) MUST also override this method, or
the two will call each other forever. The shipped providers that use
that delegation (Bedrock) do override both — keep that pairing.
"""
embedding = await self.embed(text)
return embedding, self._estimate_tokens([text])
@@ -86,6 +92,11 @@ class Provider(ABC):
Returns ``(embeddings, token_count)``; the default estimates. See
:meth:`embed_with_usage`.
IMPORTANT (recursion invariant): this default calls ``self.embed_batch``.
A provider that overrides ``embed_batch()`` to delegate to
``embed_batch_with_usage()`` (Mistral, OpenAI, Ollama do) MUST also
override this method, or the two recurse infinitely. Keep the pairing.
"""
embeddings = await self.embed_batch(texts)
return embeddings, self._estimate_tokens(texts)