test(usage): close round-5 nits (empty doc_types, consistency tidy-ups)

Round-5 claude-review (merge-ready; all nits):

- 🟡 Added test_empty_doc_types_normalizes_to_null pinning doc_types=[] → None
  in record_search_usage metadata (matches the None case).
- 🟡 record_search_usage docstring now notes nc_semantic_search_answer always
  meters with doc_types=None (it exposes no doc_types parameter).
- 🟢 BM25HybridSearchAlgorithm.__init__ now sets query_embedding /
  query_token_count alongside _embedded_query, so all three cache fields are
  instance attributes from construction (was relying on the class-level
  SearchAlgorithm defaults).
- 🟢 Ollama embed_batch_with_usage caches _dimension inline (mirrors
  OpenAI/Mistral), so the dimension is set via any embed path.
- 🟢 record_indexing_usage documents the independent-record / partial-failure
  semantics under SUM aggregation.

Deck #67.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-08 01:53:26 +02:00
co-authored by Claude Opus 4.8
parent df03d33fd4
commit 141663bb07
5 changed files with 34 additions and 4 deletions
+5
View File
@@ -159,6 +159,11 @@ class OllamaProvider(Provider):
data = response.json()
all_embeddings.extend(data["embeddings"])
# Cache the dimension inline (mirrors OpenAI/Mistral) so it is set
# via any embed path, not only an explicit _detect_dimension() call.
if self._dimension is None and data["embeddings"]:
self._dimension = len(data["embeddings"][0])
prompt_eval = data.get("prompt_eval_count")
total_tokens += (
round(prompt_eval)