test(usage): close round-5 nits (empty doc_types, consistency tidy-ups)

Round-5 claude-review (merge-ready; all nits):

- 🟡 Added test_empty_doc_types_normalizes_to_null pinning doc_types=[] → None
  in record_search_usage metadata (matches the None case).
- 🟡 record_search_usage docstring now notes nc_semantic_search_answer always
  meters with doc_types=None (it exposes no doc_types parameter).
- 🟢 BM25HybridSearchAlgorithm.__init__ now sets query_embedding /
  query_token_count alongside _embedded_query, so all three cache fields are
  instance attributes from construction (was relying on the class-level
  SearchAlgorithm defaults).
- 🟢 Ollama embed_batch_with_usage caches _dimension inline (mirrors
  OpenAI/Mistral), so the dimension is set via any embed path.
- 🟢 record_indexing_usage documents the independent-record / partial-failure
  semantics under SUM aggregation.

Deck #67.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-08 01:53:26 +02:00
co-authored by Claude Opus 4.8
parent df03d33fd4
commit 141663bb07
5 changed files with 34 additions and 4 deletions
@@ -75,6 +75,21 @@ async def test_none_token_count_records_zero(store_spy):
assert kwargs["metadata"]["doc_types"] is None
@pytest.mark.unit
async def test_empty_doc_types_normalizes_to_null(store_spy):
"""An empty doc_types list normalizes to None, same as a None input, so a
metadata->'doc_types' IS NULL query counts the all-types case consistently."""
await semantic.record_search_usage(
enabled=True,
user_id="alice",
fusion="rrf",
doc_types=[],
token_count=5,
)
kwargs = store_spy.record_usage_event.await_args.kwargs
assert kwargs["metadata"]["doc_types"] is None
@pytest.mark.unit
async def test_doc_types_metadata_is_bounded(store_spy):
"""A large doc_types list is truncated to the metadata cap."""