fix(vector): normalize doc_id to str + add Qdrant keyword payload indexes

Production was logging two cascading classes of Qdrant errors against the
welcomed-malamute deployment:

1. HTTP 400 — "Bad request: Index required but not found for \"doc_id\" of
   one of the following types: [keyword]". The collection was created via
   create_collection() with no payload indexes, so any FieldCondition
   filter on doc_id failed at the Qdrant layer (placeholder writes/reads,
   eviction, search context lookups).

2. Compounding the missing index, producers wrote a mix of int and str
   doc_ids: webhook_parser stringified node_id, scanner stringified note
   IDs, news IDs, and deck card IDs — but the file scanner passed the
   numeric file_id through unchanged. A keyword index would not have
   covered both kinds even if it had existed.

This change:

- Normalizes doc_id to str at every producer site (scanner.py:459,
  DocumentTask.doc_id, indexed_*_ids reads from Qdrant).
- Tightens str|int annotations to str across placeholder.py,
  eviction.py, search/verification.py, search/context.py,
  SearchResult.id, and the auth/api visualization endpoints.
- Defensive str() coercion on doc_id reads in semantic.py /
  bm25_hybrid.py / vector/visualization.py for the transition window
  before the backfill runs.
- Adds an idempotent startup migration in get_qdrant_client():
  - _ensure_keyword_payload_indexes creates KEYWORD indexes for
    doc_id, user_id, and doc_type (tolerates "already exists" 400s).
  - _backfill_doc_id_to_string scrolls the collection once and rewrites
    int doc_ids to str. Skipped after a quick sample shows no legacy
    int payloads.
- Public API preserved: SemanticSearchResult.id stays int via explicit
  int(r.id) narrowing in server/semantic.py — surfaces a TypeError with
  actionable context if a future doc_type ships non-numeric ids.
- Documents the startup migration in docs/configuration.md.

Tests: 11 new unit tests in tests/unit/vector/test_qdrant_client.py
covering happy path / already-exists / unrelated-400 for the index
helpers, and sample-skip / mixed-batch rewrite / payload=None edge cases
for the backfill. 889 unit tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-05-08 19:30:51 +02:00
co-authored by Claude Opus 4.7
parent 0690378915
commit 719b3b5034
17 changed files with 493 additions and 66 deletions
+9 -7
View File
@@ -224,12 +224,11 @@ def configure_semantic_tools(mcp: FastMCP):
search_results = verified_results[:limit]
# Convert SearchResult objects to SemanticSearchResult for response.
# SearchResult.id is typed `int | str` for forward-compat with future
# doc_types, but every currently indexed type uses numeric ids and
# the MCP response model narrows to `int`. Casting here makes the
# narrowing explicit and surfaces any future string-id type as a
# loud failure at the boundary instead of silently widening the
# public API.
# SearchResult.id is `str` (Qdrant keyword-indexed payload), but
# every currently indexed type uses numeric ids and the MCP response
# model narrows to `int`. Casting here makes the narrowing explicit
# and surfaces any future non-numeric-id type as a loud failure at
# the boundary instead of silently widening the public API.
results = []
for r in search_results:
try:
@@ -304,7 +303,10 @@ def configure_semantic_tools(mcp: FastMCP):
chunk_context = await get_chunk_with_context(
nc_client=client,
user_id=username,
doc_id=result.id,
# SemanticSearchResult.id is the int-narrowed
# public form; get_chunk_with_context queries
# Qdrant where doc_id is keyword-indexed as str.
doc_id=str(result.id),
doc_type=result.doc_type,
chunk_start=result.chunk_start_offset,
chunk_end=result.chunk_end_offset,