refactor(search): address PR #750 round 3 review feedback
- _verify_deck_cards: hoist int(board_id|stack_id|doc_id) out of the generic except Exception into an explicit try/except (TypeError, ValueError) before the network call, mirroring _verify_news_items. Malformed payloads now log a specific warning instead of "unexpected error". - _verify_news_items: add TODO(perf) above the get_items(batch_size=-1) call to mark the known fetch-all cost as a future profiling target. - SemanticSearchResult.id: revert from int|str back to int. The internal SearchResult.id stays int|str for forward-compat; the MCP response model narrows at the boundary. server/semantic.py casts r.id to int when constructing the response so future string-id types fail loudly here instead of silently widening the public API. - nc_semantic_search: replace the terse "extra for access filtering" comment with an ADR-019 NOTE block explaining the 2x over-fetch trade-off and the ghost-density under-delivery case (self-heals via lazy eviction). - tests/integration/test_verify_on_read.py: extend the module docstring to call out that only the note verifier is exercised against real Nextcloud, while file/deck_card/news_item are unit-only — documenting the suite split for future contributors. - ADR-019: rewrite "Module shape", "Verifier registry", example verifier, and "Deduplication" sections to match the shipped BatchVerifier interface (was per-id Verifier in the original draft). Add a "Why batch?" paragraph explaining the design choice. Update implementation checklist — every item is now [x] with corrected verifier names (plural) and the eviction module path (vector/eviction.py). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
21e5608a39
commit
aa4b9498a1
@@ -189,12 +189,31 @@ async def _verify_deck_cards(
|
||||
accessible.add(doc_id)
|
||||
return
|
||||
|
||||
# Parse defensively before the network call so a malformed payload
|
||||
# produces a specific log line, not a generic "unexpected error" from
|
||||
# the catch-all ``except Exception`` below. Mirrors ``_verify_news_items``.
|
||||
try:
|
||||
board_id_int = int(board_id)
|
||||
stack_id_int = int(stack_id)
|
||||
card_id_int = int(doc_id)
|
||||
except (TypeError, ValueError) as e:
|
||||
logger.warning(
|
||||
"Non-numeric deck metadata for card %s "
|
||||
"(board_id=%r, stack_id=%r): %s; keeping result",
|
||||
doc_id,
|
||||
board_id,
|
||||
stack_id,
|
||||
e,
|
||||
)
|
||||
accessible.add(doc_id)
|
||||
return
|
||||
|
||||
async with semaphore:
|
||||
try:
|
||||
await client.deck.get_card(
|
||||
board_id=int(board_id),
|
||||
stack_id=int(stack_id),
|
||||
card_id=int(doc_id),
|
||||
board_id=board_id_int,
|
||||
stack_id=stack_id_int,
|
||||
card_id=card_id_int,
|
||||
)
|
||||
accessible.add(doc_id)
|
||||
except HTTPStatusError as e:
|
||||
@@ -236,6 +255,11 @@ async def _verify_news_items(
|
||||
|
||||
async with semaphore:
|
||||
try:
|
||||
# TODO(perf): if profiling shows this fetch dominates query latency
|
||||
# for news-heavy users, cache the per-request item set or push for
|
||||
# a per-item News API endpoint. The shared semaphore protects
|
||||
# against runaway concurrent fetches, but the payload itself can
|
||||
# be large (News auto-purge cap is in the thousands of items).
|
||||
items = await client.news.get_items(batch_size=-1, get_read=True)
|
||||
except HTTPStatusError as e:
|
||||
# If the News API itself is gone (app disabled, user lost access),
|
||||
|
||||
Reference in New Issue
Block a user