Pulls the remaining review feedback into one commit:
- Remove the double _normalize_contact_data call: the helper now assumes
canonical keys, and create_contact normalises before calling it
(update_contact already did). Docstring states the invariant.
- Drop phone/organization from _SUPPORTED_CONTACT_KEYS; they never reach
the unknown-key check post-normalisation.
- Tighten generics to dict[str, Any] / list[str] across helpers and
ContactsClient signatures.
- Comment both URL-merge sites noting only the first URL is written.
- Log a warning when fn is missing from contact_data.
- Test coverage for _wrap_contact_field dropping value-less dicts and
for the fn-missing warning; _vcard helper now mirrors the real call
chain (normalise → build).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PR #719 review raised a claim that these three fields fall through to
the "keep unchanged" catch-all in _merge_vcard_properties. The existing
elif branches for NICKNAME/BDAY/CATEGORIES already prevent that, but
the behaviour wasn't pinned by a test. Add a focused TestMergeVcardProperties
class that calls the merge helper directly and asserts:
- Existing NICKNAME/BDAY/CATEGORIES lines are overwritten by new values.
- When the existing vCard has none of these lines, update adds them.
- A URL update doesn't clobber unrelated ORG/NOTE/TEL properties.
If the primary update path ever regresses for these fields, these tests
will catch it immediately.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Type-annotate _wrap_contact_field signature; drop stale "url" mention
from its docstring (url is handled by the list-coercion helper, not
this one).
- Split the shape-coercion helper so comma-splitting only applies to
categories: _as_str_list (no split) for org/nickname/url,
_split_categories (comma split) for CATEGORIES. Fixes the case where
organization="Smith, Jones & Associates" was mangled into a two-
component ORG.
- Share _normalize_contact_data between create and update so
_merge_vcard_properties only sees canonical keys; add URL handlers in
both update branches so the primary update path no longer drops URL
silently.
- Annotate the Contact(**kwargs) type:ignore with the reason
(pythonvCard4 typeshed doesn't accept **dict[str, Any]).
- Add tests/unit/client/test_contacts.py (pure unit, no HTTP) covering
the #716 round-trip, comma-in-org regression, invalid-bday warning,
tel/phone precedence, categories string-vs-list behaviour, and direct
_normalize_contact_data cases.
- Extend the MCP workflow test with an update-with-url step asserting
the URL handler in _merge_vcard_properties.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Fixed 8 type checker errors across the codebase:
- vector/scanner.py: Handle None scroll results with null-safe iteration
- search/{bm25_hybrid,semantic}.py: Add None checks for result.payload
- auth/{unified_verifier,webhook_routes}.py: Assert non-None auth credentials
- client/webdav.py: Add None checks before int() conversions
- providers/openai.py: Assert embedding_model is not None
- search/algorithms.py: Explicitly type doc_types set and cast values
- observability/logging_config.py: Match parent class signature (log_data)
Also fixed test_create_tag_creates_system_tag to match WebDAV implementation
(was testing OCS API endpoint, now tests correct WebDAV endpoint with
Content-Location header).
Type checker: 0 errors (down from 8), 20 warnings (ignored)
Tests: All 192 unit tests passing
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
- Add get_file_info() to get file info including file ID via PROPFIND
- Add create_tag() to create system tags via OCS API
- Add get_or_create_tag() for idempotent tag creation
- Add assign_tag_to_file() to assign tags to files via WebDAV
- Add remove_tag_from_file() to remove tags from files
Also refactors RAG evaluation:
- Add indexed_manual_pdf fixture using existing nc_client/nc_mcp_client
- Remove manual tag creation steps from workflow (now handled by fixture)
- Add comprehensive unit tests for new WebDAV methods
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
This commit fixes two critical issues with PDF processing:
1. **Text extraction mismatch (context expansion bug)**:
- Indexing used pymupdf4llm.to_markdown() producing markdown text
- Context expansion used page.get_text() producing plain text
- Different text formats caused character offset misalignment
- Search would find correct chunk, but expansion showed wrong section
- Fixed by making context.py use pymupdf4llm.to_markdown() consistently
2. **Diagnostic logging for page number assignment**:
- Added logging to verify page_boundaries exist in metadata
- Added logging to verify assign_page_numbers() assigns values
- Helps diagnose why page numbers show as null in search results
3. **mime_type storage bug**:
- Fixed incorrect field reference in processor.py:405
- Was using file_metadata.get("content_type", "")
- Should use content_type from WebDAV response
Changes:
- nextcloud_mcp_server/search/context.py: Use pymupdf4llm.to_markdown()
for PDF text extraction to match indexing method
- nextcloud_mcp_server/vector/processor.py: Add diagnostic logging for
page boundaries and assignment, fix mime_type storage
- tests/unit/client/test_webdav.py: Fix import sorting
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>