Files
mcp-nextcloud/docs/ADR-026-pluggable-database-backend.md
T
Chris CoutinhoandClaude Opus 4.7 f2b7bf132f fix(storage): address PR #798 review feedback (credentials, asyncpg extra, TLS, pool)
Round-2 fixes after the bot review on PR #798 plus two user follow-ups
(self-signed Postgres support; asyncpg should be a PyPI extra). Folded
into the same PR rather than a follow-up since the work is still
unmerged.

Security
--------
- Mask database credentials in all 5 log call sites (storage.py × 4,
  migrations.py × 1) via a new `mask_db_password()` helper in config.py.
  Uses SQLAlchemy's `make_url(...).render_as_string(hide_password=True)`
  with a regex fallback so the masking path never raises.
- New `tests/unit/test_storage_logging.py` asserts a sentinel password
  never appears in `caplog` during `RefreshTokenStorage.initialize()`.

Distribution
------------
- `asyncpg` moved to `[project.optional-dependencies] postgres` so a
  vanilla `pip install nextcloud-mcp-server` no longer pulls in the
  ~5 MB C extension. The Docker image runs `uv sync --extra postgres`,
  so containerized deployments are unchanged.
- When `DATABASE_URL=postgresql+asyncpg://...` is set on a venv missing
  the extra, `RefreshTokenStorage.initialize()` raises a friendly
  RuntimeError pointing at `[postgres]` rather than the generic
  ModuleNotFoundError.

TLS for the Postgres backend
----------------------------
- New `DATABASE_VERIFY_SSL` + `DATABASE_CA_BUNDLE` env vars mirror the
  existing `NEXTCLOUD_VERIFY_SSL` / `NEXTCLOUD_CA_BUNDLE` pattern
  (validators in Settings.__post_init__, `get_database_ssl()` helper
  alongside `get_nextcloud_ssl_verify()`). `DATABASE_VERIFY_SSL=false`
  wins over `DATABASE_CA_BUNDLE` for incident-response convenience.
- Default is **None** rather than True — keeps PR #798's behavior
  intact for cluster-internal Postgres that runs without TLS. Operators
  opt into verify-full or supply a private CA. ADR-026 records the
  reasoning vs the Nextcloud HTTPS default.
- Engine factory in `storage.py` passes `ssl` via `connect_args` only
  when `get_database_ssl()` returns non-None; otherwise asyncpg's
  default (`prefer`) applies.
- Storage logs which TLS mode is active at INFO (no secret material).

Configurable connection pool
----------------------------
- `DATABASE_POOL_SIZE` (default 10) and `DATABASE_MAX_OVERFLOW`
  (default 20) replace the hardcoded engine values. With many replicas
  this can blow past managed-Postgres `max_connections=100`; tune down
  for large fleets.
- gte-1 / gte-0 validators in __post_init__ reject 0/negative pool
  sizes at startup with the offending value in the error.

Consistency polish
------------------
- Migration 006: convert raw `op.execute("ALTER TABLE ... ADD COLUMN")`
  to `op.batch_alter_table(...).add_column(sa.Column("nonce", sa.Text))`
  for stylistic consistency with the rewritten 001-005. Downgrade now
  drops the column instead of being a no-op.
- `registered_webhooks.created_at` standardized from `sa.Float` to
  `sa.BigInteger` (all other `*_at` columns); `store_webhook()` casts
  `time.time()` → `int`.
- `is_sqlite_url()` made case-insensitive.

Testing
-------
- New `tests/integration/test_storage_postgres.py::test_cleanup_expired_roundtrip`
  exercises `cleanup_expired_tokens`, `cleanup_expired_sessions`, and
  `cleanup_expired_browser_sessions` — relies on DELETE rowcount,
  historically dialect-tricky.
- `tests/unit/test_ssl_config.py` extended with `TestDatabaseSSLSettings`
  + `TestGetDatabaseSSL` classes (9 new tests) mirroring the existing
  Nextcloud SSL tests one-for-one.

Docs
----
- `docs/configuration.md` Centralized-Storage section grew the four new
  env vars + a homelab example with a private CA.
- `docs/ADR-026` grew Distribution, TLS, and `alembic/env.py` async-pattern
  subsections explaining the non-obvious design choices.

Helm chart counterpart in cbcoutinho/helm-charts PR #34 (separate
commit on `feat/nextcloud-mcp-server-database-url`).

Verification
------------
- `uv run pytest tests/unit/` — 1025 passed.
- `TEST_DATABASE_URL=... uv run pytest tests/integration/test_storage_postgres.py -m postgres` — 6 passed (including new cleanup test).
- `uv run ruff check && uv run ruff format --check && uv run ty check -- nextcloud_mcp_server` — clean.

Tracked on Astrolabe Cloud POC board, card #99.

---

_This PR was generated with the help of AI, and reviewed by a Human_

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 18:53:45 +02:00

11 KiB

ADR-026: Pluggable database backend (DATABASE_URL)

Status

Accepted — 2026-05-16

Context

RefreshTokenStorage (in nextcloud_mcp_server/auth/storage.py) holds all of the MCP server's persistent state: refresh tokens, OAuth client credentials, OAuth sessions, browser sessions, app passwords, login-flow sessions, audit logs, and webhook registrations. Until this ADR it was backed by a single SQLite file, with the path configured by TOKEN_STORAGE_DB.

This works well for single-user deployments but blocks horizontal scaling in Kubernetes:

  • Every pod needs its own PVC (ReadWriteOnce) or a ReadWriteMany volume.
  • Tokens stored on pod A are invisible to pod B, so a Service can only route traffic to one pod at a time.
  • Restart / re-deploy cycles either drop the volume (token loss) or require coordinated PVC handling.
  • Backup, encryption-at-rest, and multi-region replication become per-pod concerns rather than centrally managed DB concerns.

We needed a way for pods to be stateless and share a centralized store without giving up the zero-config SQLite path that single-user installs and local development rely on.

Decision

Introduce a DATABASE_URL setting that accepts any SQLAlchemy async URL, with sqlite+aiosqlite:///... remaining the default. The runtime keeps a single linear migration history and a single RefreshTokenStorage class — the backend is selected purely by the URL.

Resolution order

get_database_url() (in nextcloud_mcp_server/config.py) returns:

  1. DATABASE_URL if set — wins over everything.
  2. Otherwise sqlite+aiosqlite:///{get_token_db_path()}, so the legacy TOKEN_STORAGE_DB env var and the process-local ephemeral tempfile fallback both keep working unchanged.

Why SQLAlchemy Core + async engine, not an ABC with parallel drivers

Two alternatives were considered:

Option Why rejected
Define a Storage ABC with SQLiteStorage (aiosqlite) and PostgresStorage (asyncpg) implementations Doubles the surface area — every schema change has to land in two backends, with two sets of migrations, two SQL dialects, two upsert idioms. Diverges over time.
Switch to a full SQLAlchemy ORM (declarative models) Larger refactor; the existing explicit-SQL style is intentional and well-understood by reviewers.
Keep RefreshTokenStorage and put SQLAlchemy Core under it (chosen) One method body per operation, one migration history (Alembic is already SQLAlchemy-based). The URL drives dialect, pool, and DDL.

A thin compatibility shim (_DBConn / _Cursor / _Row in storage.py) adapts the async with aiosqlite.connect(...) as db: async with db.execute(...) as cursor: ... idiom to SQLAlchemy AsyncEngine / AsyncConnection. Existing method bodies needed only their connection context-manager swapped; ? placeholders are rewritten to named binds on the fly. The seven INSERT OR REPLACE statements were rewritten as portable INSERT ... ON CONFLICT (...) DO UPDATE (SQLite ≥ 3.24, Postgres ≥ 9.5; we already require SQLite ≥ 3.35 elsewhere).

Distribution: asyncpg is an optional extra, bundled in Docker

asyncpg carries a compiled C extension (~5 MB plus a build toolchain on source installs) — too heavy a default for the pip install nextcloud-mcp-server audience, the majority of whom run the SQLite path. It is shipped as a PyPI optional dependency::

pip install 'nextcloud-mcp-server[postgres]'

The published Docker image runs uv sync --extra postgres so the container always has the driver, matching the HA-deployment audience that exercises the Postgres backend. When DATABASE_URL=postgresql+asyncpg://... is set on a venv without the extra installed, RefreshTokenStorage raises a clear actionable error before the engine is built — operators see "install with [postgres] extra" rather than a generic ModuleNotFoundError: No module named 'asyncpg'.

Alembic env.py runs the async engine inside a worker thread

nextcloud_mcp_server/alembic/env.py uses async_engine_from_config(...) + anyio.run(run_async_migrations), and the runtime invokes it from RefreshTokenStorage.initialize() via anyio.to_thread.run_sync(upgrade_database, ...). This is intentional:

  • Alembic wants a synchronous entry point (upgrade_database()), but async_engine_from_config returns an async engine.
  • Running anyio.run() directly inside an already-running event loop would deadlock; we have to be on a different thread.
  • to_thread.run_sync puts the call on a worker thread, which has no running event loop — anyio.run() is then free to spin up its own.

The pattern is non-obvious; this note exists so a future maintainer doesn't try to "simplify" it back into the main loop.

TLS for the Postgres backend

Two settings mirror the existing NEXTCLOUD_VERIFY_SSL / NEXTCLOUD_CA_BUNDLE pattern: DATABASE_VERIFY_SSL and DATABASE_CA_BUNDLE. get_database_ssl() (in nextcloud_mcp_server/config.py) returns the value to pass to asyncpg via SQLAlchemy's connect_args={"ssl": ...}.

The default is deliberately less strict than the Nextcloud HTTPS default: DATABASE_VERIFY_SSL defaults to None rather than True. When both env vars are unset we omit the ssl kwarg entirely and asyncpg's default (prefer) applies — TLS if the server offers it, no certificate validation. The reasoning:

  • Cluster-internal Postgres (CNPG via a Service, RDS over a private VPC, PgBouncer sidecar) is the common HA pattern and frequently runs without TLS or with cert hostnames asyncpg wouldn't match anyway.
  • The HTTPS analogy doesn't carry over: the Nextcloud client talks to external hostnames over public networks where verify-full is the right default. The database client talks to a controlled peer.
  • Just-shipped PR #798 had no TLS knobs and worked against cluster-local Postgres-test; flipping the default to True here would break that flow on upgrade.

Operators in production with a managed Postgres opt in with DATABASE_VERIFY_SSL=true. Homelab operators with a private CA set DATABASE_CA_BUNDLE=/path/to/ca.pem (which implies verify=true). DATABASE_VERIFY_SSL=false is the escape hatch for incident response — it wins over DATABASE_CA_BUNDLE so an operator can quickly silence cert errors without editing the secret store.

Encryption stays in Python (Fernet), not the DB

The DB only ever sees ciphertext for sensitive columns (encrypted_token, encrypted_client_secret, encrypted_password, encrypted_poll_token). The Fernet key remains a TOKEN_ENCRYPTION_KEY env var, applied in Python before INSERT and after SELECT. This means:

  • Switching backends does not invalidate or re-key existing data.
  • Postgres-level features like pgcrypto are not required.
  • Operators rotating the encryption key still go through the existing Python path.

DDL portability

All Alembic migrations were rewritten from raw op.execute("CREATE TABLE ...") strings to op.create_table() / op.create_index() calls with SQLAlchemy types. Notable choices:

  • All *_at / expiration / timestamp columns use sa.BigInteger — Postgres INTEGER is 32-bit and unix epochs are already past that range. SQLite treats BIGINT and INTEGER identically (dynamic typing) so this is backwards compatible.
  • BLOBsa.LargeBinary (becomes BYTEA on Postgres).
  • BOOLEAN DEFAULT FALSEsa.Boolean, server_default=sa.false().
  • Existing SQLite deployments are at revision 006 and skip the rewritten migrations entirely — content rewrites are safe.

No data migration, no shipped Postgres

Two scope decisions worth recording:

  1. Clean cutover, no SQLite → Postgres data migration tool. Tokens are reissued on the next login; webhooks re-register on the next sync tick. Acceptable because the ephemeral-default already implies this, and the data being preserved (audit logs, OAuth sessions) is either short-lived or reconstructible.
  2. Bring-your-own database. The MCP server consumes a DATABASE_URL; it does not provision Postgres itself. Operators use CNPG, RDS, the project's existing Helm chart with a sub-chart, etc. The postgres-test service in docker-compose.yml exists only for integration tests and manual HA smoke testing — it is gated on the postgres profile and is not the recommended production pattern.

CLI changes

The nextcloud-mcp-server db {upgrade,downgrade,current,history} commands gain a --database-url / -u flag (env DATABASE_URL) alongside the existing --database-path / -d (env TOKEN_STORAGE_DB). -u wins over -d; both fall back to get_database_url().

Consequences

Positive

  • MCP server pods become stateless. A Kubernetes Deployment can run with replicas: 3 behind a Service, with all pods pointed at the same Postgres URL — tokens written by pod A are immediately visible to pod B.
  • Centralized DB operations (backup, restore, replication, encryption at rest, monitoring) are handled by the operator's existing Postgres infrastructure rather than duplicated per-pod.
  • No regression for single-user / local-development / docker-compose installs — the SQLite tempfile path is unchanged and remains the default when DATABASE_URL is unset.
  • Test coverage doubles automatically: every test that uses the temp_storage fixture now runs against both SQLite and Postgres when TEST_DATABASE_URL is exported.

Negative

  • One more thing operators have to think about for HA deployments (Postgres connection string, credentials secret, network policies).
  • Adds SQLAlchemy + asyncpg to the runtime dependency set. SQLAlchemy was already transitively present via Alembic; asyncpg is genuinely new.
  • The compatibility shim in storage.py is a small piece of bespoke code that future contributors need to understand. The alternative — rewriting every method body to SQLAlchemy idioms — was rejected as too risky for this PR but might be revisited.

Neutral

  • The Alembic migration history was content-rewritten but its revision graph is unchanged (still 001 → 006), so existing SQLite deployments do not re-run anything.
  • TOKEN_STORAGE_DB still works exactly as before; deployments that already set it require no changes.

Verification

  1. uv run pytest tests/unit/ — SQLite path unchanged (1012 tests).
  2. docker compose --profile postgres up -d postgres-test then TEST_DATABASE_URL=postgresql+asyncpg://mcp:mcp@localhost:5433/mcp uv run pytest tests/unit/test_app_password_storage.py tests/unit/test_webhook_storage.py — every test runs once per backend.
  3. Manual end-to-end smoke against mcp-login-flow with a Postgres URL (commands in /home/chris/.claude/plans/spicy-enchanting-flurry.md → Verification).
  4. k8s HA validation (after merge in homelab-argocd): replicas: 3, confirm session continuity through the Service.