# ADR-026: Pluggable database backend (DATABASE_URL) ## Status Accepted — 2026-05-16 ## Context `RefreshTokenStorage` (in `nextcloud_mcp_server/auth/storage.py`) holds all of the MCP server's persistent state: refresh tokens, OAuth client credentials, OAuth sessions, browser sessions, app passwords, login-flow sessions, audit logs, and webhook registrations. Until this ADR it was backed by a single SQLite file, with the path configured by `TOKEN_STORAGE_DB`. This works well for single-user deployments but blocks horizontal scaling in Kubernetes: - Every pod needs its own PVC (ReadWriteOnce) or a ReadWriteMany volume. - Tokens stored on pod A are invisible to pod B, so a Service can only route traffic to one pod at a time. - Restart / re-deploy cycles either drop the volume (token loss) or require coordinated PVC handling. - Backup, encryption-at-rest, and multi-region replication become per-pod concerns rather than centrally managed DB concerns. We needed a way for pods to be stateless and share a centralized store without giving up the zero-config SQLite path that single-user installs and local development rely on. ## Decision Introduce a `DATABASE_URL` setting that accepts any SQLAlchemy async URL, with `sqlite+aiosqlite:///...` remaining the default. The runtime keeps a single linear migration history and a single `RefreshTokenStorage` class — the backend is selected purely by the URL. ### Resolution order `get_database_url()` (in `nextcloud_mcp_server/config.py`) returns: 1. `DATABASE_URL` if set — wins over everything. 2. Otherwise `sqlite+aiosqlite:///{get_token_db_path()}`, so the legacy `TOKEN_STORAGE_DB` env var and the process-local ephemeral tempfile fallback both keep working unchanged. ### Why SQLAlchemy Core + async engine, not an ABC with parallel drivers Two alternatives were considered: | Option | Why rejected | |---|---| | Define a `Storage` ABC with `SQLiteStorage` (aiosqlite) and `PostgresStorage` (asyncpg) implementations | Doubles the surface area — every schema change has to land in two backends, with two sets of migrations, two SQL dialects, two upsert idioms. Diverges over time. | | Switch to a full SQLAlchemy ORM (declarative models) | Larger refactor; the existing explicit-SQL style is intentional and well-understood by reviewers. | | **Keep `RefreshTokenStorage` and put SQLAlchemy Core under it** *(chosen)* | One method body per operation, one migration history (Alembic is already SQLAlchemy-based). The URL drives dialect, pool, and DDL. | A thin compatibility shim (`_DBConn` / `_Cursor` / `_Row` in `storage.py`) adapts the `async with aiosqlite.connect(...) as db: async with db.execute(...) as cursor: ...` idiom to SQLAlchemy `AsyncEngine` / `AsyncConnection`. Existing method bodies needed only their connection context-manager swapped; `?` placeholders are rewritten to named binds on the fly. The seven `INSERT OR REPLACE` statements were rewritten as portable `INSERT ... ON CONFLICT (...) DO UPDATE` (SQLite ≥ 3.24, Postgres ≥ 9.5; we already require SQLite ≥ 3.35 elsewhere). ### Distribution: asyncpg is an optional extra, bundled in Docker `asyncpg` carries a compiled C extension (~5 MB plus a build toolchain on source installs) — too heavy a default for the `pip install nextcloud-mcp-server` audience, the majority of whom run the SQLite path. It is shipped as a PyPI optional dependency:: pip install 'nextcloud-mcp-server[postgres]' The published Docker image runs `uv sync --extra postgres` so the container always has the driver, matching the HA-deployment audience that exercises the Postgres backend. When `DATABASE_URL=postgresql+asyncpg://...` is set on a venv without the extra installed, `RefreshTokenStorage` raises a clear actionable error before the engine is built — operators see "install with `[postgres]` extra" rather than a generic `ModuleNotFoundError: No module named 'asyncpg'`. ### Alembic env.py runs the async engine inside a worker thread `nextcloud_mcp_server/alembic/env.py` uses `async_engine_from_config(...)` + `anyio.run(run_async_migrations)`, and the runtime invokes it from `RefreshTokenStorage.initialize()` via `anyio.to_thread.run_sync(upgrade_database, ...)`. This is intentional: - Alembic wants a synchronous entry point (`upgrade_database()`), but `async_engine_from_config` returns an async engine. - Running `anyio.run()` directly inside an already-running event loop would deadlock; we have to be on a different thread. - `to_thread.run_sync` puts the call on a worker thread, which has no running event loop — `anyio.run()` is then free to spin up its own. The pattern is non-obvious; this note exists so a future maintainer doesn't try to "simplify" it back into the main loop. ### TLS for the Postgres backend Two settings mirror the existing `NEXTCLOUD_VERIFY_SSL` / `NEXTCLOUD_CA_BUNDLE` pattern: `DATABASE_VERIFY_SSL` and `DATABASE_CA_BUNDLE`. `get_database_ssl()` (in `nextcloud_mcp_server/config.py`) returns the value to pass to asyncpg via SQLAlchemy's `connect_args={"ssl": ...}`. The default is deliberately **less strict than the Nextcloud HTTPS default**: `DATABASE_VERIFY_SSL` defaults to `None` rather than `True`. When both env vars are unset we omit the `ssl` kwarg entirely and asyncpg's default (`prefer`) applies — TLS if the server offers it, no certificate validation. The reasoning: - Cluster-internal Postgres (CNPG via a Service, RDS over a private VPC, PgBouncer sidecar) is the common HA pattern and frequently runs without TLS or with cert hostnames asyncpg wouldn't match anyway. - The HTTPS analogy doesn't carry over: the Nextcloud client talks to *external* hostnames over public networks where verify-full is the right default. The database client talks to a controlled peer. - Just-shipped PR #798 had no TLS knobs and worked against cluster-local Postgres-test; flipping the default to `True` here would break that flow on upgrade. Operators in production with a managed Postgres opt in with `DATABASE_VERIFY_SSL=true`. Homelab operators with a private CA set `DATABASE_CA_BUNDLE=/path/to/ca.pem` (which implies `verify=true`). `DATABASE_VERIFY_SSL=false` is the escape hatch for incident response — it wins over `DATABASE_CA_BUNDLE` so an operator can quickly silence cert errors without editing the secret store. ### Encryption stays in Python (Fernet), not the DB The DB only ever sees ciphertext for sensitive columns (`encrypted_token`, `encrypted_client_secret`, `encrypted_password`, `encrypted_poll_token`). The Fernet key remains a `TOKEN_ENCRYPTION_KEY` env var, applied in Python before INSERT and after SELECT. This means: - Switching backends does not invalidate or re-key existing data. - Postgres-level features like `pgcrypto` are not required. - Operators rotating the encryption key still go through the existing Python path. ### DDL portability All Alembic migrations were rewritten from raw `op.execute("CREATE TABLE ...")` strings to `op.create_table()` / `op.create_index()` calls with SQLAlchemy types. Notable choices: - All `*_at` / expiration / timestamp columns use `sa.BigInteger` — Postgres `INTEGER` is 32-bit and unix epochs are already past that range. SQLite treats `BIGINT` and `INTEGER` identically (dynamic typing) so this is backwards compatible. - `BLOB` → `sa.LargeBinary` (becomes `BYTEA` on Postgres). - `BOOLEAN DEFAULT FALSE` → `sa.Boolean, server_default=sa.false()`. - Existing SQLite deployments are at revision `006` and skip the rewritten migrations entirely — content rewrites are safe. ### No data migration, no shipped Postgres Two scope decisions worth recording: 1. **Clean cutover, no SQLite → Postgres data migration tool.** Tokens are reissued on the next login; webhooks re-register on the next sync tick. Acceptable because the ephemeral-default already implies this, and the data being preserved (audit logs, OAuth sessions) is either short-lived or reconstructible. 2. **Bring-your-own database.** The MCP server consumes a `DATABASE_URL`; it does not provision Postgres itself. Operators use CNPG, RDS, the project's existing Helm chart with a sub-chart, etc. The `postgres-test` service in `docker-compose.yml` exists only for integration tests and manual HA smoke testing — it is gated on the `postgres` profile and is not the recommended production pattern. ### CLI changes The `nextcloud-mcp-server db {upgrade,downgrade,current,history}` commands gain a `--database-url / -u` flag (env `DATABASE_URL`) alongside the existing `--database-path / -d` (env `TOKEN_STORAGE_DB`). `-u` wins over `-d`; both fall back to `get_database_url()`. ## Consequences ### Positive - MCP server pods become stateless. A Kubernetes Deployment can run with `replicas: 3` behind a Service, with all pods pointed at the same Postgres URL — tokens written by pod A are immediately visible to pod B. - Centralized DB operations (backup, restore, replication, encryption at rest, monitoring) are handled by the operator's existing Postgres infrastructure rather than duplicated per-pod. - No regression for single-user / local-development / docker-compose installs — the SQLite tempfile path is unchanged and remains the default when `DATABASE_URL` is unset. - Test coverage doubles automatically: every test that uses the `temp_storage` fixture now runs against both SQLite and Postgres when `TEST_DATABASE_URL` is exported. ### Negative - One more thing operators have to think about for HA deployments (Postgres connection string, credentials secret, network policies). - Adds SQLAlchemy + asyncpg to the runtime dependency set. SQLAlchemy was already transitively present via Alembic; asyncpg is genuinely new. - The compatibility shim in `storage.py` is a small piece of bespoke code that future contributors need to understand. The alternative — rewriting every method body to SQLAlchemy idioms — was rejected as too risky for this PR but might be revisited. ### Neutral - The Alembic migration history was content-rewritten but its revision graph is unchanged (still `001 → 006`), so existing SQLite deployments do not re-run anything. - `TOKEN_STORAGE_DB` still works exactly as before; deployments that already set it require no changes. ## Related - [ADR-022 Login Flow v2](ADR-022-deployment-mode-consolidation.md) — defines the per-user app password storage that this ADR centralizes. - [ADR-002 Vector sync authentication](ADR-002-vector-sync-authentication.md) — explains the offline-access tokens that benefit most from HA storage. ## Verification 1. `uv run pytest tests/unit/` — SQLite path unchanged (1012 tests). 2. `docker compose --profile postgres up -d postgres-test` then `TEST_DATABASE_URL=postgresql+asyncpg://mcp:mcp@localhost:5433/mcp uv run pytest tests/unit/test_app_password_storage.py tests/unit/test_webhook_storage.py` — every test runs once per backend. 3. Manual end-to-end smoke against `mcp-login-flow` with a Postgres URL (commands in `/home/chris/.claude/plans/spicy-enchanting-flurry.md` → Verification). 4. k8s HA validation (after merge in `homelab-argocd`): `replicas: 3`, confirm session continuity through the Service.