fix: make procrastinate ingest queue opt-in (default to in-process anyio)
An unset INGEST_QUEUE auto-derived "postgres" whenever DATABASE_URL was PostgreSQL, silently starting the procrastinate ingest worker (schema migration, reclaim cron, deferred jobs) on every Postgres-backed tenant — even though none had opted into the api/worker split. Observed on tenant-blackbox-demo (:0.98.0): ~600 "Deferred 1 job" log lines / 24h. Resolve an unset INGEST_QUEUE to "memory" (the in-process anyio queue) regardless of the database backend. procrastinate is now strictly opt-in via an explicit INGEST_QUEUE=postgres; the existing guard still rejects postgres against a SQLite DATABASE_URL. Docs + unit test updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
1e615c2bf1
commit
ad211ee2da
+11
-5
@@ -799,9 +799,10 @@ EMBEDDING_GATEWAY_TOKEN_URL=...
|
||||
EMBEDDING_GATEWAY_CLIENT_ID=...
|
||||
EMBEDDING_GATEWAY_CLIENT_SECRET=...
|
||||
|
||||
# Ingest queue backend. Default (unset) auto-derives from DATABASE_URL:
|
||||
# - PostgreSQL DATABASE_URL → "postgres" (the procrastinate queue)
|
||||
# - SQLite / unset → "memory" (the in-process anyio queue)
|
||||
# Ingest queue backend. Default (unset) is "memory" — the in-process anyio
|
||||
# queue — *regardless of DATABASE_URL*. procrastinate is strictly opt-in: set
|
||||
# INGEST_QUEUE=postgres to split ingest into a separate worker (requires a
|
||||
# PostgreSQL DATABASE_URL). A Postgres DATABASE_URL alone never enables it.
|
||||
INGEST_QUEUE=postgres # memory | postgres
|
||||
# Process role (informational; the worker is launched via the `worker` command):
|
||||
MCP_ROLE=all # api | worker | all (default)
|
||||
@@ -810,8 +811,13 @@ TENANT_ID=<uuid> # per-tenant identity (used in collection naming)
|
||||
|
||||
### Postgres ingest queue + worker (api/worker split)
|
||||
|
||||
When `INGEST_QUEUE=postgres` (a PostgreSQL `DATABASE_URL`), the scanner **defers**
|
||||
one job per changed document into the app's Postgres via
|
||||
This is **opt-in**. By default (`INGEST_QUEUE=memory`) the scanner processes
|
||||
changed documents in-process via anyio task groups in the API pod — no
|
||||
procrastinate, no separate worker, even when `DATABASE_URL` is Postgres.
|
||||
|
||||
When you explicitly set `INGEST_QUEUE=postgres` (against a PostgreSQL
|
||||
`DATABASE_URL`), the scanner instead **defers** one job per changed document
|
||||
into the app's Postgres via
|
||||
[procrastinate](https://procrastinate.readthedocs.io); a separate **worker**
|
||||
process drains the queue (fetch → chunk → embed → upsert Qdrant). Run the two
|
||||
roles as separate Deployments from the same image:
|
||||
|
||||
Reference in New Issue
Block a user