fix: address PR #836 round-2 review (connect/timeout/observability)
🟡 Document why ProcrastinateTaskProducer.connect() uses `await app.open_async()` (AwaitableContext: await opens a long-lived pool, closed by drain()) and add a connect()/drain() lifecycle unit test (InMemoryConnector) asserting the pool is opened by connect and closed by drain — previously untested. 🟡 get_procrastinate_conninfo: forward connect_timeout from DATABASE_URL or default 10s so an unreachable DB can't hang worker/API startup indefinitely; warn only on other dropped query params. + tests. 🟢 INGEST_DELETE_SUCCEEDED_JOBS (default true) makes the worker's succeeded-job deletion configurable for audit retention. 🟢 Worker startup logs via logger.info (structured/OTel) instead of click.echo. 🟢 INGEST_STALLED_JOB_SECONDS (default 300) makes the crash-reclaim threshold tunable for slow embedding backends; reclaim reads it per-run. The broad `except` in _apply_ingest_queue_schema_open is kept deliberately: procrastinate wraps psycopg errors, so narrowing to psycopg.errors.* would miss the wrapped DDL-conflict and turn a benign concurrent-apply race into a failure; the presence re-check re-raises genuine errors. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
cfdef3c2c5
commit
820b98dac1
@@ -56,11 +56,11 @@ _NAMESPACE = "ingest"
|
||||
INGEST_TASK_NAME = f"{_NAMESPACE}:process_document"
|
||||
|
||||
# A crashed worker leaves its job in ``doing``; reclaim it once its (per-worker)
|
||||
# heartbeat is this many seconds stale. Sized well above the longest expected
|
||||
# ``process_document`` (PDF render + embedding) so a slow-but-live worker — whose
|
||||
# heartbeat stays current during a long job — is never reclaimed out from under
|
||||
# itself.
|
||||
_STALLED_AFTER_SECONDS = 300
|
||||
# heartbeat is this many seconds stale. The default is sized well above the
|
||||
# longest expected ``process_document`` (PDF render + embedding) so a slow-but-
|
||||
# live worker — whose heartbeat stays current during a long job — is never
|
||||
# reclaimed out from under itself. Operators on slow embedding backends can tune
|
||||
# it via INGEST_STALLED_JOB_SECONDS (read per-run in reclaim_stalled_ingest_jobs).
|
||||
|
||||
|
||||
# Tasks are defined as plain functions and registered onto a *fresh* Blueprint
|
||||
@@ -125,9 +125,10 @@ async def reclaim_stalled_ingest_jobs(context: JobContext, timestamp: int) -> No
|
||||
"""
|
||||
manager = context.app.job_manager
|
||||
retry_at = datetime.now(tz=timezone.utc)
|
||||
stalled_after = get_settings().ingest_stalled_job_seconds
|
||||
reclaimed = 0
|
||||
for job in await manager.get_stalled_jobs(
|
||||
queue=INGEST_QUEUE_NAME, seconds_since_heartbeat=_STALLED_AFTER_SECONDS
|
||||
queue=INGEST_QUEUE_NAME, seconds_since_heartbeat=stalled_after
|
||||
):
|
||||
if job.id is None:
|
||||
continue
|
||||
@@ -307,6 +308,12 @@ class ProcrastinateTaskProducer:
|
||||
@classmethod
|
||||
async def connect(cls) -> ProcrastinateTaskProducer:
|
||||
app = get_procrastinate_app()
|
||||
# ``App.open_async()`` returns procrastinate's dual-mode AwaitableContext:
|
||||
# ``await``-ing it opens the connector pool and leaves it open (vs the
|
||||
# ``async with`` form, which closes on block exit). The producer's pool is
|
||||
# long-lived — owned by the server lifespan and torn down once in
|
||||
# ``drain()`` (close_async) on shutdown — so the bare ``await`` is correct
|
||||
# here, unlike the scoped ``async with`` used for one-shot schema apply.
|
||||
await app.open_async()
|
||||
return cls(app)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user