fix(vector): retry transient embed errors so a pod rollover drops 0 docs
From card 309 (OHR-Bench smoke-test triage): during a backend-pod rollover the
embedding endpoint was briefly unreachable, and openai.APIConnectionError /
ConnectError propagated unretried (the provider only retried 429). Documents
exhausted the 3 in-process retries and were dropped for that scan cycle.
Broaden the provider-level retry to the transient set -- APIConnectionError,
APITimeoutError, 429, and 5xx -- on the existing exponential backoff (2s->60s,
5 attempts), so a few seconds of retry rides through the rollover. Permanent
4xx (auth, bad request) still re-raise immediately. Generalize the shared
_retry helper (retry_on_rate_limit -> retry_on_transient, predicate renamed to
should_retry, accurate log label) with a back-compat alias; Mistral gets 429+5xx
for parity. The production gateway path inherits this via GatewayProvider, which
delegates to the decorated OpenAIProvider methods.
Add astrolabe_vector_ingest_dropped_total{reason}, incremented when a document
exhausts retries, classified (connection|timeout|rate_limit|server|qdrant|other)
by _drop_reason so the embed-drop rate is alertable per cause. Dropped docs are
NOT marked failed, so the next full scan re-picks them (re-queue via scan loop).
Refs: Deck board 12 card 309 (AC #1 no permanently-dropped docs; embed-drop
metric for AC #5).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
457c115ef4
commit
258ee96f4c
@@ -13,7 +13,7 @@ import logging
|
||||
from mistralai.client import Mistral
|
||||
from mistralai.client.errors import SDKError
|
||||
|
||||
from ._retry import retry_on_rate_limit
|
||||
from ._retry import retry_on_transient
|
||||
from .base import Provider
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
@@ -30,13 +30,17 @@ BATCH_SIZE = 64
|
||||
_NO_EMBEDDING_MODEL_MSG = "Embedding not supported - no embedding_model configured"
|
||||
|
||||
|
||||
def _is_rate_limit(exc: BaseException) -> bool:
|
||||
"""True only for HTTP 429 SDKErrors."""
|
||||
return getattr(exc, "status_code", None) == 429
|
||||
def _is_transient(exc: BaseException) -> bool:
|
||||
"""Retry HTTP 429 (rate limit) and 5xx (server/transient) SDKErrors."""
|
||||
status = getattr(exc, "status_code", None)
|
||||
return status == 429 or (isinstance(status, int) and status >= 500)
|
||||
|
||||
|
||||
_retry_429 = retry_on_rate_limit(
|
||||
SDKError, is_rate_limit=_is_rate_limit, provider_name="Mistral"
|
||||
_retry_transient = retry_on_transient(
|
||||
SDKError,
|
||||
should_retry=_is_transient,
|
||||
provider_name="Mistral",
|
||||
label="transient error",
|
||||
)
|
||||
|
||||
|
||||
@@ -90,7 +94,7 @@ class MistralProvider(Provider):
|
||||
def supports_generation(self) -> bool:
|
||||
return False
|
||||
|
||||
@_retry_429
|
||||
@_retry_transient
|
||||
async def embed(self, text: str) -> list[float]:
|
||||
"""Generate an embedding for a single text."""
|
||||
if not self.supports_embeddings:
|
||||
@@ -169,7 +173,7 @@ class MistralProvider(Provider):
|
||||
|
||||
return all_embeddings, total_tokens
|
||||
|
||||
@_retry_429
|
||||
@_retry_transient
|
||||
async def _embed_batch_request(
|
||||
self, batch: list[str]
|
||||
) -> tuple[list[list[float]], int]:
|
||||
|
||||
Reference in New Issue
Block a user