feat(usage): rename metrics → tokens_embedded/pages_embedded + export token cost to Prometheus
Billing product model finalized (Deck #281): bill pages externally, record tokens internally. Rename the data-plane metric literals to match the now- canonical contract (Deck #284) — the control plane's METRIC_EVENT_NAMES is already renamed, so the old names would be unmapped and never sync to Stripe. Rename (values unchanged): - embeddings_queries → tokens_embedded (value = real token count, already emitted by this PR; the unit upstream providers bill on). - pages_chunks → pages_embedded (value kept as len(chunk_texts) interim; TODO(#282): real normalized "pages indexed" count — real pages for paginated types, chars/tokens-per-page constant otherwise — is deferred to the instrumentation card, this only lands the name/contract). - All literals, log strings, docstrings, comments, the migration comment, and tests renamed; grep confirms zero old strings remain. Observability (new): export embedding token cost to Prometheus as astrolabe_embedding_tokens_total{provider,operation} (operation = index|query) so the billed cost unit is visible in Grafana, not just the per-tenant billing DB. Dedicated counter (doesn't inflate the existing chunk/request metrics) and always-on (independent of USAGE_METERING_ENABLED, so OSS/self-host gets it). Wired on both the indexing batch embed and the search query embed (query inside the per-request cache-miss branch, so reused embeddings aren't double-counted). Note: the rename orphans any pre-existing embeddings_queries/pages_chunks rows in tenant app DBs (CP no longer maps them) — acceptable; pipeline is inert with throwaway dev/sandbox data. Deck #284 (folded into PR #875). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
ddefb03701
commit
973f80e7b9
@@ -338,6 +338,16 @@ embedding_chars_total = Counter(
|
||||
["kind", "provider"],
|
||||
)
|
||||
|
||||
# Token consumption — the billed cost unit (mirrors the tokens_embedded billing
|
||||
# measure, Deck #67). On a dedicated counter (not folded into the chunk/request
|
||||
# metrics above) so query embeds don't inflate indexing dashboards; labelled by
|
||||
# operation = index | query. Always emitted, independent of USAGE_METERING_ENABLED.
|
||||
embedding_tokens_total = Counter(
|
||||
"astrolabe_embedding_tokens_total",
|
||||
"Total embedding tokens consumed (provider-reported or estimated)",
|
||||
["provider", "operation"], # operation: index | query
|
||||
)
|
||||
|
||||
# --- Chunking & indexed-by-type -----------------------------------------------
|
||||
|
||||
document_chunks_total = Counter(
|
||||
@@ -727,6 +737,25 @@ def record_embedding(
|
||||
embedding_chars_total.labels(kind=kind, provider=provider).inc(chars)
|
||||
|
||||
|
||||
def record_embedding_tokens(provider: str, operation: str, tokens: int) -> None:
|
||||
"""Export embedding token consumption to Prometheus.
|
||||
|
||||
Mirrors the ``tokens_embedded`` billing measure (Deck #67) as an always-on
|
||||
observability signal — emitted regardless of ``USAGE_METERING_ENABLED`` so
|
||||
OSS/self-host deployments still see token cost in Grafana.
|
||||
|
||||
Args:
|
||||
provider: Provider family (mistral | openai | bedrock | ollama | simple).
|
||||
operation: ``"index"`` (chunk-batch embedding) or ``"query"`` (search
|
||||
query embedding).
|
||||
tokens: Token count for this embedding request (no-op when ``<= 0``).
|
||||
"""
|
||||
if tokens > 0:
|
||||
embedding_tokens_total.labels(provider=provider, operation=operation).inc(
|
||||
tokens
|
||||
)
|
||||
|
||||
|
||||
def record_document_chunks(doc_type: str, count: int) -> None:
|
||||
"""
|
||||
Record the number of chunks produced for a document.
|
||||
|
||||
Reference in New Issue
Block a user