feat(usage): rename metrics → tokens_embedded/pages_embedded + export token cost to Prometheus

Billing product model finalized (Deck #281): bill pages externally, record
tokens internally. Rename the data-plane metric literals to match the now-
canonical contract (Deck #284) — the control plane's METRIC_EVENT_NAMES is
already renamed, so the old names would be unmapped and never sync to Stripe.

Rename (values unchanged):
- embeddings_queries → tokens_embedded (value = real token count, already
  emitted by this PR; the unit upstream providers bill on).
- pages_chunks → pages_embedded (value kept as len(chunk_texts) interim;
  TODO(#282): real normalized "pages indexed" count — real pages for paginated
  types, chars/tokens-per-page constant otherwise — is deferred to the
  instrumentation card, this only lands the name/contract).
- All literals, log strings, docstrings, comments, the migration comment, and
  tests renamed; grep confirms zero old strings remain.

Observability (new): export embedding token cost to Prometheus as
astrolabe_embedding_tokens_total{provider,operation} (operation = index|query)
so the billed cost unit is visible in Grafana, not just the per-tenant billing
DB. Dedicated counter (doesn't inflate the existing chunk/request metrics) and
always-on (independent of USAGE_METERING_ENABLED, so OSS/self-host gets it).
Wired on both the indexing batch embed and the search query embed (query inside
the per-request cache-miss branch, so reused embeddings aren't double-counted).

Note: the rename orphans any pre-existing embeddings_queries/pages_chunks rows
in tenant app DBs (CP no longer maps them) — acceptable; pipeline is inert with
throwaway dev/sandbox data.

Deck #284 (folded into PR #875).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-08 13:17:53 +02:00
co-authored by Claude Opus 4.8
parent ddefb03701
commit 973f80e7b9
15 changed files with 129 additions and 45 deletions
@@ -338,6 +338,16 @@ embedding_chars_total = Counter(
["kind", "provider"],
)
# Token consumption — the billed cost unit (mirrors the tokens_embedded billing
# measure, Deck #67). On a dedicated counter (not folded into the chunk/request
# metrics above) so query embeds don't inflate indexing dashboards; labelled by
# operation = index | query. Always emitted, independent of USAGE_METERING_ENABLED.
embedding_tokens_total = Counter(
"astrolabe_embedding_tokens_total",
"Total embedding tokens consumed (provider-reported or estimated)",
["provider", "operation"], # operation: index | query
)
# --- Chunking & indexed-by-type -----------------------------------------------
document_chunks_total = Counter(
@@ -727,6 +737,25 @@ def record_embedding(
embedding_chars_total.labels(kind=kind, provider=provider).inc(chars)
def record_embedding_tokens(provider: str, operation: str, tokens: int) -> None:
"""Export embedding token consumption to Prometheus.
Mirrors the ``tokens_embedded`` billing measure (Deck #67) as an always-on
observability signal — emitted regardless of ``USAGE_METERING_ENABLED`` so
OSS/self-host deployments still see token cost in Grafana.
Args:
provider: Provider family (mistral | openai | bedrock | ollama | simple).
operation: ``"index"`` (chunk-batch embedding) or ``"query"`` (search
query embedding).
tokens: Token count for this embedding request (no-op when ``<= 0``).
"""
if tokens > 0:
embedding_tokens_total.labels(provider=provider, operation=operation).inc(
tokens
)
def record_document_chunks(doc_type: str, count: int) -> None:
"""
Record the number of chunks produced for a document.