feat(usage): rename metrics → tokens_embedded/pages_embedded + export token cost to Prometheus

Billing product model finalized (Deck #281): bill pages externally, record
tokens internally. Rename the data-plane metric literals to match the now-
canonical contract (Deck #284) — the control plane's METRIC_EVENT_NAMES is
already renamed, so the old names would be unmapped and never sync to Stripe.

Rename (values unchanged):
- embeddings_queries → tokens_embedded (value = real token count, already
  emitted by this PR; the unit upstream providers bill on).
- pages_chunks → pages_embedded (value kept as len(chunk_texts) interim;
  TODO(#282): real normalized "pages indexed" count — real pages for paginated
  types, chars/tokens-per-page constant otherwise — is deferred to the
  instrumentation card, this only lands the name/contract).
- All literals, log strings, docstrings, comments, the migration comment, and
  tests renamed; grep confirms zero old strings remain.

Observability (new): export embedding token cost to Prometheus as
astrolabe_embedding_tokens_total{provider,operation} (operation = index|query)
so the billed cost unit is visible in Grafana, not just the per-tenant billing
DB. Dedicated counter (doesn't inflate the existing chunk/request metrics) and
always-on (independent of USAGE_METERING_ENABLED, so OSS/self-host gets it).
Wired on both the indexing batch embed and the search query embed (query inside
the per-request cache-miss branch, so reused embeddings aren't double-counted).

Note: the rename orphans any pre-existing embeddings_queries/pages_chunks rows
in tenant app DBs (CP no longer maps them) — acceptable; pipeline is inert with
throwaway dev/sandbox data.

Deck #284 (folded into PR #875).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chris Coutinho
2026-06-08 13:17:53 +02:00
co-authored by Claude Opus 4.8
parent ddefb03701
commit 973f80e7b9
15 changed files with 129 additions and 45 deletions
+32 -1
View File
@@ -12,7 +12,10 @@ from __future__ import annotations
import pytest
from nextcloud_mcp_server.config import Settings
from nextcloud_mcp_server.observability.metrics import record_embedding
from nextcloud_mcp_server.observability.metrics import (
record_embedding,
record_embedding_tokens,
)
pytestmark = pytest.mark.unit
@@ -120,3 +123,31 @@ class TestRecordEmbedding:
assert metric_sample(
"astrolabe_embedding_requests_total", {**labels, "status": "error"}
) == pytest.approx(1.0)
class TestRecordEmbeddingTokens:
"""astrolabe_embedding_tokens_total — token cost split by index/query."""
def test_index_increments_by_token_count(self, metric_sample):
labels = {"provider": "tok-prov", "operation": "index"}
before = metric_sample("astrolabe_embedding_tokens_total", labels)
record_embedding_tokens("tok-prov", "index", 4242)
assert metric_sample(
"astrolabe_embedding_tokens_total", labels
) == pytest.approx(before + 4242)
def test_query_operation_is_separate_series(self, metric_sample):
labels = {"provider": "tok-prov", "operation": "query"}
before = metric_sample("astrolabe_embedding_tokens_total", labels)
record_embedding_tokens("tok-prov", "query", 7)
assert metric_sample(
"astrolabe_embedding_tokens_total", labels
) == pytest.approx(before + 7)
def test_zero_or_negative_is_noop(self, metric_sample):
labels = {"provider": "tok-noop", "operation": "index"}
record_embedding_tokens("tok-noop", "index", 0)
record_embedding_tokens("tok-noop", "index", -3)
assert metric_sample(
"astrolabe_embedding_tokens_total", labels
) == pytest.approx(0.0)