From 5dca9dccf17527f2bd260aaa2726d3b568ddeea2 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Mon, 28 Sep 2026 08:57:14 -0700 Subject: [PATCH 01/42] Add the MetricMapping and TelemetryDestination kinds Modelplane collects from everything it installs and normalizes it onto one modelplane_* surface, and two things about that are the operator's: where the telemetry goes, and what to do about a component Modelplane ships no rules for. These are the two kinds the telemetry design gives them. Both hold OpenTelemetry collector configuration that Modelplane passes through unread. A MetricMapping carries OTTL statements rendered into every inference cluster's transform processor, and passthrough for keeping a component's own metric names flowing during a migration. A TelemetryDestination carries the collector's exporters and extensions, a secretRef so credentials stay out of the object, and collector: External for a platform that already runs a fleet collector and hands over its endpoint. Passing configuration through unread is the point rather than an omission. An operator writes the collector's own language, documented upstream, so any exporter or authenticator it gains works without a Modelplane release, and a field-by-field schema would either restate all of it or quietly cap what a destination can be. Signed-off-by: Dennis Ramdass --- apis/metricmappings/composition.yaml | 13 +++ apis/metricmappings/definition.yaml | 83 ++++++++++++++++++ apis/telemetrydestinations/composition.yaml | 13 +++ apis/telemetrydestinations/definition.yaml | 97 +++++++++++++++++++++ crossplane-project.yaml | 8 ++ flake.nix | 2 + uv.lock | 40 +++++++++ 7 files changed, 256 insertions(+) create mode 100644 apis/metricmappings/composition.yaml create mode 100644 apis/metricmappings/definition.yaml create mode 100644 apis/telemetrydestinations/composition.yaml create mode 100644 apis/telemetrydestinations/definition.yaml diff --git a/apis/metricmappings/composition.yaml b/apis/metricmappings/composition.yaml new file mode 100644 index 000000000..8359dc414 --- /dev/null +++ b/apis/metricmappings/composition.yaml @@ -0,0 +1,13 @@ +apiVersion: apiextensions.crossplane.io/v1 +kind: Composition +metadata: + name: metricmappings.modelplane.ai +spec: + compositeTypeRef: + apiVersion: modelplane.ai/v1alpha1 + kind: MetricMapping + mode: Pipeline + pipeline: + - functionRef: + name: modelplane-modelplanecompose-metric-mapping + step: compose-metric-mapping diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml new file mode 100644 index 000000000..def56a04b --- /dev/null +++ b/apis/metricmappings/definition.yaml @@ -0,0 +1,83 @@ +apiVersion: apiextensions.crossplane.io/v2 +kind: CompositeResourceDefinition +metadata: + name: metricmappings.modelplane.ai +spec: + group: modelplane.ai + names: + categories: [crossplane, modelplane] + kind: MetricMapping + plural: metricmappings + shortNames: [mm] + scope: Cluster + versions: + - name: v1alpha1 + served: true + referenceable: true + additionalPrinterColumns: + - name: CLUSTERS + type: integer + jsonPath: .status.clusters + - name: AGE + type: date + jsonPath: .metadata.creationTimestamp + schema: + openAPIV3Schema: + type: object + required: [spec] + properties: + spec: + type: object + description: >- + How one component's metrics become part of the modelplane_* + surface. Modelplane renders every MetricMapping into each + inference cluster's collector, so a mapping is written once on + the control plane and reaches the whole fleet. + + A mapping naming a component Modelplane already provides + statements for is additive: its statements run after the + built-in ones and its flags apply. That is how passthrough is + turned on for an engine that needs no statements of its own. + properties: + statements: + type: array + description: >- + OTTL statements, rendered into the collector's transform + processor beside Modelplane's own. Modelplane does not + interpret them: what you write here is the collector's own + configuration language, documented by OpenTelemetry, and it + is the same thing Modelplane writes for vLLM. + + Statements select through their own where clauses, so + nothing declares which engine a deployment runs. + items: + type: string + maxLength: 2048 + maxItems: 128 + passthrough: + type: boolean + default: false + description: >- + Send this component's own metric names onward as well as the + modelplane_* ones they become. + + Off by default, because a series the statements did not + rename is one whose meaning Modelplane cannot vouch for + across engines, and it costs the same to carry as one that + was renamed. On, for reading an engine's raw names during a + migration or while debugging that engine. + + A passed-through series is still merged across a + deployment's replicas, so it keeps the labels the engine + gave it and carries no pod identity. + status: + type: object + properties: + clusters: + type: integer + description: >- + How many inference clusters have taken these statements. + conditions: + type: array + items: + type: object diff --git a/apis/telemetrydestinations/composition.yaml b/apis/telemetrydestinations/composition.yaml new file mode 100644 index 000000000..b345a3d7b --- /dev/null +++ b/apis/telemetrydestinations/composition.yaml @@ -0,0 +1,13 @@ +apiVersion: apiextensions.crossplane.io/v1 +kind: Composition +metadata: + name: telemetrydestinations.modelplane.ai +spec: + compositeTypeRef: + apiVersion: modelplane.ai/v1alpha1 + kind: TelemetryDestination + mode: Pipeline + pipeline: + - functionRef: + name: modelplane-modelplanecompose-telemetry-destination + step: compose-telemetry-destination diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml new file mode 100644 index 000000000..4c40dbe33 --- /dev/null +++ b/apis/telemetrydestinations/definition.yaml @@ -0,0 +1,97 @@ +apiVersion: apiextensions.crossplane.io/v2 +kind: CompositeResourceDefinition +metadata: + name: telemetrydestinations.modelplane.ai +spec: + group: modelplane.ai + names: + categories: [crossplane, modelplane] + kind: TelemetryDestination + plural: telemetrydestinations + shortNames: [td] + scope: Cluster + versions: + - name: v1alpha1 + served: true + referenceable: true + additionalPrinterColumns: + - name: COLLECTOR + type: string + jsonPath: .spec.collector + - name: AGE + type: date + jsonPath: .metadata.creationTimestamp + schema: + openAPIV3Schema: + type: object + required: [spec] + properties: + spec: + type: object + required: [exporters] + description: >- + Where the fleet's telemetry goes. Modelplane composes no + collectors until a TelemetryDestination exists: neither tier + stores anything, so collecting with nowhere to export would + spend GPU-cluster memory on samples nobody reads. Creating one + turns collection on everywhere at once. + + There is no per-deployment opt-out. A ModelDeployment's author + owns neither the destination nor its bill. + properties: + collector: + type: string + default: Composed + enum: [Composed, External] + description: >- + Whether Modelplane runs the fleet collector. Composed (the + default) puts one on the control plane, and every inference + cluster exports to it. External composes none, for a + platform that already operates one: each inference cluster + then exports to the endpoint below directly. + + External gives up the single egress point, one place to + change the destination, and the control plane's own series + reaching the fleet without a path of their own. Whoever + imposed the endpoint has usually provided them already. + exporters: + type: object + x-kubernetes-preserve-unknown-fields: true + description: >- + The OpenTelemetry collector's exporters block, passed + through unread. Modelplane validates that it parses and + reports whether the destination accepts writes; it does not + model what an exporter is. + + So any exporter the collector provides works, with its TLS, + retry and queue settings intact, and a destination keeps + working when the collector gains an exporter Modelplane has + never heard of. + extensions: + type: object + x-kubernetes-preserve-unknown-fields: true + description: >- + The collector's extensions block, passed through unread, + for the authenticator an exporter references. Bearer token, + basic auth, OIDC and SigV4 all work, because none of them + is modelled here. + secretRef: + type: object + required: [name] + description: >- + A Secret whose keys Modelplane mounts into the collector as + environment variables, so configuration above refers to + ${env:TOKEN} and the credential itself never appears in this + object or in kubectl output. + properties: + name: + type: string + maxLength: 253 + description: Name of the Secret, in Modelplane's namespace. + status: + type: object + properties: + conditions: + type: array + items: + type: object diff --git a/crossplane-project.yaml b/crossplane-project.yaml index b4eb40d39..532b15c58 100644 --- a/crossplane-project.yaml +++ b/crossplane-project.yaml @@ -75,6 +75,14 @@ spec: tarball: name: compose-model-service pathPrefix: _output/functions/compose-model-service + - source: Tarball + tarball: + name: compose-metric-mapping + pathPrefix: _output/functions/compose-metric-mapping + - source: Tarball + tarball: + name: compose-telemetry-destination + pathPrefix: _output/functions/compose-telemetry-destination - source: Tarball tarball: name: compose-usages diff --git a/flake.nix b/flake.nix index 56c8b85ab..02704e802 100644 --- a/flake.nix +++ b/flake.nix @@ -57,6 +57,8 @@ "compose-inference-class" "compose-inference-cluster" "compose-inference-gateway" + "compose-metric-mapping" + "compose-telemetry-destination" "compose-nebius-cluster" "compose-serving-stack" "compose-vultr-cluster" diff --git a/uv.lock b/uv.lock index b72267a7c..9fcb8dba1 100644 --- a/uv.lock +++ b/uv.lock @@ -15,6 +15,7 @@ members = [ "compose-inference-class", "compose-inference-cluster", "compose-inference-gateway", + "compose-metric-mapping", "compose-model-cache", "compose-model-deployment", "compose-model-endpoint", @@ -23,6 +24,7 @@ members = [ "compose-model-service", "compose-nebius-cluster", "compose-serving-stack", + "compose-telemetry-destination", "compose-usages", "compose-vultr-cluster", "crossplane-models", @@ -208,6 +210,25 @@ requires-dist = [ { name = "grpcio", specifier = ">=1.73.1" }, ] +[[package]] +name = "compose-metric-mapping" +version = "0.0.0" +source = { editable = "functions/compose-metric-mapping" } +dependencies = [ + { name = "click" }, + { name = "crossplane-function-sdk-python" }, + { name = "crossplane-models" }, + { name = "grpcio" }, +] + +[package.metadata] +requires-dist = [ + { name = "click", specifier = ">=8.1.0" }, + { name = "crossplane-function-sdk-python", specifier = ">=0.14.0" }, + { name = "crossplane-models", editable = "schemas/python" }, + { name = "grpcio", specifier = ">=1.73.1" }, +] + [[package]] name = "compose-model-cache" version = "0.0.0" @@ -364,6 +385,25 @@ requires-dist = [ { name = "pyyaml", specifier = ">=6.0" }, ] +[[package]] +name = "compose-telemetry-destination" +version = "0.0.0" +source = { editable = "functions/compose-telemetry-destination" } +dependencies = [ + { name = "click" }, + { name = "crossplane-function-sdk-python" }, + { name = "crossplane-models" }, + { name = "grpcio" }, +] + +[package.metadata] +requires-dist = [ + { name = "click", specifier = ">=8.1.0" }, + { name = "crossplane-function-sdk-python", specifier = ">=0.14.0" }, + { name = "crossplane-models", editable = "schemas/python" }, + { name = "grpcio", specifier = ">=1.73.1" }, +] + [[package]] name = "compose-usages" version = "0.0.0" From f13387d8d3d28eee60492329bee51287201a4899 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Mon, 28 Sep 2026 08:57:14 -0700 Subject: [PATCH 02/42] Report whether a mapping and a destination actually work Neither kind composes anything: compose-serving-stack is what renders them into a collector. So both functions could mark their XR Ready and look no further, which for a data resource is nearly honest. It is also how a fleet ends up producing no telemetry with every object green. A MetricMapping reaches every inference cluster, so the function resolves them and reports the count, and is not Ready when there are none. It also catches a mapping carrying neither statements nor passthrough, which changes nothing and otherwise looks exactly like one that works. A TelemetryDestination gets the one check Modelplane can make without modelling what an exporter is. An exporter's auth block names an authenticator by extension name, and the collector refuses to start when no extension defines it, so a typo there stops telemetry with the failure two layers from its cause. The same for a credential Secret that does not exist, and for a destination carrying no exporters at all. Neither reads further in. Validating the statements or the exporters would be Modelplane modelling the collector's configuration, which is what these kinds exist to avoid. Signed-off-by: Dennis Ramdass --- .../function/__init__.py | 13 + .../compose-metric-mapping/function/fn.py | 106 ++++++++ .../compose-metric-mapping/function/main.py | 55 +++++ .../compose-metric-mapping/pyproject.toml | 26 ++ .../compose-metric-mapping/tests/__init__.py | 13 + .../compose-metric-mapping/tests/test_fn.py | 169 +++++++++++++ .../function/__init__.py | 13 + .../function/fn.py | 128 ++++++++++ .../function/main.py | 55 +++++ .../pyproject.toml | 26 ++ .../tests/__init__.py | 13 + .../tests/test_fn.py | 231 ++++++++++++++++++ 12 files changed, 848 insertions(+) create mode 100644 functions/compose-metric-mapping/function/__init__.py create mode 100644 functions/compose-metric-mapping/function/fn.py create mode 100644 functions/compose-metric-mapping/function/main.py create mode 100644 functions/compose-metric-mapping/pyproject.toml create mode 100644 functions/compose-metric-mapping/tests/__init__.py create mode 100644 functions/compose-metric-mapping/tests/test_fn.py create mode 100644 functions/compose-telemetry-destination/function/__init__.py create mode 100644 functions/compose-telemetry-destination/function/fn.py create mode 100644 functions/compose-telemetry-destination/function/main.py create mode 100644 functions/compose-telemetry-destination/pyproject.toml create mode 100644 functions/compose-telemetry-destination/tests/__init__.py create mode 100644 functions/compose-telemetry-destination/tests/test_fn.py diff --git a/functions/compose-metric-mapping/function/__init__.py b/functions/compose-metric-mapping/function/__init__.py new file mode 100644 index 000000000..ebf4b2ad4 --- /dev/null +++ b/functions/compose-metric-mapping/function/__init__.py @@ -0,0 +1,13 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. diff --git a/functions/compose-metric-mapping/function/fn.py b/functions/compose-metric-mapping/function/fn.py new file mode 100644 index 000000000..5c0d9468d --- /dev/null +++ b/functions/compose-metric-mapping/function/fn.py @@ -0,0 +1,106 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Compose a MetricMapping. + +A MetricMapping carries OTTL statements that the collector on every +inference cluster renders into its transform processor. Modelplane does not +interpret them: what an operator writes here is the collector's own +configuration language. + +This function composes nothing. The collector is composed by +compose-serving-stack, which reads every MetricMapping. What this function +does is tell an operator whether the mapping reaches anything, because a +mapping that reaches no cluster looks identical to one that works. +""" + +import grpc +from crossplane.function import logging, request, resource, response +from crossplane.function.proto.v1 import run_function_pb2 as fnv1 +from crossplane.function.proto.v1 import run_function_pb2_grpc as grpcv1 +from models.ai.modelplane.metricmapping import v1alpha1 + +CONDITION_TYPE_ACCEPTED = "Accepted" +CONDITION_REASON_AVAILABLE = "Available" +CONDITION_REASON_WAITING_FOR_CLUSTERS = "WaitingForClusters" +CONDITION_REASON_NO_CLUSTERS = "NoClusters" +CONDITION_REASON_NO_STATEMENTS = "NoStatements" + + +class FunctionRunner(grpcv1.FunctionRunnerServiceServicer): + """A FunctionRunner handles gRPC RunFunctionRequests.""" + + def __init__(self) -> None: + """Create a new FunctionRunner.""" + self.log = logging.get_logger() + + async def RunFunction( + self, req: fnv1.RunFunctionRequest, _: grpc.aio.ServicerContext | None + ) -> fnv1.RunFunctionResponse: # ty: ignore[invalid-method-override] # the generated grpc servicer base is untyped + """Run the function.""" + log = self.log.bind(tag=req.meta.tag) + log.info("Running function") + + rsp = response.to(req) + xr = v1alpha1.MetricMapping(**resource.struct_to_dict(req.observed.composite.resource)) + + # Every inference cluster renders every mapping, so the count of + # clusters is the count that took these statements. + response.require_resources( + rsp, + name="clusters", + api_version="modelplane.ai/v1alpha1", + kind="InferenceCluster", + ) + if "clusters" not in req.required_resources: + _not_ready(rsp, CONDITION_REASON_WAITING_FOR_CLUSTERS, "Waiting for the inference clusters to resolve") + return rsp + + clusters = len(list(request.get_required_resources(req, "clusters"))) + resource.update_status(rsp.desired.composite, v1alpha1.Status(clusters=clusters)) + + # A mapping carrying neither statements nor passthrough does nothing at + # all, which is worth saying rather than reporting ready. + if not xr.spec.statements and not xr.spec.passthrough: + _not_ready( + rsp, + CONDITION_REASON_NO_STATEMENTS, + "No statements and passthrough is off, so this mapping changes nothing", + ) + return rsp + + if clusters == 0: + _not_ready(rsp, CONDITION_REASON_NO_CLUSTERS, "No inference cluster to render these statements into") + return rsp + + response.set_conditions( + rsp, + resource.Condition( + typ=CONDITION_TYPE_ACCEPTED, + status="True", + reason=CONDITION_REASON_AVAILABLE, + message=f"Rendered into {clusters} inference cluster(s)", + ), + ) + rsp.desired.composite.ready = fnv1.READY_TRUE + return rsp + + +def _not_ready(rsp: fnv1.RunFunctionResponse, reason: str, message: str) -> None: + """Report a mapping that isn't reaching anything, and why.""" + response.set_conditions( + rsp, + resource.Condition(typ=CONDITION_TYPE_ACCEPTED, status="False", reason=reason, message=message), + ) + rsp.desired.composite.ready = fnv1.READY_FALSE diff --git a/functions/compose-metric-mapping/function/main.py b/functions/compose-metric-mapping/function/main.py new file mode 100644 index 000000000..2e8441dac --- /dev/null +++ b/functions/compose-metric-mapping/function/main.py @@ -0,0 +1,55 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""The composition function's main CLI.""" + +import click +from crossplane.function import logging, runtime + +from function import fn + + +@click.command() +@click.option("--debug", "-d", is_flag=True, help="Emit debug logs.") +@click.option( + "--address", + default="0.0.0.0:9443", + show_default=True, + help="Address at which to listen for gRPC connections", +) +@click.option("--tls-certs-dir", help="Serve using mTLS certificates.", envvar="TLS_SERVER_CERTS_DIR") +@click.option( + "--insecure", + is_flag=True, + help="Run without mTLS credentials. If you supply this flag --tls-certs-dir will be ignored.", +) +def cli(debug: bool, address: str, tls_certs_dir: str, insecure: bool) -> None: + """A Crossplane composition function.""" + try: + level = logging.Level.INFO + if debug: + level = logging.Level.DEBUG + logging.configure(level=level) + runtime.serve( + fn.FunctionRunner(), + address, + creds=runtime.load_credentials(tls_certs_dir), + insecure=insecure, + ) + except Exception as e: + click.echo(f"Cannot run function: {e}") + + +if __name__ == "__main__": + cli() diff --git a/functions/compose-metric-mapping/pyproject.toml b/functions/compose-metric-mapping/pyproject.toml new file mode 100644 index 000000000..526492339 --- /dev/null +++ b/functions/compose-metric-mapping/pyproject.toml @@ -0,0 +1,26 @@ +[build-system] +requires = ["uv_build>=0.11.0,<0.12"] +build-backend = "uv_build" + +[project] +name = "compose-metric-mapping" +version = "0.0.0" +description = "Mark a MetricMapping as ready." +requires-python = ">=3.11,<3.14" +license = "Apache-2.0" +dependencies = [ + "crossplane-function-sdk-python>=0.14.0", + "click>=8.1.0", + "grpcio>=1.73.1", + "crossplane-models", +] + +[tool.uv.sources] +crossplane-models = { workspace = true } + +[project.scripts] +function = "function.main:cli" + +[tool.uv.build-backend] +module-name = "function" +module-root = "" diff --git a/functions/compose-metric-mapping/tests/__init__.py b/functions/compose-metric-mapping/tests/__init__.py new file mode 100644 index 000000000..ebf4b2ad4 --- /dev/null +++ b/functions/compose-metric-mapping/tests/__init__.py @@ -0,0 +1,13 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. diff --git a/functions/compose-metric-mapping/tests/test_fn.py b/functions/compose-metric-mapping/tests/test_fn.py new file mode 100644 index 000000000..76209d80b --- /dev/null +++ b/functions/compose-metric-mapping/tests/test_fn.py @@ -0,0 +1,169 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Tests for the compose-metric-mapping function.""" + +import dataclasses +import unittest + +from crossplane.function import logging, resource +from crossplane.function.proto.v1 import run_function_pb2 as fnv1 +from function import fn +from google.protobuf import duration_pb2 as durationpb +from google.protobuf import json_format +from google.protobuf import struct_pb2 as structpb + + +@dataclasses.dataclass +class Case: + """A test case for compose-metric-mapping.""" + + name: str + req: fnv1.RunFunctionRequest + want: fnv1.RunFunctionResponse + + +def setUpModule() -> None: + logging.configure(level=logging.Level.DISABLED) + + +class TestFunctionRunner(unittest.IsolatedAsyncioTestCase): + """Tests for FunctionRunner.RunFunction.""" + + @classmethod + def setUpClass(cls) -> None: + cls.runner = fn.FunctionRunner() + + async def test_compose(self) -> None: + """The function reports whether a mapping reaches any cluster.""" + mapping = { + "apiVersion": "modelplane.ai/v1alpha1", + "kind": "MetricMapping", + "metadata": {"name": "my-engine"}, + "spec": { + "statements": ['set(name, "modelplane_requests_waiting") where name == "my_engine_queued"'], + }, + } + cluster = resource.dict_to_struct( + {"apiVersion": "modelplane.ai/v1alpha1", "kind": "InferenceCluster", "metadata": {"name": "prod-us-east"}} + ) + + def req(xr: dict, clusters: list | None) -> fnv1.RunFunctionRequest: + r = fnv1.RunFunctionRequest( + observed=fnv1.State(composite=fnv1.Resource(resource=resource.dict_to_struct(xr))), + ) + if clusters is not None: + r.required_resources["clusters"].items.extend([fnv1.Resource(resource=c) for c in clusters]) + return r + + def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.RunFunctionResponse: + composite = fnv1.Resource(ready=ready) + if status is not None: + composite.resource.CopyFrom(resource.dict_to_struct(status)) + return fnv1.RunFunctionResponse( + meta=fnv1.ResponseMeta(ttl=durationpb.Duration(seconds=60)), + desired=fnv1.State(composite=composite), + conditions=[cond], + context=structpb.Struct(), + requirements=fnv1.Requirements( + resources={ + "clusters": fnv1.ResourceSelector(api_version="modelplane.ai/v1alpha1", kind="InferenceCluster") + } + ), + ) + + no_statements = {**mapping, "spec": {}} + passthrough_only = {**mapping, "spec": {"passthrough": True}} + + cases = [ + Case( + name="ready, and says how many clusters took the statements", + req=req(mapping, [cluster, cluster]), + want=want( + fnv1.READY_TRUE, + {"status": {"clusters": 2}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Rendered into 2 inference cluster(s)", + ), + ), + ), + Case( + name="ready on passthrough alone, which is a mapping with no statements to write", + req=req(passthrough_only, [cluster]), + want=want( + fnv1.READY_TRUE, + {"status": {"clusters": 1}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Rendered into 1 inference cluster(s)", + ), + ), + ), + Case( + name="not ready when no cluster exists to render into", + req=req(mapping, []), + want=want( + fnv1.READY_FALSE, + {"status": {"clusters": 0}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="NoClusters", + message="No inference cluster to render these statements into", + ), + ), + ), + Case( + name="not ready when the mapping would change nothing", + req=req(no_statements, [cluster]), + want=want( + fnv1.READY_FALSE, + {"status": {"clusters": 1}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="NoStatements", + message="No statements and passthrough is off, so this mapping changes nothing", + ), + ), + ), + Case( + name="waits for the clusters to resolve", + req=req(mapping, None), + want=want( + fnv1.READY_FALSE, + None, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="WaitingForClusters", + message="Waiting for the inference clusters to resolve", + ), + ), + ), + ] + + for case in cases: + with self.subTest(case.name): + got = await self.runner.RunFunction(case.req, None) + self.assertEqual( + json_format.MessageToDict(case.want), + json_format.MessageToDict(got), + "-want, +got", + ) diff --git a/functions/compose-telemetry-destination/function/__init__.py b/functions/compose-telemetry-destination/function/__init__.py new file mode 100644 index 000000000..ebf4b2ad4 --- /dev/null +++ b/functions/compose-telemetry-destination/function/__init__.py @@ -0,0 +1,13 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. diff --git a/functions/compose-telemetry-destination/function/fn.py b/functions/compose-telemetry-destination/function/fn.py new file mode 100644 index 000000000..309a7b912 --- /dev/null +++ b/functions/compose-telemetry-destination/function/fn.py @@ -0,0 +1,128 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Compose a TelemetryDestination. + +A TelemetryDestination carries the collector's exporters and extensions +verbatim, and compose-serving-stack renders them into the collector it +composes. Modelplane does not model what an exporter is, so there is little +here to validate and the little there is matters: an exporter naming an +authenticator that no extension defines makes a collector refuse to start, +and that failure surfaces as telemetry silently never arriving. +""" + +import grpc +from crossplane.function import logging, request, resource, response +from crossplane.function.proto.v1 import run_function_pb2 as fnv1 +from crossplane.function.proto.v1 import run_function_pb2_grpc as grpcv1 +from models.ai.modelplane.telemetrydestination import v1alpha1 + +CONDITION_TYPE_ACCEPTED = "Accepted" +CONDITION_REASON_AVAILABLE = "Available" +CONDITION_REASON_NO_EXPORTERS = "NoExporters" +CONDITION_REASON_UNKNOWN_AUTHENTICATOR = "UnknownAuthenticator" +CONDITION_REASON_WAITING_FOR_SECRET = "WaitingForSecret" +CONDITION_REASON_SECRET_NOT_FOUND = "SecretNotFound" + +_SECRET_KEY = "secret" + + +class FunctionRunner(grpcv1.FunctionRunnerServiceServicer): + """A FunctionRunner handles gRPC RunFunctionRequests.""" + + def __init__(self) -> None: + """Create a new FunctionRunner.""" + self.log = logging.get_logger() + + async def RunFunction( + self, req: fnv1.RunFunctionRequest, _: grpc.aio.ServicerContext | None + ) -> fnv1.RunFunctionResponse: # ty: ignore[invalid-method-override] # the generated grpc servicer base is untyped + """Run the function.""" + log = self.log.bind(tag=req.meta.tag) + log.info("Running function") + + rsp = response.to(req) + xr = v1alpha1.TelemetryDestination(**resource.struct_to_dict(req.observed.composite.resource)) + + exporters = xr.spec.exporters or {} + if not exporters: + _not_ready(rsp, CONDITION_REASON_NO_EXPORTERS, "No exporters, so collected telemetry has nowhere to go") + return rsp + + # An exporter's auth block names an authenticator by extension name. The + # collector refuses to start when it names one no extension defines, and + # a collector that never starts looks exactly like a fleet that produces + # nothing, so it is worth catching on the object instead. + extensions = set((xr.spec.extensions or {}).keys()) + missing = sorted(_authenticators(exporters) - extensions) + if missing: + _not_ready( + rsp, + CONDITION_REASON_UNKNOWN_AUTHENTICATOR, + f"No extension defines {', '.join(missing)}, so the collector would refuse to start", + ) + return rsp + + if xr.spec.secretRef is not None: + response.require_resources( + rsp, + name=_SECRET_KEY, + api_version="v1", + kind="Secret", + match_name=xr.spec.secretRef.name, + ) + if _SECRET_KEY not in req.required_resources: + _not_ready(rsp, CONDITION_REASON_WAITING_FOR_SECRET, "Waiting for the credential Secret to resolve") + return rsp + if not list(request.get_required_resources(req, _SECRET_KEY)): + _not_ready( + rsp, + CONDITION_REASON_SECRET_NOT_FOUND, + f"Secret {xr.spec.secretRef.name} does not exist, so the collector has no credential to send with", + ) + return rsp + + resource.update_status(rsp.desired.composite, v1alpha1.Status()) + response.set_conditions( + rsp, + resource.Condition( + typ=CONDITION_TYPE_ACCEPTED, + status="True", + reason=CONDITION_REASON_AVAILABLE, + message=f"Exporting through {', '.join(sorted(exporters))}", + ), + ) + rsp.desired.composite.ready = fnv1.READY_TRUE + return rsp + + +def _authenticators(exporters: dict) -> set[str]: + """Every authenticator an exporter references, by extension name.""" + names: set[str] = set() + for cfg in exporters.values(): + if not isinstance(cfg, dict): + continue + auth = cfg.get("auth") + if isinstance(auth, dict) and isinstance(auth.get("authenticator"), str): + names.add(auth["authenticator"]) + return names + + +def _not_ready(rsp: fnv1.RunFunctionResponse, reason: str, message: str) -> None: + """Report a destination nothing can send through, and why.""" + response.set_conditions( + rsp, + resource.Condition(typ=CONDITION_TYPE_ACCEPTED, status="False", reason=reason, message=message), + ) + rsp.desired.composite.ready = fnv1.READY_FALSE diff --git a/functions/compose-telemetry-destination/function/main.py b/functions/compose-telemetry-destination/function/main.py new file mode 100644 index 000000000..2e8441dac --- /dev/null +++ b/functions/compose-telemetry-destination/function/main.py @@ -0,0 +1,55 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""The composition function's main CLI.""" + +import click +from crossplane.function import logging, runtime + +from function import fn + + +@click.command() +@click.option("--debug", "-d", is_flag=True, help="Emit debug logs.") +@click.option( + "--address", + default="0.0.0.0:9443", + show_default=True, + help="Address at which to listen for gRPC connections", +) +@click.option("--tls-certs-dir", help="Serve using mTLS certificates.", envvar="TLS_SERVER_CERTS_DIR") +@click.option( + "--insecure", + is_flag=True, + help="Run without mTLS credentials. If you supply this flag --tls-certs-dir will be ignored.", +) +def cli(debug: bool, address: str, tls_certs_dir: str, insecure: bool) -> None: + """A Crossplane composition function.""" + try: + level = logging.Level.INFO + if debug: + level = logging.Level.DEBUG + logging.configure(level=level) + runtime.serve( + fn.FunctionRunner(), + address, + creds=runtime.load_credentials(tls_certs_dir), + insecure=insecure, + ) + except Exception as e: + click.echo(f"Cannot run function: {e}") + + +if __name__ == "__main__": + cli() diff --git a/functions/compose-telemetry-destination/pyproject.toml b/functions/compose-telemetry-destination/pyproject.toml new file mode 100644 index 000000000..d36739309 --- /dev/null +++ b/functions/compose-telemetry-destination/pyproject.toml @@ -0,0 +1,26 @@ +[build-system] +requires = ["uv_build>=0.11.0,<0.12"] +build-backend = "uv_build" + +[project] +name = "compose-telemetry-destination" +version = "0.0.0" +description = "Mark a TelemetryDestination as ready." +requires-python = ">=3.11,<3.14" +license = "Apache-2.0" +dependencies = [ + "crossplane-function-sdk-python>=0.14.0", + "click>=8.1.0", + "grpcio>=1.73.1", + "crossplane-models", +] + +[tool.uv.sources] +crossplane-models = { workspace = true } + +[project.scripts] +function = "function.main:cli" + +[tool.uv.build-backend] +module-name = "function" +module-root = "" diff --git a/functions/compose-telemetry-destination/tests/__init__.py b/functions/compose-telemetry-destination/tests/__init__.py new file mode 100644 index 000000000..ebf4b2ad4 --- /dev/null +++ b/functions/compose-telemetry-destination/tests/__init__.py @@ -0,0 +1,13 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. diff --git a/functions/compose-telemetry-destination/tests/test_fn.py b/functions/compose-telemetry-destination/tests/test_fn.py new file mode 100644 index 000000000..d5710ba63 --- /dev/null +++ b/functions/compose-telemetry-destination/tests/test_fn.py @@ -0,0 +1,231 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Tests for the compose-metric-mapping function.""" + +import dataclasses +import unittest + +from crossplane.function import logging, resource +from crossplane.function.proto.v1 import run_function_pb2 as fnv1 +from function import fn +from google.protobuf import duration_pb2 as durationpb +from google.protobuf import json_format +from google.protobuf import struct_pb2 as structpb + + +@dataclasses.dataclass +class Case: + """A test case for compose-metric-mapping.""" + + name: str + req: fnv1.RunFunctionRequest + want: fnv1.RunFunctionResponse + + +def setUpModule() -> None: + logging.configure(level=logging.Level.DISABLED) + + +class TestFunctionRunner(unittest.IsolatedAsyncioTestCase): + """Tests for FunctionRunner.RunFunction.""" + + @classmethod + def setUpClass(cls) -> None: + cls.runner = fn.FunctionRunner() + + async def test_compose(self) -> None: + """The function reports whether a destination can actually be sent through.""" + exporters = { + "otlphttp": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}} + } + extensions = {"bearertokenauth": {"token": "${env:OTLP_TOKEN}"}} + + def xr(spec: dict) -> dict: + return { + "apiVersion": "modelplane.ai/v1alpha1", + "kind": "TelemetryDestination", + "metadata": {"name": "default"}, + "spec": spec, + } + + def req(spec: dict, secrets: list | None = None) -> fnv1.RunFunctionRequest: + r = fnv1.RunFunctionRequest( + observed=fnv1.State(composite=fnv1.Resource(resource=resource.dict_to_struct(xr(spec)))), + ) + if secrets is not None: + r.required_resources["secret"].items.extend([fnv1.Resource(resource=s) for s in secrets]) + return r + + def want( + ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition, secret: str | None = None + ) -> fnv1.RunFunctionResponse: + composite = fnv1.Resource(ready=ready) + if status is not None: + composite.resource.CopyFrom(resource.dict_to_struct(status)) + rsp = fnv1.RunFunctionResponse( + meta=fnv1.ResponseMeta(ttl=durationpb.Duration(seconds=60)), + desired=fnv1.State(composite=composite), + conditions=[cond], + context=structpb.Struct(), + ) + if secret is not None: + rsp.requirements.resources["secret"].api_version = "v1" + rsp.requirements.resources["secret"].kind = "Secret" + rsp.requirements.resources["secret"].match_name = secret + return rsp + + cases = [ + Case( + name="ready, naming the exporters it sends through", + req=req({"exporters": exporters, "extensions": extensions}), + want=want( + fnv1.READY_TRUE, + {"status": {}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Exporting through otlphttp", + ), + ), + ), + Case( + name="ready with an exporter that references no authenticator at all", + req=req( + {"exporters": {"prometheusremotewrite": {"endpoint": "https://prom.acme.example/api/v1/write"}}} + ), + want=want( + fnv1.READY_TRUE, + {"status": {}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Exporting through prometheusremotewrite", + ), + ), + ), + Case( + name="ready once the credential Secret exists", + req=req( + {"exporters": exporters, "extensions": extensions, "secretRef": {"name": "telemetry-credentials"}}, + secrets=[ + resource.dict_to_struct( + {"apiVersion": "v1", "kind": "Secret", "metadata": {"name": "telemetry-credentials"}} + ) + ], + ), + want=want( + fnv1.READY_TRUE, + {"status": {}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Exporting through otlphttp", + ), + secret="telemetry-credentials", + ), + ), + Case( + name="waits for the credential Secret to resolve", + req=req( + {"exporters": exporters, "extensions": extensions, "secretRef": {"name": "telemetry-credentials"}}, + ), + want=want( + fnv1.READY_FALSE, + None, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="WaitingForSecret", + message="Waiting for the credential Secret to resolve", + ), + secret="telemetry-credentials", + ), + ), + Case( + name="not ready when an exporter names an authenticator nothing defines", + req=req({"exporters": exporters}), + want=want( + fnv1.READY_FALSE, + None, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="UnknownAuthenticator", + message="No extension defines bearertokenauth, so the collector would refuse to start", + ), + ), + ), + Case( + name="an exporter that is not a mapping is left to the collector to reject", + req=req({"exporters": {"otlphttp": "https://otel.acme.example"}}), + want=want( + fnv1.READY_TRUE, + {"status": {}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Exporting through otlphttp", + ), + ), + ), + Case( + name="not ready with no exporters at all", + req=req({"exporters": {}}), + want=want( + fnv1.READY_FALSE, + None, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="NoExporters", + message="No exporters, so collected telemetry has nowhere to go", + ), + ), + ), + Case( + name="not ready when the credential Secret is missing", + req=req( + {"exporters": exporters, "extensions": extensions, "secretRef": {"name": "telemetry-credentials"}}, + secrets=[], + ), + want=want( + fnv1.READY_FALSE, + None, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_FALSE, + reason="SecretNotFound", + message=( + "Secret telemetry-credentials does not exist, " + "so the collector has no credential to send with" + ), + ), + secret="telemetry-credentials", + ), + ), + ] + + for case in cases: + with self.subTest(case.name): + got = await self.runner.RunFunction(case.req, None) + self.assertEqual( + json_format.MessageToDict(case.want), + json_format.MessageToDict(got), + "-want, +got", + ) From 7cea2bd2456603f54fec617a1a60d125546bd85e Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Mon, 28 Sep 2026 08:57:14 -0700 Subject: [PATCH 03/42] Regenerate schemas for the telemetry kinds x-kubernetes-preserve-unknown-fields generates as dict[str, Any], which is what a block passed through unread should be, and the collector enum generates as a Literal with its default. Signed-off-by: Dennis Ramdass --- schemas/.lock.json | 13 +- .../ai/modelplane/metricmapping/__init__.py | 0 .../ai/modelplane/metricmapping/v1alpha1.py | 124 +++++++++++++++++ .../telemetrydestination/__init__.py | 0 .../telemetrydestination/v1alpha1.py | 130 ++++++++++++++++++ 5 files changed, 260 insertions(+), 7 deletions(-) create mode 100644 schemas/python/models/ai/modelplane/metricmapping/__init__.py create mode 100644 schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py create mode 100644 schemas/python/models/ai/modelplane/telemetrydestination/__init__.py create mode 100644 schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py diff --git a/schemas/.lock.json b/schemas/.lock.json index f6003c682..981541a13 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,15 +1,14 @@ { "packages": { - "fs://apis": "a3c8867fa02ab5c25944ea69b769d69c5b8de1cf1d21ca61fc9c558159c7bf25", + "fs://apis": "cdfe7459a9629105e306637ccee45d6c6c252d1fc61680f08056b3c2d7f77f7e", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", - "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.8.1": "sha256:ca2e9e3b2e3a8b6ca44a9700d5abf7abd733cfa388d1afe9bb2bf7c847cd39ef", - "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.8.1": "sha256:ebb1bcd8dc9a7e60e97a324652fbd9d1b0609ddbb1107db518ce999cbc1113bd", - "xpkg://xpkg.upbound.io/upbound/provider-aws-eks:v2.8.1": "sha256:45ba27a14d6f8c4b9acdced0c48bec5650aab22bfc275af8bf1f5c770da3fa24", - "xpkg://xpkg.upbound.io/upbound/provider-aws-iam:v2.8.1": "sha256:ef802f20c76dc7d513d811beb7a9f8b4f5360eaf0d133c71bda289d02f766472", + "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", + "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", + "xpkg://xpkg.upbound.io/upbound/provider-aws-eks:v2.6.0": "sha256:5d144b19e188cb96c918aa7e4ccbc6759b8733ff09bd8ce412723c669aa3f763", + "xpkg://xpkg.upbound.io/upbound/provider-aws-iam:v2.6.0": "sha256:dbc5288589ccb302d527565680477f08477c280fc5c616dda95dfd558108a038", "xpkg://xpkg.upbound.io/upbound/provider-azure-containerservice:v2.6.0": "sha256:7d8a9bb3eb168e6eef0694253a23321fa98acad8b897ade168a4f0bb89985fed", "xpkg://xpkg.upbound.io/upbound/provider-azure-network:v2.6.0": "sha256:0d83bc4964488e5602b56dd7680e4bae0fbd5fa94b3a9af403d8431ded45635f", - "xpkg://xpkg.upbound.io/upbound/provider-civo:v1.0.0": "sha256:9cb7a795930d0b467a44d6b083560c48d977d802fb6be73e0ad909d9700874b7", - "xpkg://xpkg.upbound.io/upbound/provider-family-aws:v2.8.1": "sha256:b9c8e06e52c30c664a70997cc881e8e6b7823d7806a75319364930fcaf54df46", + "xpkg://xpkg.upbound.io/upbound/provider-family-aws:v2.6.0": "sha256:9fbe222866e9b763dae2db4de1599a52273125a75e8fe6328abbdc827695a74c", "xpkg://xpkg.upbound.io/upbound/provider-family-azure:v2.6.0": "sha256:1f2f597d5702ccb241429f0d8942f3c21e0ea55fffcef4ff9dca76b32bdc5dff", "xpkg://xpkg.upbound.io/upbound/provider-family-gcp:v2.6.0": "sha256:2e33bfd0f501155e7be63f13471d2f87f00519f99b916b7cd63ff40faa517145", "xpkg://xpkg.upbound.io/upbound/provider-gcp-cloudplatform:v2.6.0": "sha256:f1fe8bc55c474464642303e6fa8608c83e369b42ff12bb8a60a3e2d77339a52b", diff --git a/schemas/python/models/ai/modelplane/metricmapping/__init__.py b/schemas/python/models/ai/modelplane/metricmapping/__init__.py new file mode 100644 index 000000000..e69de29bb diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py new file mode 100644 index 000000000..531d7e152 --- /dev/null +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -0,0 +1,124 @@ +# generated by datamodel-codegen: +# filename: workdir/modelplane_ai_v1alpha1_metricmapping.yaml + +from __future__ import annotations + +from typing import Literal + +from pydantic import AwareDatetime, BaseModel, Field, RootModel, constr + +from ....io.k8s.apimachinery.pkg.apis.meta import v1 + + +class CompositionRef(BaseModel): + name: str + + +class CompositionRevisionRef(BaseModel): + name: str + + +class CompositionRevisionSelector(BaseModel): + matchLabels: dict[str, str] + + +class CompositionSelector(BaseModel): + matchLabels: dict[str, str] + + +class ResourceRef(BaseModel): + apiVersion: str + kind: str + name: str | None = None + namespace: str | None = None + + +class Crossplane(BaseModel): + compositionRef: CompositionRef | None = None + compositionRevisionRef: CompositionRevisionRef | None = None + compositionRevisionSelector: CompositionRevisionSelector | None = None + compositionSelector: CompositionSelector | None = None + compositionUpdatePolicy: Literal['Automatic', 'Manual'] | None = None + resourceRefs: list[ResourceRef] | None = None + + +class Statement(RootModel[constr(max_length=2048)]): + root: constr(max_length=2048) + + +class Spec(BaseModel): + crossplane: Crossplane | None = None + """ + Configures how Crossplane will reconcile this composite resource + """ + passthrough: bool | None = False + """ + Send this component's own metric names onward as well as the modelplane_* ones they become. + Off by default, because a series the statements did not rename is one whose meaning Modelplane cannot vouch for across engines, and it costs the same to carry as one that was renamed. On, for reading an engine's raw names during a migration or while debugging that engine. + A passed-through series is still merged across a deployment's replicas, so it keeps the labels the engine gave it and carries no pod identity. + """ + statements: list[Statement] | None = Field(None, max_length=128) + """ + OTTL statements, rendered into the collector's transform processor beside Modelplane's own. Modelplane does not interpret them: what you write here is the collector's own configuration language, documented by OpenTelemetry, and it is the same thing Modelplane writes for vLLM. + Statements select through their own where clauses, so nothing declares which engine a deployment runs. + """ + + +class Condition(BaseModel): + lastTransitionTime: AwareDatetime + message: str | None = None + observedGeneration: int | None = None + reason: str + status: str + type: str + + +class Status(BaseModel): + clusters: int | None = None + """ + How many inference clusters have taken these statements. + """ + conditions: list[Condition] | None = None + """ + Conditions of the resource. + """ + + +class MetricMapping(BaseModel): + apiVersion: Literal['modelplane.ai/v1alpha1'] | None = 'modelplane.ai/v1alpha1' + """ + APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + """ + kind: Literal['MetricMapping'] | None = 'MetricMapping' + """ + Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + """ + metadata: v1.ObjectMeta | None = None + """ + Standard object's metadata. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#metadata + """ + spec: Spec + """ + How one component's metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each inference cluster's collector, so a mapping is written once on the control plane and reaches the whole fleet. + A mapping naming a component Modelplane already provides statements for is additive: its statements run after the built-in ones and its flags apply. That is how passthrough is turned on for an engine that needs no statements of its own. + """ + status: Status | None = None + + +class MetricMappingList(BaseModel): + apiVersion: str | None = None + """ + APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + """ + items: list[MetricMapping] + """ + List of metricmappings. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md + """ + kind: str | None = None + """ + Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + """ + metadata: v1.ListMeta | None = None + """ + Standard list metadata. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + """ \ No newline at end of file diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/__init__.py b/schemas/python/models/ai/modelplane/telemetrydestination/__init__.py new file mode 100644 index 000000000..e69de29bb diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py new file mode 100644 index 000000000..3adecebd4 --- /dev/null +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -0,0 +1,130 @@ +# generated by datamodel-codegen: +# filename: workdir/modelplane_ai_v1alpha1_telemetrydestination.yaml + +from __future__ import annotations + +from typing import Any, Literal + +from pydantic import AwareDatetime, BaseModel, constr + +from ....io.k8s.apimachinery.pkg.apis.meta import v1 + + +class CompositionRef(BaseModel): + name: str + + +class CompositionRevisionRef(BaseModel): + name: str + + +class CompositionRevisionSelector(BaseModel): + matchLabels: dict[str, str] + + +class CompositionSelector(BaseModel): + matchLabels: dict[str, str] + + +class ResourceRef(BaseModel): + apiVersion: str + kind: str + name: str | None = None + namespace: str | None = None + + +class Crossplane(BaseModel): + compositionRef: CompositionRef | None = None + compositionRevisionRef: CompositionRevisionRef | None = None + compositionRevisionSelector: CompositionRevisionSelector | None = None + compositionSelector: CompositionSelector | None = None + compositionUpdatePolicy: Literal['Automatic', 'Manual'] | None = None + resourceRefs: list[ResourceRef] | None = None + + +class SecretRef(BaseModel): + name: constr(max_length=253) + """ + Name of the Secret, in Modelplane's namespace. + """ + + +class Spec(BaseModel): + collector: Literal['Composed', 'External'] | None = 'Composed' + """ + Whether Modelplane runs the fleet collector. Composed (the default) puts one on the control plane, and every inference cluster exports to it. External composes none, for a platform that already operates one: each inference cluster then exports to the endpoint below directly. + External gives up the single egress point, one place to change the destination, and the control plane's own series reaching the fleet without a path of their own. Whoever imposed the endpoint has usually provided them already. + """ + crossplane: Crossplane | None = None + """ + Configures how Crossplane will reconcile this composite resource + """ + exporters: dict[str, Any] + """ + The OpenTelemetry collector's exporters block, passed through unread. Modelplane validates that it parses and reports whether the destination accepts writes; it does not model what an exporter is. + So any exporter the collector provides works, with its TLS, retry and queue settings intact, and a destination keeps working when the collector gains an exporter Modelplane has never heard of. + """ + extensions: dict[str, Any] | None = None + """ + The collector's extensions block, passed through unread, for the authenticator an exporter references. Bearer token, basic auth, OIDC and SigV4 all work, because none of them is modelled here. + """ + secretRef: SecretRef | None = None + """ + A Secret whose keys Modelplane mounts into the collector as environment variables, so configuration above refers to ${env:TOKEN} and the credential itself never appears in this object or in kubectl output. + """ + + +class Condition(BaseModel): + lastTransitionTime: AwareDatetime + message: str | None = None + observedGeneration: int | None = None + reason: str + status: str + type: str + + +class Status(BaseModel): + conditions: list[Condition] | None = None + """ + Conditions of the resource. + """ + + +class TelemetryDestination(BaseModel): + apiVersion: Literal['modelplane.ai/v1alpha1'] | None = 'modelplane.ai/v1alpha1' + """ + APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + """ + kind: Literal['TelemetryDestination'] | None = 'TelemetryDestination' + """ + Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + """ + metadata: v1.ObjectMeta | None = None + """ + Standard object's metadata. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#metadata + """ + spec: Spec + """ + Where the fleet's telemetry goes. Modelplane composes no collectors until a TelemetryDestination exists: neither tier stores anything, so collecting with nowhere to export would spend GPU-cluster memory on samples nobody reads. Creating one turns collection on everywhere at once. + There is no per-deployment opt-out. A ModelDeployment's author owns neither the destination nor its bill. + """ + status: Status | None = None + + +class TelemetryDestinationList(BaseModel): + apiVersion: str | None = None + """ + APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + """ + items: list[TelemetryDestination] + """ + List of telemetrydestinations. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md + """ + kind: str | None = None + """ + Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + """ + metadata: v1.ListMeta | None = None + """ + Standard list metadata. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + """ \ No newline at end of file From 6ed19041ae597d43ee08267fbc74cc0dabe77270 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Mon, 28 Sep 2026 08:57:14 -0700 Subject: [PATCH 04/42] Add the telemetry guide What an operator reads: the metric surface, where to point a TelemetryDestination, what to write for an engine Modelplane ships no statements for, and migrating off the per-cluster Prometheus, with the rename table and the compatibility rules that keep an existing dashboard working while its panels are rewritten. draft: true until the kinds it describes are composing and collecting, so it stays out of the published site. The Vale vocabulary comes with it; every word it adds is one only this guide uses. Signed-off-by: Dennis Ramdass --- docs/content/guides/telemetry.md | 246 ++++++++++++++++++ .../config/vocabularies/Modelplane/accept.txt | 15 ++ 2 files changed, 261 insertions(+) create mode 100644 docs/content/guides/telemetry.md diff --git a/docs/content/guides/telemetry.md b/docs/content/guides/telemetry.md new file mode 100644 index 000000000..974a2acac --- /dev/null +++ b/docs/content/guides/telemetry.md @@ -0,0 +1,246 @@ +--- +title: Telemetry +weight: 20 +draft: true +aliases: +- /guides/collecting-engine-metrics/ +description: Collect normalized metrics across the fleet and send them anywhere that speaks OTLP. +--- + +{{< hint warning >}} +**Draft.** This page documents [the metrics design][design], which isn't built yet. It's +here to check the experience reads well before it's implemented, and it's excluded from the +site by `draft: true`. It replaces [Collecting engine metrics]({{< ref +"guides/collecting-engine-metrics.md" >}}) when the per-cluster Prometheus stack is +removed, and takes that page's URL with it. + +[design]: https://github.com/modelplaneai/modelplane/pull/363 +{{< /hint >}} + +Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from +every component Modelplane installs, which is more than your engines. It renames each +component's series to a single `modelplane_*` vocabulary and pushes to a collector on your +control plane. That collector is your fleet's +single egress point, and it sends to any backend that speaks OTLP. + +Modelplane has no API for this: nothing to write, and nothing to keep in sync as your +deployments change. + +## What you get + +Every series carries `cluster`. A series about a deployment also carries `deployment`, +`namespace`, `model`, and `engine`. Some of what you can read: + +| Metric | Means | +| --- | --- | +| `modelplane_frontend_ttft_seconds` | Time to the first token, measured at the gateway | +| `modelplane_frontend_tpot_seconds` | Time per output token, measured at the gateway | +| `modelplane_frontend_request_duration_seconds` | What the caller waited, end to end | +| `modelplane_request_queue_seconds` | How long a request waited before the engine started | +| `modelplane_requests_waiting` | Queue depth per engine | +| `modelplane_kv_cache_utilization_ratio` | KV-cache occupancy, averaged over replicas | +| `modelplane_kv_cache_utilization_ratio_max` | KV-cache occupancy of the busiest replica | +| `modelplane_tokens_total` | Tokens in and out, by `direction` | +| `modelplane_replica_gpus` | GPUs a replica holds | +| `modelplane_gpu_seconds_total` | GPU-time bound to serving | + +Latency appears twice on purpose. The `frontend_` series are what your caller experienced, +measured at the gateway. The engine's own series are what the engine spent. When the +frontend number is slow and the engine number isn't, the problem is routing, queueing, or +the network rather than the model. + +Saturation gauges come as a pair. The average is what you plan capacity against; the `_max` +is what you alert on, because three replicas at 0.3 and one at 0.99 average to something +comfortable while the fourth evicts and recomputes. A high `_max` beside +`modelplane_requests_preempted_total` climbing is one replica thrashing. + + +No series names a pod. Replicas are interchangeable, so they're summed before the metrics +leave the cluster; a rolling update would otherwise leave a dead series behind for every pod +it replaced. + + +## Sending it somewhere + +Create a `TelemetryDestination` naming whatever you already run: + +```yaml +apiVersion: modelplane.ai/v1alpha1 +kind: TelemetryDestination +metadata: + name: default +spec: + exporters: + otlphttp: + endpoint: https://otel.example.internal +``` + +`spec.exporters` is the OpenTelemetry collector's own exporters block, so any exporter the +collector provides works here, with its usual TLS and retry settings. + +Put credentials in a Secret and name it with `secretRef`. Modelplane mounts its keys into +the collector as environment variables, so your config refers to `${env:OTLP_TOKEN}` and the +token never appears in `kubectl get -o yaml`: + +```yaml +spec: + secretRef: + name: telemetry-credentials + extensions: + bearertokenauth: + token: ${env:OTLP_TOKEN} + exporters: + otlphttp: + endpoint: https://otel.example.internal + auth: + authenticator: bearertokenauth +``` + +If you run Prometheus, export to that instead and query the fleet there: + +```yaml +spec: + exporters: + prometheusremotewrite: + endpoint: https://prom.example.internal/api/v1/write +``` + +Until you create one, Modelplane composes no collectors: nothing here stores anything, so +collecting with nowhere to send it would spend GPU-cluster memory on samples nobody reads. +Creating a destination turns collection on everywhere at once, and there's no per-deployment +opt-out. + +Your clusters reach the control plane, and only the control plane reaches your backend. A +cluster with no route to your observability stack still reports, and the backend's +credential lives in one place instead of on every GPU cluster. + +## Computing rates, quantiles, and ratios + +A collector transforms each measurement as it passes it on. It holds no history, so it +produces no rates and no quantiles. Your backend does that. A fleet-wide p99: + +```promql +histogram_quantile(0.99, sum by (le) ( + rate(modelplane_frontend_ttft_seconds_bucket{model="Qwen/Qwen3-8B"}[5m]))) +``` + +Modelplane provides these as queries and Grafana dashboards rather than as precomputed +series. To precompute them, export to Prometheus and write recording rules there. + +## Engines + +vLLM and SGLang need no configuration. + +Any other OpenAI-compatible engine reports its top-line numbers with no configuration +either. The gateway measures those, not the engine, so `modelplane_frontend_*` and the token +counters work for an engine Modelplane has never seen. + +To normalize that engine's own metrics as well, create a `MetricMapping`: + +```yaml +apiVersion: modelplane.ai/v1alpha1 +kind: MetricMapping +metadata: + name: my-engine +spec: + statements: + - set(name, "modelplane_requests_waiting") + where name == "my_engine_queued_requests" +``` + +`spec.statements` are OTTL, the collector's own transform language. Modelplane renders them +into every cluster's collector, so you write them once. An engine with no statements is +still collected, under its own names. + +One engine needs a flag. SGLang publishes `/metrics` only when it runs with +`--enable-metrics`, so add it to the engine args. vLLM needs nothing. + +## Why engine latency and gateway latency differ + +`modelplane_request_ttft_seconds` comes from the engine, and engines bucket their +histograms differently. vLLM resolves down to a millisecond. SGLang resolves to a hundred +of them. A quantile across both is wrong, not approximate. Use the engine series to compare +one engine against itself, and the `frontend_` series for anything fleet-wide. + +Some measurements don't translate at all. SGLang's inter-token latency isn't vLLM's time per +output token, so neither is renamed onto a shared name. The gateway measures time per output +token for both. + +## Migrating from a hand-written `PodMonitor` + +[Collecting engine metrics]({{< ref "guides/collecting-engine-metrics.md" >}}) had you write +a `PodMonitor` and reach the in-cluster Prometheus over a `port-forward`. Both are gone, and +this page replaces that one. Three steps, and two of them fail quietly if you skip them. + +**Keep your Prometheus, and point a destination at it.** Collection becomes a push, so your +store stops scraping and starts receiving. Same Prometheus, same retention, same Grafana: + +```yaml +spec: + exporters: + prometheusremotewrite: + endpoint: http://prometheus.monitoring.svc:9090/api/v1/write +``` + +**Delete the monitors you wrote.** A `PodMonitor` or `ScrapeConfig` pointed at your engines +keeps working against your own Prometheus, so nothing appears to break and you collect +everything twice, under `vllm:*` and under `modelplane_*`, paying for both. One written +against Modelplane's Prometheus stops being read by anything, because the operator goes with +the stack. + +**Rewrite your dashboard queries.** Names change, and so do three labels: `model_name` +becomes `model`, pod labels are gone because replicas are summed before they leave the +cluster, and every series now carries `cluster`. + +| Was | Is | +| --- | --- | +| `vllm:time_to_first_token_seconds` | `modelplane_request_ttft_seconds` | +| `vllm:e2e_request_latency_seconds` | `modelplane_request_duration_seconds` | +| `vllm:request_queue_time_seconds` | `modelplane_request_queue_seconds` | +| `vllm:request_prefill_time_seconds` | `modelplane_request_prefill_seconds` | +| `vllm:request_decode_time_seconds` | `modelplane_request_decode_seconds` | +| `vllm:num_requests_running` | `modelplane_requests_running` | +| `vllm:num_requests_waiting` | `modelplane_requests_waiting` | +| `vllm:kv_cache_usage_perc` | `modelplane_kv_cache_utilization_ratio` | +| `vllm:num_preemptions_total` | `modelplane_requests_preempted_total` | +| `vllm:prefix_cache_hits_total` | `modelplane_prefix_cache_hits_total` | +| `vllm:prompt_tokens_total` | `modelplane_tokens_total{direction="input"}` | +| `vllm:generation_tokens_total` | `modelplane_tokens_total{direction="output"}` | +| `vllm:request_success_total{finished_reason}` | `modelplane_responses_total{reason}` | +| `DCGM_FI_DEV_FB_USED` | `modelplane_gpu_memory_used_bytes` | +| `DCGM_FI_DEV_GPU_TEMP` | `modelplane_gpu_temperature_celsius` | +| `DCGM_FI_DEV_POWER_USAGE` | `modelplane_gpu_power_watts` | +| `DCGM_FI_PROF_PIPE_TENSOR_ACTIVE` | `modelplane_gpu_tensor_active_ratio` | +| `envoy_cluster_upstream_rq_time` | `modelplane_frontend_request_duration_seconds` | +| `envoy_cluster_upstream_rq_xx` | `modelplane_requests_total{status}` | + +Two have no direct replacement. `vllm:inter_token_latency_seconds` isn't renamed, because +SGLang publishes a metric of the same name measuring something else; use +`modelplane_frontend_tpot_seconds`, which the gateway measures the same way for every +engine. `DCGM_FI_DEV_GPU_UTIL` isn't renamed either, because it only tells you the card +wasn't idle; use `modelplane_gpu_compute_active_ratio` and +`modelplane_gpu_tensor_active_ratio`. + +You can also skip the rewrite for now. Modelplane provides compatibility recording rules +that rebuild the old names from the new ones, so your existing dashboards keep working +untouched: + +```yaml +- record: vllm:time_to_first_token_seconds_bucket + expr: label_replace(modelplane_request_ttft_seconds_bucket, + "model_name", "$1", "model", "(.*)") +``` + +Load it into the Prometheus you already run and nothing on a dashboard changes. It's one +rule evaluation per metric over series your backend already holds, so it costs far less than +collecting everything twice. It covers the names in the table above, and it's meant to be +deleted once your panels use the new ones. + +For a raw series with no `modelplane_*` name at all, set `passthrough: true` on a +`MetricMapping` for that engine and its own names stay readable on the cluster. vLLM and +SGLang have no mapping of their own, so write one carrying just the flag: a mapping for an +engine Modelplane already knows adds to the built-in statements rather than replacing them. + +These series are still merged across replicas, so they keep `model_name` and your old +grouping works, but they carry no pod label. + diff --git a/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt b/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt index c0efe996f..7fbfda46e 100644 --- a/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt +++ b/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt @@ -21,6 +21,8 @@ EKSCluster ServingStack ModelCache ModelCaches +MetricMapping +MetricMappings ModelDeployment ModelDeployments ModelEndpoint @@ -29,6 +31,8 @@ ModelReplica ModelReplicas ModelService ModelServices +TelemetryDestination +TelemetryDestinations # Spec fields and config keys clusterSelector @@ -580,3 +584,14 @@ DGX BasePOD SuperPOD Run:ai +OTLP +OpenTelemetry +Grafana +quantile +quantiles +precompute +top-line +p99 +OTTL +per-deployment +opt-out From 0946764c8ef3cf2b0ee739ac8aef0c3995e67f5b Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Mon, 28 Sep 2026 09:21:05 -0700 Subject: [PATCH 05/42] Compose the OpenTelemetry collector The MetricMapping and TelemetryDestination kinds accept OTTL statements and exporter config, but nothing renders a collector from them, so a fleet that declares both still has no telemetry pipeline. Compose one collector per InferenceCluster from those kinds: a ConfigMap carrying the rendered YAML, a Deployment running it, and the RBAC the Prometheus receiver needs to discover pods. The receiver scrapes engines, the gateway and the substrate; the transform processor applies the built-in rename statements plus whatever the MetricMappings add, and the exporters come straight from the TelemetryDestination. The Deployment carries a checksum of the rendered config so a mapping or destination change restarts the collector. The checksum is a sha256 of the config rather than Python's hash(), which is seeded per process and would redeploy on every reconcile. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 295 ++++++++++++++++++ .../compose-serving-stack/function/fn.py | 77 ++++- .../function/stacks/__init__.py | 5 +- .../function/stacks/metrics.py | 106 +++++++ .../tests/test_collector.py | 124 ++++++++ .../compose-serving-stack/tests/test_fn.py | 15 +- 6 files changed, 618 insertions(+), 4 deletions(-) create mode 100644 functions/compose-serving-stack/function/collector.py create mode 100644 functions/compose-serving-stack/function/stacks/metrics.py create mode 100644 functions/compose-serving-stack/tests/test_collector.py diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py new file mode 100644 index 000000000..8d921c9ec --- /dev/null +++ b/functions/compose-serving-stack/function/collector.py @@ -0,0 +1,295 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""The OpenTelemetry collector each inference cluster runs. + +It scrapes everything Modelplane installs, renames what it scraped onto the +modelplane_* surface, merges each deployment's replicas into one series, and +exports to the destination. + +Rendered here rather than through the OpenTelemetry Operator: its +OpenTelemetryCollector CRD would be one more operator on every GPU cluster to +install, own and upgrade, for a Deployment and a ConfigMap that this composes +directly. The receiver does its own Kubernetes service discovery, so nothing +here needs the operator's target allocator either. +""" + +import hashlib +from typing import Any + +import yaml + +NAMESPACE = "modelplane-system" +NAME = "modelplane-collector" + +# Pinned rather than floating: a collector that silently changed what it +# renames on a chart bump would move the metric surface under an operator's +# dashboards. +IMAGE = "otel/opentelemetry-collector-contrib:0.139.0" + +# The label Modelplane stamps on every serving pod, and the port name it gives +# the engine's metrics. Both matter: an engine container's port is unnamed by +# default, and matching by number would find the pd-sidecar on a disaggregated +# pod rather than the engine behind it. +_SERVING_LABEL = "modelplane_ai_serving" +_METRICS_PORT = "http" + +_SCRAPE_INTERVAL = "15s" +_SUBSTRATE_INTERVAL = "30s" + + +def _relabel_pod_identity() -> list[dict[str, Any]]: + """Carry Modelplane's identity from the pod's labels onto every series. + + A recorded series inherits these, so a MetricMapping's statements need no + labels of their own. + """ + return [ + {"source_labels": [f"__meta_kubernetes_pod_label_{src}"], "target_label": dst} + for src, dst in ( + ("modelplane_ai_deployment", "deployment"), + ("modelplane_ai_model", "model"), + ("modelplane_ai_engine", "engine"), + ("modelplane_ai_role", "role"), + ) + ] + [{"source_labels": ["__meta_kubernetes_namespace"], "target_label": "namespace"}] + + +def _scrape_configs() -> list[dict[str, Any]]: + """What to scrape on an inference cluster. + + The gateway's GenAI metrics sit on the ext-proc sidecar's admin port rather + than the proxy's, so the front door needs a target of its own. + """ + return [ + { + "job_name": "modelplane-engines", + "scrape_interval": _SCRAPE_INTERVAL, + "kubernetes_sd_configs": [{"role": "pod"}], + "relabel_configs": [ + {"source_labels": [f"__meta_kubernetes_pod_label_{_SERVING_LABEL}"], "action": "keep", "regex": "true"}, + { + "source_labels": ["__meta_kubernetes_pod_container_port_name"], + "action": "keep", + "regex": _METRICS_PORT, + }, + *_relabel_pod_identity(), + ], + }, + { + "job_name": "modelplane-gateway", + "scrape_interval": _SCRAPE_INTERVAL, + "kubernetes_sd_configs": [{"role": "pod"}], + "relabel_configs": [ + { + "source_labels": ["__meta_kubernetes_pod_label_gateway_envoyproxy_io_owning_gateway_name"], + "action": "keep", + "regex": ".+", + }, + {"source_labels": ["__meta_kubernetes_pod_container_port_name"], "action": "keep", "regex": "metrics"}, + ], + }, + { + "job_name": "modelplane-substrate", + "scrape_interval": _SUBSTRATE_INTERVAL, + "kubernetes_sd_configs": [{"role": "pod"}], + "relabel_configs": [ + { + "source_labels": ["__meta_kubernetes_pod_annotation_prometheus_io_scrape"], + "action": "keep", + "regex": "true", + }, + {"source_labels": ["__meta_kubernetes_namespace"], "target_label": "namespace"}, + ], + }, + ] + + +def _transform(statements: list[str]) -> dict[str, Any]: + """Modelplane's renames, then whatever the MetricMappings add.""" + return {"metric_statements": [{"context": "metric", "statements": statements}]} + + +def config( + cluster: str, + statements: list[str], + exporters: dict[str, Any], + extensions: dict[str, Any], + *, + keep_raw: bool, +) -> str: + """The collector's configuration, as YAML. + + Only modelplane_* leaves the cluster unless a MetricMapping asked to keep an + engine's own names: a series the statements did not rename is one whose + meaning Modelplane cannot vouch for across engines, and it costs the same to + carry as one that was renamed. + """ + processors: dict[str, Any] = { + # cluster is stamped here rather than downstream: one receiver on the + # control plane sees a merged stream and cannot tell senders apart. + "resource/cluster": {"attributes": [{"key": "cluster", "value": cluster, "action": "upsert"}]}, + "transform/modelplane": _transform(statements), + # A pod's identity is a resource attribute, where a metric processor + # cannot reach it. Strip and merge the resources first, or the + # aggregation below combines nothing. + "groupbyattrs/replicas": {"keys": ["cluster", "namespace", "deployment", "model", "engine", "role"]}, + "batch": {"timeout": "10s"}, + } + pipeline = ["resource/cluster", "transform/modelplane", "groupbyattrs/replicas"] + if not keep_raw: + processors["filter/modelplane"] = { + "metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}, + } + pipeline.append("filter/modelplane") + pipeline.append("batch") + + service: dict[str, Any] = { + "pipelines": {"metrics": {"receivers": ["prometheus"], "processors": pipeline, "exporters": sorted(exporters)}}, + } + cfg: dict[str, Any] = { + "receivers": {"prometheus": {"config": {"scrape_configs": _scrape_configs()}}}, + "processors": processors, + "exporters": exporters, + "service": service, + } + if extensions: + cfg["extensions"] = extensions + # An authenticator the service does not list is one the collector will + # not load, and an exporter referencing it then fails at startup. + service["extensions"] = sorted(extensions) + return yaml.safe_dump(cfg, sort_keys=False) + + +def _digest(rendered: str) -> str: + """A stable hash of the rendered config, so the pod restarts when it changes.""" + return hashlib.sha256(rendered.encode()).hexdigest()[:16] + + +def objects( + cluster: str, + statements: list[str], + exporters: dict[str, Any], + extensions: dict[str, Any], + secret_name: str | None, + *, + keep_raw: bool, +) -> list[tuple[str, dict[str, Any], str | None]]: + """The collector as (key, manifest, readiness CEL) triples.""" + labels = {"app.kubernetes.io/name": NAME, "app.kubernetes.io/managed-by": "modelplane"} + rendered = config(cluster, statements, exporters, extensions, keep_raw=keep_raw) + volumes: list[dict[str, Any]] = [{"name": "config", "configMap": {"name": NAME}}] + mounts: list[dict[str, Any]] = [{"name": "config", "mountPath": "/conf"}] + env_from: list[dict[str, Any]] = [] + if secret_name: + # Mounted both ways. An environment variable is fixed for the life of a + # process, so a rotated credential would need a restart to be read; a + # mounted file is refreshed in place and an authenticator reading one + # picks the new credential up without one. + volumes.append({"name": "credentials", "secret": {"secretName": secret_name}}) + mounts.append({"name": "credentials", "mountPath": "/etc/modelplane/telemetry", "readOnly": True}) + env_from.append({"secretRef": {"name": secret_name}}) + + return [ + ( + "collector-serviceaccount", + { + "apiVersion": "v1", + "kind": "ServiceAccount", + "metadata": {"name": NAME, "namespace": NAMESPACE, "labels": labels}, + }, + None, + ), + ( + "collector-clusterrole", + { + "apiVersion": "rbac.authorization.k8s.io/v1", + "kind": "ClusterRole", + "metadata": {"name": NAME, "labels": labels}, + # Read-only, and only what Kubernetes service discovery needs. + "rules": [ + { + "apiGroups": [""], + "resources": ["pods", "services", "endpoints", "nodes", "nodes/metrics"], + "verbs": ["get", "list", "watch"], + }, + {"nonResourceURLs": ["/metrics"], "verbs": ["get"]}, + ], + }, + None, + ), + ( + "collector-clusterrolebinding", + { + "apiVersion": "rbac.authorization.k8s.io/v1", + "kind": "ClusterRoleBinding", + "metadata": {"name": NAME, "labels": labels}, + "roleRef": {"apiGroup": "rbac.authorization.k8s.io", "kind": "ClusterRole", "name": NAME}, + "subjects": [{"kind": "ServiceAccount", "name": NAME, "namespace": NAMESPACE}], + }, + None, + ), + ( + "collector-config", + { + "apiVersion": "v1", + "kind": "ConfigMap", + "metadata": {"name": NAME, "namespace": NAMESPACE, "labels": labels}, + "data": {"collector.yaml": rendered}, + }, + None, + ), + ( + "collector", + { + "apiVersion": "apps/v1", + "kind": "Deployment", + "metadata": {"name": NAME, "namespace": NAMESPACE, "labels": labels}, + "spec": { + "replicas": 1, + "selector": {"matchLabels": {"app.kubernetes.io/name": NAME}}, + "template": { + "metadata": { + "labels": {"app.kubernetes.io/name": NAME}, + # The config is a file, so a changed ConfigMap does + # not restart the pod on its own. Hashed with + # sha256 rather than hash(), whose string seed is + # randomised per process: the annotation would + # differ on every reconcile and redeploy the + # collector forever. + "annotations": {"modelplane.ai/config-hash": _digest(rendered)}, + }, + "spec": { + "serviceAccountName": NAME, + "containers": [ + { + "name": "collector", + "image": IMAGE, + "args": ["--config=/conf/collector.yaml"], + "volumeMounts": mounts, + **({"envFrom": env_from} if env_from else {}), + "resources": { + "requests": {"cpu": "100m", "memory": "256Mi"}, + "limits": {"memory": "512Mi"}, + }, + } + ], + "volumes": volumes, + }, + }, + }, + }, + "object.status.readyReplicas > 0", + ), + ] diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index 9f77439e0..258e327cc 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -35,10 +35,12 @@ """ import grpc -from crossplane.function import logging, resource, response +from crossplane.function import logging, request, resource, response from crossplane.function.proto.v1 import run_function_pb2 as fnv1 from crossplane.function.proto.v1 import run_function_pb2_grpc as grpcv1 from models.ai.modelplane.infrastructure.servingstack import v1alpha1 +from models.ai.modelplane.metricmapping import v1alpha1 as mmv1alpha1 +from models.ai.modelplane.telemetrydestination import v1alpha1 as tdv1alpha1 from models.io.crossplane.m.helm.providerconfig import v1beta1 as helmpcv1beta1 from models.io.crossplane.m.helm.release import v1beta1 as helmv1beta1 from models.io.crossplane.m.kubernetes.object import v1alpha1 as k8sobjv1alpha1 @@ -48,7 +50,7 @@ from models.io.crossplane.protection.usage import v1beta1 as usagev1beta1 from models.io.k8s.apimachinery.pkg.apis.meta import v1 as metav1 -from function import gateway, stacks +from function import collector, gateway, stacks # Label key every rendered Release and Object carries, valued with its # composed-resource key, so Usage resourceSelectors can name any @@ -320,6 +322,7 @@ def compose(self) -> None: rendered = self.compose_components(components) rendered += self.compose_gateway() rendered += self.compose_gateway_pki() + rendered += self.compose_collector() self.compose_component_usages(components) self.compose_gateway_usages() self.write_status() @@ -744,6 +747,76 @@ def compose_gateway_pki(self) -> list[str]: rendered.append("gateway-client-auth") return rendered + def compose_collector(self) -> list[str]: + """Compose the collector that gathers this cluster's telemetry. + + Nothing until a TelemetryDestination exists. Neither collector stores + anything, so collecting with nowhere to export is GPU-cluster memory and + CPU spent on samples nobody will ever read; a fleet that has not said + where its telemetry goes gets none composed. + + Returns the composed-resource keys it rendered, for readiness. + """ + response.require_resources( + self.rsp, + name="destinations", + api_version="modelplane.ai/v1alpha1", + kind="TelemetryDestination", + ) + response.require_resources( + self.rsp, + name="mappings", + api_version="modelplane.ai/v1alpha1", + kind="MetricMapping", + ) + if "destinations" not in self.req.required_resources or "mappings" not in self.req.required_resources: + return [] + + destinations = list(request.get_required_resources(self.req, "destinations")) + if not destinations: + return [] + dest = tdv1alpha1.TelemetryDestination.model_validate(destinations[0]) + if len(destinations) > 1: + # Which one wins would otherwise be whichever the API server listed + # first, and a fleet would export somewhere nobody chose. + response.warning( + self.rsp, + f"{len(destinations)} TelemetryDestinations exist; using " + f"{_name(dest.metadata)}. Telemetry has one destination per fleet.", + ) + + statements: list[str] = list(stacks.METRIC_STATEMENTS) + keep_raw = False + for m in request.get_required_resources(self.req, "mappings"): + mapping = mmv1alpha1.MetricMapping.model_validate(m) + statements += [str(st.root) for st in mapping.spec.statements or []] + keep_raw = keep_raw or bool(mapping.spec.passthrough) + + pc_observed = self.provider_configs_observed() + pc = _pc_name(self.xr) + rendered: list[str] = [] + for key, manifest, cel in collector.objects( + cluster=_name(self.xr.metadata), + statements=statements, + exporters=dict(dest.spec.exporters or {}), + extensions=dict(dest.spec.extensions or {}), + secret_name=dest.spec.secretRef.name if dest.spec.secretRef else None, + keep_raw=keep_raw, + ): + if not (pc_observed or key in self.req.observed.resources): + continue + resource.update( + self.rsp.desired.resources[key], + _k8s_object( + pc, + manifest, + metadata=metav1.ObjectMeta(labels={_LABEL_RESOURCE: key}), + ready_when=cel, + ), + ) + rendered.append(key) + return rendered + def compose_gateway_usages(self) -> None: """Compose Usages ordering the hand-rendered gateway teardown. diff --git a/functions/compose-serving-stack/function/stacks/__init__.py b/functions/compose-serving-stack/function/stacks/__init__.py index 6e8d240a6..6d07236db 100644 --- a/functions/compose-serving-stack/function/stacks/__init__.py +++ b/functions/compose-serving-stack/function/stacks/__init__.py @@ -21,12 +21,14 @@ stack's own file. See design/serving-stack-generation.md. """ -from function.stacks import common, components, dynamo, standard +from function.stacks import common, components, dynamo, metrics, standard from function.stacks.clouds import civo, existing, nebius, vultr from function.stacks.clouds.generated.aicr import aks, eks, gke from function.stacks.components import Chart, Cloud, Component, Manifests, Stack +from function.stacks.metrics import METRIC_STATEMENTS __all__ = [ + "METRIC_STATEMENTS", "Chart", "Cloud", "Component", @@ -35,6 +37,7 @@ "clouds", "components", "join", + "metrics", "stacks", ] diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py new file mode 100644 index 000000000..39c87de23 --- /dev/null +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -0,0 +1,106 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""What Modelplane calls each metric the components it installs emit. + +Reviewed, pinned data, the same discipline as the component lists. A +MetricMapping is the extension point for a component Modelplane ships no +statements for; these are the ones it does. + +Renaming is only safe where the measurements agree. SGLang's +inter_token_latency is not vLLM's time per output token, so neither is +renamed onto a shared name and the gateway supplies that measurement for +both. A histogram is renamed only where its bucket boundaries match, which is +why SGLang's latency histograms are absent here: they resolve to a hundred +milliseconds where vLLM's resolve to one, and a quantile across the two is +wrong rather than approximate. +""" + +# The front door, which measures every request it proxies under the +# OpenTelemetry GenAI conventions, for whatever engine is behind it. These are +# the SLO metrics: one component, one bucket layout, so a fleet quantile over +# them is sound. +_GATEWAY = { + "gen_ai_server_request_duration_seconds": "modelplane_frontend_request_duration_seconds", + "gen_ai_server_time_to_first_token_seconds": "modelplane_frontend_ttft_seconds", + "gen_ai_server_time_per_output_token_seconds": "modelplane_frontend_tpot_seconds", +} + +# The engines, which explain what the gateway measured. +_VLLM = { + "vllm:time_to_first_token_seconds": "modelplane_request_ttft_seconds", + "vllm:e2e_request_latency_seconds": "modelplane_request_duration_seconds", + "vllm:request_queue_time_seconds": "modelplane_request_queue_seconds", + "vllm:request_prefill_time_seconds": "modelplane_request_prefill_seconds", + "vllm:request_decode_time_seconds": "modelplane_request_decode_seconds", + "vllm:request_prompt_tokens": "modelplane_request_input_tokens", + "vllm:request_generation_tokens": "modelplane_request_output_tokens", + "vllm:num_requests_running": "modelplane_requests_running", + "vllm:num_requests_waiting": "modelplane_requests_waiting", + "vllm:kv_cache_usage_perc": "modelplane_kv_cache_utilization_ratio", + "vllm:num_preemptions_total": "modelplane_requests_preempted_total", + "vllm:prefix_cache_hits_total": "modelplane_prefix_cache_hits_total", + "vllm:prefix_cache_queries_total": "modelplane_prefix_cache_lookups_total", +} + +_SGLANG = { + "sglang:queue_time_seconds": "modelplane_request_queue_seconds", + "sglang:num_running_reqs": "modelplane_requests_running", + "sglang:num_queue_reqs": "modelplane_requests_waiting", + "sglang:token_usage": "modelplane_kv_cache_utilization_ratio", + "sglang:num_retracted_requests_total": "modelplane_requests_preempted_total", + "sglang:num_retracted_input_tokens_total": "modelplane_tokens_recomputed_total", + "sglang:prompt_tokens_histogram": "modelplane_request_input_tokens", + "sglang:generation_tokens_histogram": "modelplane_request_output_tokens", +} + +# The endpoint picker. A router's queue is a different measurement from an +# engine's, so it keeps a name of its own. +_PICKER = { + "llm_d_epp_scheduler_e2e_duration_seconds": "modelplane_route_decision_seconds", +} + +# The GPUs, through whichever vendor's exporter the stack installed. DCGM +# reports energy in millijoules, which the unit in the name says it is not. +_GPU = { + "DCGM_FI_DEV_FB_USED": "modelplane_gpu_memory_used_bytes", + "DCGM_FI_PROF_GR_ENGINE_ACTIVE": "modelplane_gpu_compute_active_ratio", + "DCGM_FI_PROF_PIPE_TENSOR_ACTIVE": "modelplane_gpu_tensor_active_ratio", + "DCGM_FI_PROF_DRAM_ACTIVE": "modelplane_gpu_memory_bandwidth_ratio", + "DCGM_FI_DEV_GPU_TEMP": "modelplane_gpu_temperature_celsius", + "DCGM_FI_DEV_POWER_USAGE": "modelplane_gpu_power_watts", +} + + +def _rename(pairs: dict[str, str]) -> list[str]: + return [f'set(name, "{new}") where name == "{old}"' for old, new in pairs.items()] + + +def statements() -> list[str]: + """Every rename Modelplane ships, as OTTL.""" + return ( + _rename(_GATEWAY) + + _rename(_VLLM) + + _rename(_SGLANG) + + _rename(_PICKER) + + _rename(_GPU) + # DCGM counts energy in millijoules. The scale is here rather than left + # to a query, because a name ending in _joules_total that holds + # millijoules is the kind of thing nobody notices until a bill. + + ['set(value_double, value_double / 1000) where name == "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION"'] + + _rename({"DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION": "modelplane_energy_joules_total"}) + ) + + +METRIC_STATEMENTS = statements() diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py new file mode 100644 index 000000000..69dcb66b7 --- /dev/null +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -0,0 +1,124 @@ +# Copyright 2026 The Modelplane Authors. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Tests for the collector this stack composes.""" + +import unittest + +import yaml +from function import collector, stacks + +_EXPORTERS = {"otlphttp": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}}} +_EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/token"}} + + +def _config(*, keep_raw: bool = False, extensions: dict | None = None) -> dict: + return yaml.safe_load( + collector.config( + "prod-us-east", + list(stacks.METRIC_STATEMENTS), + _EXPORTERS, + _EXTENSIONS if extensions is None else extensions, + keep_raw=keep_raw, + ) + ) + + +class TestConfig(unittest.TestCase): + """The collector configuration this renders.""" + + def test_pipeline_order(self) -> None: + """groupbyattrs runs before the merge, or the merge combines nothing. + + A pod's identity is a resource attribute, which a metric processor + can't see, so the resources have to be stripped and merged first. + """ + procs = _config()["service"]["pipelines"]["metrics"]["processors"] + self.assertLess(procs.index("transform/modelplane"), procs.index("groupbyattrs/replicas")) + self.assertEqual(procs[-1], "batch") + + def test_only_modelplane_leaves_the_cluster(self) -> None: + """A series the statements didn't rename is dropped, unless asked for.""" + self.assertIn("filter/modelplane", _config()["processors"]) + self.assertNotIn("filter/modelplane", _config(keep_raw=True)["processors"]) + + def test_cluster_is_stamped_here(self) -> None: + """One receiver downstream sees a merged stream and can't tell senders apart.""" + attrs = _config()["processors"]["resource/cluster"]["attributes"] + self.assertEqual(attrs, [{"key": "cluster", "value": "prod-us-east", "action": "upsert"}]) + + def test_engine_scrape_selects_the_port_by_name(self) -> None: + """Matching by number would find the pd-sidecar on a disaggregated pod.""" + jobs = {j["job_name"]: j for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"]} + keeps = [r for r in jobs["modelplane-engines"]["relabel_configs"] if r.get("action") == "keep"] + self.assertIn("__meta_kubernetes_pod_container_port_name", [k["source_labels"][0] for k in keeps]) + + def test_gateway_has_a_target_of_its_own(self) -> None: + """Its GenAI metrics are on the ext-proc sidecar, not the proxy's port.""" + jobs = [j["job_name"] for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"]] + self.assertIn("modelplane-gateway", jobs) + + def test_extensions_are_declared_to_the_service(self) -> None: + """An authenticator the service doesn't list is one the collector won't load.""" + self.assertEqual(_config()["service"]["extensions"], ["bearertokenauth"]) + self.assertNotIn("extensions", _config(extensions={})["service"]) + + def test_energy_is_scaled_before_it_is_renamed(self) -> None: + """DCGM counts millijoules, and the name says joules.""" + statements = stacks.METRIC_STATEMENTS + scale = next(i for i, s in enumerate(statements) if "value_double / 1000" in s) + rename = next(i for i, s in enumerate(statements) if "modelplane_energy_joules_total" in s) + self.assertLess(scale, rename) + + def test_sglang_latency_histograms_are_not_renamed(self) -> None: + """Their buckets resolve to 100ms where vLLM's resolve to 1ms.""" + joined = " ".join(stacks.METRIC_STATEMENTS) + self.assertNotIn("sglang:time_to_first_token_seconds", joined) + self.assertNotIn("sglang:inter_token_latency", joined) + + +class TestObjects(unittest.TestCase): + """The manifests this composes.""" + + def _objects(self, secret: str | None = None) -> dict: + return { + k: m + for k, m, _ in collector.objects( + "prod-us-east", list(stacks.METRIC_STATEMENTS), _EXPORTERS, _EXTENSIONS, secret, keep_raw=False + ) + } + + def test_config_hash_is_stable_across_processes(self) -> None: + """hash() is seeded per process, so it would redeploy on every reconcile.""" + first = self._objects()["collector"]["spec"]["template"]["metadata"]["annotations"] + second = self._objects()["collector"]["spec"]["template"]["metadata"]["annotations"] + self.assertEqual(first, second) + self.assertRegex(first["modelplane.ai/config-hash"], r"^[0-9a-f]{16}$") + + def test_credentials_mount_as_a_file_and_an_environment_variable(self) -> None: + """A rotated token in an environment variable needs a restart to be read.""" + pod = self._objects(secret="telemetry-credentials")["collector"]["spec"]["template"]["spec"] + self.assertIn("credentials", [v["name"] for v in pod["volumes"]]) + self.assertEqual(pod["containers"][0]["envFrom"], [{"secretRef": {"name": "telemetry-credentials"}}]) + + def test_no_secret_mounts_nothing(self) -> None: + pod = self._objects()["collector"]["spec"]["template"]["spec"] + self.assertEqual([v["name"] for v in pod["volumes"]], ["config"]) + self.assertNotIn("envFrom", pod["containers"][0]) + + def test_rbac_is_read_only(self) -> None: + """Service discovery needs to list pods, and nothing needs to write.""" + rules = self._objects()["collector-clusterrole"]["rules"] + verbs = {v for r in rules for v in r["verbs"]} + self.assertEqual(verbs, {"get", "list", "watch"}) diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index f041b2edb..a57eb673d 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -829,7 +829,12 @@ def _existing_dynamo_stack() -> dict[str, fnv1.Resource]: def _response(resources: dict[str, fnv1.Resource], status: dict | None = None) -> fnv1.RunFunctionResponse: - """A whole expected response: 60s TTL, empty context, the XR status.""" + """A whole expected response: 60s TTL, empty context, the XR status. + + Every response asks for the telemetry kinds, because the collector is + composed from them and the function cannot know whether any exist until + they resolve. + """ return fnv1.RunFunctionResponse( meta=fnv1.ResponseMeta(ttl=durationpb.Duration(seconds=60)), desired=fnv1.State( @@ -837,6 +842,14 @@ def _response(resources: dict[str, fnv1.Resource], status: dict | None = None) - resources=resources, ), context=structpb.Struct(), + requirements=fnv1.Requirements( + resources={ + "destinations": fnv1.ResourceSelector( + api_version="modelplane.ai/v1alpha1", kind="TelemetryDestination" + ), + "mappings": fnv1.ResourceSelector(api_version="modelplane.ai/v1alpha1", kind="MetricMapping"), + } + ), ) From 6eaf26f523fdba544f5c4afd9bb6829abf9ad4e2 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 09:37:25 -0700 Subject: [PATCH 06/42] Narrow the telemetry API to what's in use Two fields went in ahead of any demand for them. TelemetryDestination's collector enum chose between composing a fleet collector and deferring to someone else's. Nothing composes a fleet collector yet, and nothing reads the field: each inference cluster's collector exports straight to the destination's exporters, so a platform that already runs one points those exporters at its own endpoint and the enum never comes up. When a fleet collector does land, whether to compose it is a composition-level choice, not a field every user reads. MetricMapping's passthrough kept an engine's own metric names flowing during a migration off the Prometheus stack. Migrations aren't what this project should be optimising for yet, and the flag was a bool where an enum would belong if it came back. Dropping it means only modelplane_* leaves a cluster, which takes a branch out of the collector config too. Three smaller corrections alongside: the collector image was pinned 23 releases back, at 0.139.0; both XRDs were missing the category that files them under Platform in the API reference; and the guide offered Grafana dashboards that don't exist anywhere in this repo. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 21 +------ apis/telemetrydestinations/definition.yaml | 20 +------ docs/content/guides/telemetry.md | 16 +++--- .../compose-metric-mapping/function/fn.py | 10 +--- .../compose-metric-mapping/tests/test_fn.py | 17 +----- .../function/collector.py | 56 +++++++++---------- .../compose-serving-stack/function/fn.py | 3 - .../function/stacks/metrics.py | 3 +- .../tests/test_collector.py | 9 ++- schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 8 +-- .../telemetrydestination/v1alpha1.py | 5 -- 12 files changed, 48 insertions(+), 122 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index def56a04b..467fd28c8 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -5,7 +5,7 @@ metadata: spec: group: modelplane.ai names: - categories: [crossplane, modelplane] + categories: [crossplane, modelplane, platform] kind: MetricMapping plural: metricmappings shortNames: [mm] @@ -36,8 +36,7 @@ spec: A mapping naming a component Modelplane already provides statements for is additive: its statements run after the - built-in ones and its flags apply. That is how passthrough is - turned on for an engine that needs no statements of its own. + built-in ones. properties: statements: type: array @@ -54,22 +53,6 @@ spec: type: string maxLength: 2048 maxItems: 128 - passthrough: - type: boolean - default: false - description: >- - Send this component's own metric names onward as well as the - modelplane_* ones they become. - - Off by default, because a series the statements did not - rename is one whose meaning Modelplane cannot vouch for - across engines, and it costs the same to carry as one that - was renamed. On, for reading an engine's raw names during a - migration or while debugging that engine. - - A passed-through series is still merged across a - deployment's replicas, so it keeps the labels the engine - gave it and carries no pod identity. status: type: object properties: diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index 4c40dbe33..3a0273f98 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -5,7 +5,7 @@ metadata: spec: group: modelplane.ai names: - categories: [crossplane, modelplane] + categories: [crossplane, modelplane, platform] kind: TelemetryDestination plural: telemetrydestinations shortNames: [td] @@ -15,9 +15,6 @@ spec: served: true referenceable: true additionalPrinterColumns: - - name: COLLECTOR - type: string - jsonPath: .spec.collector - name: AGE type: date jsonPath: .metadata.creationTimestamp @@ -39,21 +36,6 @@ spec: There is no per-deployment opt-out. A ModelDeployment's author owns neither the destination nor its bill. properties: - collector: - type: string - default: Composed - enum: [Composed, External] - description: >- - Whether Modelplane runs the fleet collector. Composed (the - default) puts one on the control plane, and every inference - cluster exports to it. External composes none, for a - platform that already operates one: each inference cluster - then exports to the endpoint below directly. - - External gives up the single egress point, one place to - change the destination, and the control plane's own series - reaching the fleet without a path of their own. Whoever - imposed the endpoint has usually provided them already. exporters: type: object x-kubernetes-preserve-unknown-fields: true diff --git a/docs/content/guides/telemetry.md b/docs/content/guides/telemetry.md index 974a2acac..89eaa7a26 100644 --- a/docs/content/guides/telemetry.md +++ b/docs/content/guides/telemetry.md @@ -124,8 +124,9 @@ histogram_quantile(0.99, sum by (le) ( rate(modelplane_frontend_ttft_seconds_bucket{model="Qwen/Qwen3-8B"}[5m]))) ``` -Modelplane provides these as queries and Grafana dashboards rather than as precomputed -series. To precompute them, export to Prometheus and write recording rules there. +Modelplane has no dashboards of its own. What it exports is counters and histogram buckets, and +your backend derives the rates and quantiles at query time. To precompute them instead, +export to Prometheus and write recording rules there. ## Engines @@ -236,11 +237,8 @@ rule evaluation per metric over series your backend already holds, so it costs f collecting everything twice. It covers the names in the table above, and it's meant to be deleted once your panels use the new ones. -For a raw series with no `modelplane_*` name at all, set `passthrough: true` on a -`MetricMapping` for that engine and its own names stay readable on the cluster. vLLM and -SGLang have no mapping of their own, so write one carrying just the flag: a mapping for an -engine Modelplane already knows adds to the built-in statements rather than replacing them. - -These series are still merged across replicas, so they keep `model_name` and your old -grouping works, but they carry no pod label. +A series no statement renames doesn't leave the cluster. If a panel needs an engine's +own name, write a `MetricMapping` that renames it onto the `modelplane_*` surface: a +mapping for an engine Modelplane already knows adds to the built-in statements rather +than replacing them. diff --git a/functions/compose-metric-mapping/function/fn.py b/functions/compose-metric-mapping/function/fn.py index 5c0d9468d..7729fc57d 100644 --- a/functions/compose-metric-mapping/function/fn.py +++ b/functions/compose-metric-mapping/function/fn.py @@ -70,14 +70,8 @@ async def RunFunction( clusters = len(list(request.get_required_resources(req, "clusters"))) resource.update_status(rsp.desired.composite, v1alpha1.Status(clusters=clusters)) - # A mapping carrying neither statements nor passthrough does nothing at - # all, which is worth saying rather than reporting ready. - if not xr.spec.statements and not xr.spec.passthrough: - _not_ready( - rsp, - CONDITION_REASON_NO_STATEMENTS, - "No statements and passthrough is off, so this mapping changes nothing", - ) + if not xr.spec.statements: + _not_ready(rsp, CONDITION_REASON_NO_STATEMENTS, "No statements, so this mapping changes nothing") return rsp if clusters == 0: diff --git a/functions/compose-metric-mapping/tests/test_fn.py b/functions/compose-metric-mapping/tests/test_fn.py index 76209d80b..6afc75fdd 100644 --- a/functions/compose-metric-mapping/tests/test_fn.py +++ b/functions/compose-metric-mapping/tests/test_fn.py @@ -84,7 +84,6 @@ def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.R ) no_statements = {**mapping, "spec": {}} - passthrough_only = {**mapping, "spec": {"passthrough": True}} cases = [ Case( @@ -101,20 +100,6 @@ def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.R ), ), ), - Case( - name="ready on passthrough alone, which is a mapping with no statements to write", - req=req(passthrough_only, [cluster]), - want=want( - fnv1.READY_TRUE, - {"status": {"clusters": 1}}, - fnv1.Condition( - type="Accepted", - status=fnv1.STATUS_CONDITION_TRUE, - reason="Available", - message="Rendered into 1 inference cluster(s)", - ), - ), - ), Case( name="not ready when no cluster exists to render into", req=req(mapping, []), @@ -139,7 +124,7 @@ def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.R type="Accepted", status=fnv1.STATUS_CONDITION_FALSE, reason="NoStatements", - message="No statements and passthrough is off, so this mapping changes nothing", + message="No statements, so this mapping changes nothing", ), ), ), diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 8d921c9ec..14e79637c 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -33,10 +33,7 @@ NAMESPACE = "modelplane-system" NAME = "modelplane-collector" -# Pinned rather than floating: a collector that silently changed what it -# renames on a chart bump would move the metric surface under an operator's -# dashboards. -IMAGE = "otel/opentelemetry-collector-contrib:0.139.0" +IMAGE = "otel/opentelemetry-collector-contrib:0.161.0" # The label Modelplane stamps on every serving pod, and the port name it gives # the engine's metrics. Both matter: an engine container's port is unnamed by @@ -126,15 +123,12 @@ def config( statements: list[str], exporters: dict[str, Any], extensions: dict[str, Any], - *, - keep_raw: bool, ) -> str: """The collector's configuration, as YAML. - Only modelplane_* leaves the cluster unless a MetricMapping asked to keep an - engine's own names: a series the statements did not rename is one whose - meaning Modelplane cannot vouch for across engines, and it costs the same to - carry as one that was renamed. + Only modelplane_* leaves the cluster: a series the statements did not rename + is one whose meaning Modelplane cannot vouch for across engines, and it + costs the same to carry as one that was renamed. """ processors: dict[str, Any] = { # cluster is stamped here rather than downstream: one receiver on the @@ -145,15 +139,16 @@ def config( # cannot reach it. Strip and merge the resources first, or the # aggregation below combines nothing. "groupbyattrs/replicas": {"keys": ["cluster", "namespace", "deployment", "model", "engine", "role"]}, + "filter/modelplane": {"metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}}, "batch": {"timeout": "10s"}, } - pipeline = ["resource/cluster", "transform/modelplane", "groupbyattrs/replicas"] - if not keep_raw: - processors["filter/modelplane"] = { - "metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}, - } - pipeline.append("filter/modelplane") - pipeline.append("batch") + pipeline = [ + "resource/cluster", + "transform/modelplane", + "groupbyattrs/replicas", + "filter/modelplane", + "batch", + ] service: dict[str, Any] = { "pipelines": {"metrics": {"receivers": ["prometheus"], "processors": pipeline, "exporters": sorted(exporters)}}, @@ -183,12 +178,20 @@ def objects( exporters: dict[str, Any], extensions: dict[str, Any], secret_name: str | None, - *, - keep_raw: bool, ) -> list[tuple[str, dict[str, Any], str | None]]: - """The collector as (key, manifest, readiness CEL) triples.""" + """The collector as (key, manifest, readiness CEL) triples. + + Composed here rather than as a Manifests entry in the stack because the + stack is fixed at build time and every object here depends on request-time + data: the ConfigMap holds config rendered from the TelemetryDestination and + the MetricMappings, the Deployment carries that config's digest and the + destination's optional Secret mounts, and none of it exists at all until a + TelemetryDestination does. The ServiceAccount and RBAC would fit the stack, + but splitting one component across two mechanisms would put a collector's + permissions on clusters running no collector. + """ labels = {"app.kubernetes.io/name": NAME, "app.kubernetes.io/managed-by": "modelplane"} - rendered = config(cluster, statements, exporters, extensions, keep_raw=keep_raw) + rendered = config(cluster, statements, exporters, extensions) volumes: list[dict[str, Any]] = [{"name": "config", "configMap": {"name": NAME}}] mounts: list[dict[str, Any]] = [{"name": "config", "mountPath": "/conf"}] env_from: list[dict[str, Any]] = [] @@ -217,7 +220,6 @@ def objects( "apiVersion": "rbac.authorization.k8s.io/v1", "kind": "ClusterRole", "metadata": {"name": NAME, "labels": labels}, - # Read-only, and only what Kubernetes service discovery needs. "rules": [ { "apiGroups": [""], @@ -262,12 +264,10 @@ def objects( "template": { "metadata": { "labels": {"app.kubernetes.io/name": NAME}, - # The config is a file, so a changed ConfigMap does - # not restart the pod on its own. Hashed with - # sha256 rather than hash(), whose string seed is - # randomised per process: the annotation would - # differ on every reconcile and redeploy the - # collector forever. + # A changed ConfigMap doesn't restart the pod on + # its own. sha256 rather than hash(), whose string + # seed is randomised per process and would redeploy + # the collector on every reconcile. "annotations": {"modelplane.ai/config-hash": _digest(rendered)}, }, "spec": { diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index 258e327cc..ae5489380 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -786,11 +786,9 @@ def compose_collector(self) -> list[str]: ) statements: list[str] = list(stacks.METRIC_STATEMENTS) - keep_raw = False for m in request.get_required_resources(self.req, "mappings"): mapping = mmv1alpha1.MetricMapping.model_validate(m) statements += [str(st.root) for st in mapping.spec.statements or []] - keep_raw = keep_raw or bool(mapping.spec.passthrough) pc_observed = self.provider_configs_observed() pc = _pc_name(self.xr) @@ -801,7 +799,6 @@ def compose_collector(self) -> list[str]: exporters=dict(dest.spec.exporters or {}), extensions=dict(dest.spec.extensions or {}), secret_name=dest.spec.secretRef.name if dest.spec.secretRef else None, - keep_raw=keep_raw, ): if not (pc_observed or key in self.req.observed.resources): continue diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index 39c87de23..e0a244286 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -71,8 +71,7 @@ "llm_d_epp_scheduler_e2e_duration_seconds": "modelplane_route_decision_seconds", } -# The GPUs, through whichever vendor's exporter the stack installed. DCGM -# reports energy in millijoules, which the unit in the name says it is not. +# The GPUs, through whichever vendor's exporter the stack installed. _GPU = { "DCGM_FI_DEV_FB_USED": "modelplane_gpu_memory_used_bytes", "DCGM_FI_PROF_GR_ENGINE_ACTIVE": "modelplane_gpu_compute_active_ratio", diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 69dcb66b7..d31045b5e 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -23,14 +23,13 @@ _EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/token"}} -def _config(*, keep_raw: bool = False, extensions: dict | None = None) -> dict: +def _config(*, extensions: dict | None = None) -> dict: return yaml.safe_load( collector.config( "prod-us-east", list(stacks.METRIC_STATEMENTS), _EXPORTERS, _EXTENSIONS if extensions is None else extensions, - keep_raw=keep_raw, ) ) @@ -49,9 +48,9 @@ def test_pipeline_order(self) -> None: self.assertEqual(procs[-1], "batch") def test_only_modelplane_leaves_the_cluster(self) -> None: - """A series the statements didn't rename is dropped, unless asked for.""" + """A series the statements didn't rename is dropped.""" self.assertIn("filter/modelplane", _config()["processors"]) - self.assertNotIn("filter/modelplane", _config(keep_raw=True)["processors"]) + self.assertIn("filter/modelplane", _config()["service"]["pipelines"]["metrics"]["processors"]) def test_cluster_is_stamped_here(self) -> None: """One receiver downstream sees a merged stream and can't tell senders apart.""" @@ -95,7 +94,7 @@ def _objects(self, secret: str | None = None) -> dict: return { k: m for k, m, _ in collector.objects( - "prod-us-east", list(stacks.METRIC_STATEMENTS), _EXPORTERS, _EXTENSIONS, secret, keep_raw=False + "prod-us-east", list(stacks.METRIC_STATEMENTS), _EXPORTERS, _EXTENSIONS, secret ) } diff --git a/schemas/.lock.json b/schemas/.lock.json index 981541a13..5b2f5d9c2 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "cdfe7459a9629105e306637ccee45d6c6c252d1fc61680f08056b3c2d7f77f7e", + "fs://apis": "e4bb7953a7e73cb25d317f8beb70973a11cfbdd2bf5b75e570292e94bc9087a7", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index 531d7e152..c85ad7d88 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -51,12 +51,6 @@ class Spec(BaseModel): """ Configures how Crossplane will reconcile this composite resource """ - passthrough: bool | None = False - """ - Send this component's own metric names onward as well as the modelplane_* ones they become. - Off by default, because a series the statements did not rename is one whose meaning Modelplane cannot vouch for across engines, and it costs the same to carry as one that was renamed. On, for reading an engine's raw names during a migration or while debugging that engine. - A passed-through series is still merged across a deployment's replicas, so it keeps the labels the engine gave it and carries no pod identity. - """ statements: list[Statement] | None = Field(None, max_length=128) """ OTTL statements, rendered into the collector's transform processor beside Modelplane's own. Modelplane does not interpret them: what you write here is the collector's own configuration language, documented by OpenTelemetry, and it is the same thing Modelplane writes for vLLM. @@ -100,7 +94,7 @@ class MetricMapping(BaseModel): spec: Spec """ How one component's metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each inference cluster's collector, so a mapping is written once on the control plane and reaches the whole fleet. - A mapping naming a component Modelplane already provides statements for is additive: its statements run after the built-in ones and its flags apply. That is how passthrough is turned on for an engine that needs no statements of its own. + A mapping naming a component Modelplane already provides statements for is additive: its statements run after the built-in ones. """ status: Status | None = None diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index 3adecebd4..5bad09805 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -50,11 +50,6 @@ class SecretRef(BaseModel): class Spec(BaseModel): - collector: Literal['Composed', 'External'] | None = 'Composed' - """ - Whether Modelplane runs the fleet collector. Composed (the default) puts one on the control plane, and every inference cluster exports to it. External composes none, for a platform that already operates one: each inference cluster then exports to the endpoint below directly. - External gives up the single egress point, one place to change the destination, and the control plane's own series reaching the fleet without a path of their own. Whoever imposed the endpoint has usually provided them already. - """ crossplane: Crossplane | None = None """ Configures how Crossplane will reconcile this composite resource From 8ada9b6c4266536d7383aee0c3f79b504ddc2204 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 09:43:20 -0700 Subject: [PATCH 07/42] Run a value rewrite in the context that can reach the value The collector refuses to start on the config this composes. The statement that scales DCGM's energy counter reads value_double, which is a datapoint path, and every statement was rendered into a single metric-context block: segment "value_double" from path "metric.value_double" is not a valid path nor a valid OTTL keyword for the metric context Render two blocks instead, the datapoint one first. It has to be a separate block rather than an earlier line in the same one: the transform processor finishes a block over every datapoint before it starts the next, so a rename sharing the block would run after the first datapoint was scaled and leave every later datapoint unmatched, and unscaled. Verified by running `validate` in the collector image itself, against the rendered config, with and without exporter authentication. Also build the renames Modelplane provides as MetricMappings rather than as a bare list of statements, so they reach the collector through the same path an operator's mapping does. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 17 +++++- .../compose-serving-stack/function/fn.py | 11 ++-- .../function/stacks/__init__.py | 4 +- .../function/stacks/metrics.py | 55 ++++++++++++++----- .../tests/test_collector.py | 37 ++++++++----- 5 files changed, 89 insertions(+), 35 deletions(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 14e79637c..2842e3d38 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -30,6 +30,8 @@ import yaml +from function.stacks import metrics as stacks_metrics + NAMESPACE = "modelplane-system" NAME = "modelplane-collector" @@ -114,8 +116,19 @@ def _scrape_configs() -> list[dict[str, Any]]: def _transform(statements: list[str]) -> dict[str, Any]: - """Modelplane's renames, then whatever the MetricMappings add.""" - return {"metric_statements": [{"context": "metric", "statements": statements}]} + """Value rewrites first, then every rename. + + Two blocks rather than one list: a statement reaching a datapoint's value + can't run in the metric context, and the processor finishes a block over + every datapoint before it starts the next, which is what keeps a rename + from stranding the datapoints a value rewrite hasn't reached yet. + """ + return { + "metric_statements": [ + {"context": "datapoint", "statements": list(stacks_metrics.DATAPOINT_STATEMENTS)}, + {"context": "metric", "statements": statements}, + ] + } def config( diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index ae5489380..ca4b8454d 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -785,10 +785,13 @@ def compose_collector(self) -> list[str]: f"{_name(dest.metadata)}. Telemetry has one destination per fleet.", ) - statements: list[str] = list(stacks.METRIC_STATEMENTS) - for m in request.get_required_resources(self.req, "mappings"): - mapping = mmv1alpha1.MetricMapping.model_validate(m) - statements += [str(st.root) for st in mapping.spec.statements or []] + # Modelplane's own mappings first, then the operator's, which add to + # them rather than replacing them. + mappings = list(stacks.BUILTIN_MAPPINGS) + mappings += [ + mmv1alpha1.MetricMapping.model_validate(m) for m in request.get_required_resources(self.req, "mappings") + ] + statements: list[str] = [str(st.root) for mp in mappings for st in mp.spec.statements or []] pc_observed = self.provider_configs_observed() pc = _pc_name(self.xr) diff --git a/functions/compose-serving-stack/function/stacks/__init__.py b/functions/compose-serving-stack/function/stacks/__init__.py index 6d07236db..5e803c978 100644 --- a/functions/compose-serving-stack/function/stacks/__init__.py +++ b/functions/compose-serving-stack/function/stacks/__init__.py @@ -25,10 +25,10 @@ from function.stacks.clouds import civo, existing, nebius, vultr from function.stacks.clouds.generated.aicr import aks, eks, gke from function.stacks.components import Chart, Cloud, Component, Manifests, Stack -from function.stacks.metrics import METRIC_STATEMENTS +from function.stacks.metrics import BUILTIN_MAPPINGS __all__ = [ - "METRIC_STATEMENTS", + "BUILTIN_MAPPINGS", "Chart", "Cloud", "Component", diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index e0a244286..9d19a932f 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -27,6 +27,9 @@ wrong rather than approximate. """ +from models.ai.modelplane.metricmapping import v1alpha1 +from models.io.k8s.apimachinery.pkg.apis.meta import v1 as metav1 + # The front door, which measures every request it proxies under the # OpenTelemetry GenAI conventions, for whatever engine is behind it. These are # the SLO metrics: one component, one bucket layout, so a fleet quantile over @@ -86,20 +89,44 @@ def _rename(pairs: dict[str, str]) -> list[str]: return [f'set(name, "{new}") where name == "{old}"' for old, new in pairs.items()] -def statements() -> list[str]: - """Every rename Modelplane ships, as OTTL.""" - return ( - _rename(_GATEWAY) - + _rename(_VLLM) - + _rename(_SGLANG) - + _rename(_PICKER) - + _rename(_GPU) - # DCGM counts energy in millijoules. The scale is here rather than left - # to a query, because a name ending in _joules_total that holds - # millijoules is the kind of thing nobody notices until a bill. - + ['set(value_double, value_double / 1000) where name == "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION"'] - + _rename({"DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION": "modelplane_energy_joules_total"}) +def _mapping(name: str, statements: list[str]) -> v1alpha1.MetricMapping: + return v1alpha1.MetricMapping( + metadata=metav1.ObjectMeta(name=name), + spec=v1alpha1.Spec(statements=[v1alpha1.Statement(st) for st in statements]), ) -METRIC_STATEMENTS = statements() +def mappings() -> list[v1alpha1.MetricMapping]: + """The mappings Modelplane provides, as the kind an operator would write. + + Built as MetricMappings rather than as a bare list of statements so the + collector renders Modelplane's own renames through the same path as an + operator's, and a built-in that breaks breaks the path everyone uses. + """ + return [ + _mapping("modelplane-gateway", _rename(_GATEWAY)), + _mapping("modelplane-vllm", _rename(_VLLM)), + _mapping("modelplane-sglang", _rename(_SGLANG)), + _mapping("modelplane-picker", _rename(_PICKER)), + _mapping( + "modelplane-gpu", + _rename(_GPU) + _rename({"DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION": "modelplane_energy_joules_total"}), + ), + ] + + +BUILTIN_MAPPINGS = mappings() + +# A datapoint's value is out of reach of the metric context, so the one +# statement that rewrites a value rather than a name runs in its own block +# ahead of the renames. It has to be a block of its own rather than an earlier +# line: the transform processor finishes a block over every datapoint before +# starting the next, and a rename landing first would leave every datapoint +# after the first unmatched and unscaled. +# +# DCGM counts energy in millijoules. Left to a query instead, a name ending in +# _joules_total holding millijoules is the kind of thing nobody notices until a +# bill. +DATAPOINT_STATEMENTS = [ + 'set(value_double, value_double / 1000) where metric.name == "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION"' +] diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index d31045b5e..9cf052aa3 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -23,11 +23,15 @@ _EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/token"}} +def _statements() -> list[str]: + return [str(st.root) for mp in stacks.BUILTIN_MAPPINGS for st in mp.spec.statements or []] + + def _config(*, extensions: dict | None = None) -> dict: return yaml.safe_load( collector.config( "prod-us-east", - list(stacks.METRIC_STATEMENTS), + _statements(), _EXPORTERS, _EXTENSIONS if extensions is None else extensions, ) @@ -74,15 +78,27 @@ def test_extensions_are_declared_to_the_service(self) -> None: self.assertNotIn("extensions", _config(extensions={})["service"]) def test_energy_is_scaled_before_it_is_renamed(self) -> None: - """DCGM counts millijoules, and the name says joules.""" - statements = stacks.METRIC_STATEMENTS - scale = next(i for i, s in enumerate(statements) if "value_double / 1000" in s) - rename = next(i for i, s in enumerate(statements) if "modelplane_energy_joules_total" in s) - self.assertLess(scale, rename) + """DCGM counts millijoules, and the name says joules. + + The scale is a block ahead of the renames, not a line ahead. The + processor finishes a block over every datapoint before the next one + starts, so a rename sharing the block would strand every datapoint + after the first at millijoules. + """ + blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] + self.assertEqual([b["context"] for b in blocks], ["datapoint", "metric"]) + self.assertTrue(any("value_double / 1000" in st for st in blocks[0]["statements"])) + self.assertTrue(any("modelplane_energy_joules_total" in st for st in blocks[1]["statements"])) + + def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: + """value_double is a datapoint path; the collector refuses to start on it here.""" + blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] + metric_block = next(b for b in blocks if b["context"] == "metric") + self.assertFalse([st for st in metric_block["statements"] if "value_double" in st]) def test_sglang_latency_histograms_are_not_renamed(self) -> None: """Their buckets resolve to 100ms where vLLM's resolve to 1ms.""" - joined = " ".join(stacks.METRIC_STATEMENTS) + joined = " ".join(_statements()) self.assertNotIn("sglang:time_to_first_token_seconds", joined) self.assertNotIn("sglang:inter_token_latency", joined) @@ -91,12 +107,7 @@ class TestObjects(unittest.TestCase): """The manifests this composes.""" def _objects(self, secret: str | None = None) -> dict: - return { - k: m - for k, m, _ in collector.objects( - "prod-us-east", list(stacks.METRIC_STATEMENTS), _EXPORTERS, _EXTENSIONS, secret - ) - } + return {k: m for k, m, _ in collector.objects("prod-us-east", _statements(), _EXPORTERS, _EXTENSIONS, secret)} def test_config_hash_is_stable_across_processes(self) -> None: """hash() is seeded per process, so it would redeploy on every reconcile.""" From 8c758254b0eb4dd6a70f56b81e330e7f6f222514 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 09:58:12 -0700 Subject: [PATCH 08/42] Convert framebuffer memory to the unit its name claims DCGM reports DCGM_FI_DEV_FB_USED in MiB. It is renamed to modelplane_gpu_memory_used_bytes, with no conversion, so the series reads about a millionth of the memory actually in use. Scale it in the datapoint block beside the energy counter, which had the same problem in millijoules and was already handled there. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-serving-stack/function/stacks/metrics.py | 10 ++++++---- .../compose-serving-stack/tests/test_collector.py | 7 +++++++ 2 files changed, 13 insertions(+), 4 deletions(-) diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index 9d19a932f..8bb49409b 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -124,9 +124,11 @@ def mappings() -> list[v1alpha1.MetricMapping]: # starting the next, and a rename landing first would leave every datapoint # after the first unmatched and unscaled. # -# DCGM counts energy in millijoules. Left to a query instead, a name ending in -# _joules_total holding millijoules is the kind of thing nobody notices until a -# bill. +# DCGM reports energy in millijoules and framebuffer memory in MiB, and both +# are renamed onto a name that states a different unit. Left to a query +# instead, a name ending in _joules_total holding millijoules is the kind of +# thing nobody notices until a bill. DATAPOINT_STATEMENTS = [ - 'set(value_double, value_double / 1000) where metric.name == "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION"' + 'set(value_double, value_double / 1000) where metric.name == "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION"', + 'set(value_double, value_double * 1048576) where metric.name == "DCGM_FI_DEV_FB_USED"', ] diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 9cf052aa3..2f8b85cdb 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -90,6 +90,13 @@ def test_energy_is_scaled_before_it_is_renamed(self) -> None: self.assertTrue(any("value_double / 1000" in st for st in blocks[0]["statements"])) self.assertTrue(any("modelplane_energy_joules_total" in st for st in blocks[1]["statements"])) + def test_dcgm_units_are_converted_to_the_unit_the_name_claims(self) -> None: + """DCGM reports mJ and MiB; the names say joules and bytes.""" + blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] + scales = " ".join(blocks[0]["statements"]) + self.assertIn("DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION", scales) + self.assertIn("DCGM_FI_DEV_FB_USED", scales) + def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: """value_double is a datapoint path; the collector refuses to start on it here.""" blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] From a4cc7900540c206affe1ceb342a82c45b02b5132 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 10:17:24 -0700 Subject: [PATCH 09/42] Describe the metrics instead of the collector's transform language A MetricMapping held OTTL statements, which put the collector's own configuration language in a Modelplane API. It leaves nothing to validate at the apiserver, pins what an operator writes to one collector's grammar, and commits Modelplane to OTTL for as long as the kind lives. Take the metrics themselves instead: `from` is what the component emits, `to` is what Modelplane calls it, and `fromUnit` says what the source is measured in when the target name claims something else. compose-serving-stack compiles them to OTTL, which is where knowing the collector belongs. The unit field is not speculative. DCGM reports framebuffer memory in MiB and energy in millijoules, and both are renamed onto names claiming bytes and joules; we had the energy conversion and were missing the memory one. A field that asks the question catches that where a hand-written statement doesn't. Every rename Modelplane provides is expressible here, so the built-ins and an operator's mapping are now the same shape as well as the same code path. Left out: folding two metrics into one under a label, and taking a histogram's count as a counter. Both are real, neither has a caller yet. TelemetryDestination grows the same treatment on its own terms. The exporters map becomes a list of named sinks, each with the exporter `type` to send with and that exporter's `config` passed through unread, because the exporter schema is upstream and versioned separately - typing it would mean a Modelplane release per exporter, and dropping TLS, retry and queue settings that destinations actually need. What does become typed is the part that was broken: `secretRef` moves onto the sink, so two sinks no longer share one Secret and tell their keys apart by prefix, and each mounts under its own directory. Both kinds are APIs a platform engineer uses, so the guide moves out of guides/ and into the Platform section as "Monitor the Fleet", and the Prometheus-stack page it replaces is deleted rather than left to contradict it. The two pages that linked there now link here. Verified by running `validate` in the collector image against the compiled config, with two sinks, per-sink credentials, and an operator mapping carrying a unit conversion. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 60 +++++++--- apis/telemetrydestinations/definition.yaml | 92 ++++++++++----- .../guides/collecting-engine-metrics.md | 67 ----------- docs/content/models/model-service.md | 2 +- .../content/{guides => platform}/telemetry.md | 104 +++++++++++------ docs/content/recipes/qwen2.5-72b.md | 2 +- docs/data/apigroups.yaml | 2 + .../compose-metric-mapping/function/fn.py | 18 +-- .../compose-metric-mapping/tests/test_fn.py | 24 +--- .../function/collector.py | 105 +++++++++++++----- .../compose-serving-stack/function/fn.py | 6 +- .../function/stacks/metrics.py | 63 +++++------ .../tests/test_collector.py | 53 +++++++-- .../function/fn.py | 52 +++++---- .../tests/test_fn.py | 81 ++++++-------- schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 28 +++-- .../telemetrydestination/v1alpha1.py | 34 ++++-- 18 files changed, 453 insertions(+), 342 deletions(-) delete mode 100644 docs/content/guides/collecting-engine-metrics.md rename docs/content/{guides => platform}/telemetry.md (79%) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index 467fd28c8..63258af48 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -28,6 +28,7 @@ spec: properties: spec: type: object + required: [metrics] description: >- How one component's metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each @@ -35,24 +36,57 @@ spec: the control plane and reaches the whole fleet. A mapping naming a component Modelplane already provides - statements for is additive: its statements run after the - built-in ones. + renames for is additive: its renames run after the built-in + ones, and a later rename of the same metric wins. properties: - statements: + metrics: type: array + minItems: 1 + maxItems: 128 description: >- - OTTL statements, rendered into the collector's transform - processor beside Modelplane's own. Modelplane does not - interpret them: what you write here is the collector's own - configuration language, documented by OpenTelemetry, and it - is the same thing Modelplane writes for vLLM. + The metrics this component emits, and what Modelplane calls + them. - Statements select through their own where clauses, so - nothing declares which engine a deployment runs. + Rename only where the measurements agree. Two engines' + histograms sharing a name are worth less than nothing if + their buckets disagree, because a quantile across them is + wrong rather than approximate. items: - type: string - maxLength: 2048 - maxItems: 128 + type: object + required: [from, to] + properties: + from: + type: string + maxLength: 255 + description: >- + The metric's name as the component emits it, matched + exactly. Nothing here declares which engine a + deployment runs: a name that no component emits simply + matches nothing. + to: + type: string + maxLength: 255 + pattern: '^modelplane_[a-z0-9_]*[a-z0-9]$' + description: >- + What Modelplane calls it. Only modelplane_* leaves a + cluster, so a metric with no name here is one nobody + downstream can read. + fromUnit: + type: string + enum: [Millijoules, Mebibytes, Milliseconds, Nanoseconds] + description: >- + What the component measures this in, when that isn't + the unit the name claims. Modelplane converts to the + base unit: millijoules and milliseconds are divided by + a thousand, nanoseconds by a billion, and mebibytes + multiplied out to bytes. + + Say it whenever the source disagrees with the target, + even where the factor looks obvious. A name ending in + _bytes that holds mebibytes is the kind of thing + nobody notices until a capacity review, and stating + the source unit is what makes the conversion happen at + all. status: type: object properties: diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index 3a0273f98..7bf078b02 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -25,7 +25,7 @@ spec: properties: spec: type: object - required: [exporters] + required: [sinks] description: >- Where the fleet's telemetry goes. Modelplane composes no collectors until a TelemetryDestination exists: neither tier @@ -36,40 +36,80 @@ spec: There is no per-deployment opt-out. A ModelDeployment's author owns neither the destination nor its bill. properties: - exporters: - type: object - x-kubernetes-preserve-unknown-fields: true + sinks: + type: array + minItems: 1 + maxItems: 16 description: >- - The OpenTelemetry collector's exporters block, passed - through unread. Modelplane validates that it parses and - reports whether the destination accepts writes; it does not - model what an exporter is. + Where to send it. Every sink gets the whole stream, so two + sinks is two copies of the fleet's metrics, billed twice. + x-kubernetes-list-type: map + x-kubernetes-list-map-keys: [name] + items: + type: object + required: [name, type, config] + properties: + name: + type: string + maxLength: 63 + pattern: '^[a-z0-9]([-a-z0-9]*[a-z0-9])?$' + description: >- + This sink's name, unique within the destination. It + names the collector's exporter instance and the + directory its credential mounts at, so renaming one + restarts the collector. + type: + type: string + maxLength: 63 + description: >- + The collector exporter to send with, by the name + OpenTelemetry gives it: otlphttp, otlp, + prometheusremotewrite, kafka, and every other one the + collector provides. + + Not an enum, because enumerating them here would mean + a Modelplane release for each exporter the collector + gains, and the collector already refuses to start on a + name it doesn't have. + config: + type: object + x-kubernetes-preserve-unknown-fields: true + description: >- + That exporter's own configuration, passed through + unread. Modelplane does not model what an exporter is, + so TLS, retry, queue and compression settings all + work, and a sink keeps working when the collector + gains a setting Modelplane has never heard of. + secretRef: + type: object + required: [name] + description: >- + A Secret holding this sink's credential. Its keys + reach the collector as environment variables, for + config above referring to ${env:TOKEN}, and as files + under /etc/modelplane/telemetry//, for an + authenticator reading one from disk. - So any exporter the collector provides works, with its TLS, - retry and queue settings intact, and a destination keeps - working when the collector gains an exporter Modelplane has - never heard of. + Per sink rather than per destination, so two sinks + with different credentials don't have to share one + Secret and tell their keys apart by prefix. A file is + refreshed in place where an environment variable is + fixed for the life of the process, so an authenticator + reading the file picks up a rotated credential without + a restart. + properties: + name: + type: string + maxLength: 253 + description: Name of the Secret, in Modelplane's namespace. extensions: type: object x-kubernetes-preserve-unknown-fields: true description: >- The collector's extensions block, passed through unread, - for the authenticator an exporter references. Bearer token, + for the authenticator a sink references. Bearer token, basic auth, OIDC and SigV4 all work, because none of them is modelled here. - secretRef: - type: object - required: [name] - description: >- - A Secret whose keys Modelplane mounts into the collector as - environment variables, so configuration above refers to - ${env:TOKEN} and the credential itself never appears in this - object or in kubectl output. - properties: - name: - type: string - maxLength: 253 - description: Name of the Secret, in Modelplane's namespace. status: type: object properties: diff --git a/docs/content/guides/collecting-engine-metrics.md b/docs/content/guides/collecting-engine-metrics.md deleted file mode 100644 index ae6ecff51..000000000 --- a/docs/content/guides/collecting-engine-metrics.md +++ /dev/null @@ -1,67 +0,0 @@ ---- -title: Collecting engine metrics -weight: 20 -description: Scrape a vLLM engine's Prometheus metrics through the in-cluster Prometheus. ---- - -Scraping an inference engine's Prometheus metrics, shown on the smallest serving -shape: a 0.5B Qwen chat model on one NVIDIA L4. vLLM publishes metrics at -`/metrics` on its serving port with no extra flag, and Modelplane runs a -Prometheus on every workload cluster with `PodMonitor` discovery open across -namespaces, so scraping the engine is a `PodMonitor` plus a `port-forward`. The -model is only the subject; the same wiring fits any engine, with the SGLang, -leader/worker, and prefill/decode differences noted at the end. - -This was run end to end on GKE. The `InferenceClass` and `ModelDeployment` are the -exact manifests from that run, and the `PodMonitor` below scraped this deployment. -Apply the platform side first, then the ML side. - -## Platform - -{{< manifests "guides/collecting-engine-metrics/inference-class.yaml" >}} - -{{< manifests "guides/collecting-engine-metrics/inference-cluster.yaml" >}} - -## Deployment - -{{< manifests "guides/collecting-engine-metrics/model-deployment.yaml" >}} - -{{< manifests "guides/collecting-engine-metrics/model-service.yaml" >}} - -## Scraping the metrics - -The `PodMonitor` selects engine pods by the `modelplane.ai/serving` label -Modelplane stamps on them, and the `monitoring` namespace Prometheus discovers any -`PodMonitor`, so this is the whole config. The engine container port is unnamed, -so reference it by number with `targetPort`: - -{{< manifests "guides/collecting-engine-metrics/podmonitor.yaml" >}} - -The engine pods and the `PodMonitor` CRD live on the workload cluster, not the -control plane, so apply it there. The pods run in the namespace Modelplane -mirrors `ml-team` into, which `kubectl get ns -l modelplane.ai/namespace=ml-team` -finds. Then read the metrics from the in-cluster Prometheus over a -`port-forward`: - -```bash -kubectl -n monitoring port-forward svc/prometheus-prometheus 9090:9090 # workload cluster -# open http://localhost:9090, Status > Targets to confirm the scrape, then query -# e.g. vllm:num_requests_running or vllm:gpu_cache_usage_perc -``` - -### Other engine shapes - -The `PodMonitor` above fits a single-pod vLLM engine. The selector and port shift -by shape: - -- **SGLang**: exposes `/metrics` only when the engine runs with - `--enable-metrics`; otherwise it's identical (same selector, `targetPort: 8000`). -- **Leader/worker**: only the leader serves the API and carries - `modelplane.ai/serving`, so the selector above already scrapes the leader alone; - the workers expose nothing. -- **prefill/decode**: two engines, labelled `llm-d.ai/role: prefill` and - `llm-d.ai/role: decode`. The prefill engine serves on `8000`; the decode engine - sits behind the routing sidecar that takes `8000` and listens on `8001`, so - scrape decode with `targetPort: 8001`. Select each by its role label to keep - them apart. - diff --git a/docs/content/models/model-service.md b/docs/content/models/model-service.md index fec06d836..a366694c8 100644 --- a/docs/content/models/model-service.md +++ b/docs/content/models/model-service.md @@ -240,7 +240,7 @@ fails, so don't mix one into a service that OpenAI callers use. Scrape an engine's own operational paths like `/metrics` and `/health` from the replica directly. See -[Collecting engine metrics]({{< ref "/guides/collecting-engine-metrics" >}}). +[Monitor the Fleet]({{< ref "/platform/telemetry" >}}). ## Example diff --git a/docs/content/guides/telemetry.md b/docs/content/platform/telemetry.md similarity index 79% rename from docs/content/guides/telemetry.md rename to docs/content/platform/telemetry.md index 89eaa7a26..82aefbfbf 100644 --- a/docs/content/guides/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -1,21 +1,12 @@ --- -title: Telemetry -weight: 20 -draft: true +title: Monitor the Fleet +weight: 37 aliases: - /guides/collecting-engine-metrics/ +- /guides/telemetry/ description: Collect normalized metrics across the fleet and send them anywhere that speaks OTLP. --- -{{< hint warning >}} -**Draft.** This page documents [the metrics design][design], which isn't built yet. It's -here to check the experience reads well before it's implemented, and it's excluded from the -site by `draft: true`. It replaces [Collecting engine metrics]({{< ref -"guides/collecting-engine-metrics.md" >}}) when the per-cluster Prometheus stack is -removed, and takes that page's URL with it. - -[design]: https://github.com/modelplaneai/modelplane/pull/363 -{{< /hint >}} Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from every component Modelplane installs, which is more than your engines. It renames each @@ -70,27 +61,31 @@ kind: TelemetryDestination metadata: name: default spec: - exporters: - otlphttp: + sinks: + - name: primary + type: otlphttp + config: endpoint: https://otel.example.internal ``` -`spec.exporters` is the OpenTelemetry collector's own exporters block, so any exporter the -collector provides works here, with its usual TLS and retry settings. +`type` names a collector exporter, and `config` is that exporter's own configuration, so any +exporter the collector provides works here with its usual TLS and retry settings. -Put credentials in a Secret and name it with `secretRef`. Modelplane mounts its keys into -the collector as environment variables, so your config refers to `${env:OTLP_TOKEN}` and the -token never appears in `kubectl get -o yaml`: +Put credentials in a Secret and name it with the sink's `secretRef`. Modelplane mounts its +keys as environment variables, so your config refers to `${env:OTLP_TOKEN}` and the token +never appears in `kubectl get -o yaml`: ```yaml spec: - secretRef: - name: telemetry-credentials extensions: bearertokenauth: token: ${env:OTLP_TOKEN} - exporters: - otlphttp: + sinks: + - name: primary + type: otlphttp + secretRef: + name: telemetry-credentials + config: endpoint: https://otel.example.internal auth: authenticator: bearertokenauth @@ -100,11 +95,33 @@ If you run Prometheus, export to that instead and query the fleet there: ```yaml spec: - exporters: - prometheusremotewrite: + sinks: + - name: prometheus + type: prometheusremotewrite + config: + endpoint: https://prom.example.internal/api/v1/write +``` + +Name more than one sink and every one gets the whole stream. Each carries its own +credential, so a vendor and your own Prometheus don't have to share a Secret: + +```yaml +spec: + sinks: + - name: vendor + type: otlphttp + secretRef: + name: vendor-token + config: + endpoint: https://otel.vendor.example + - name: prometheus + type: prometheusremotewrite + config: endpoint: https://prom.example.internal/api/v1/write ``` +That is two copies of the fleet's metrics, billed twice. + Until you create one, Modelplane composes no collectors: nothing here stores anything, so collecting with nowhere to send it would spend GPU-cluster memory on samples nobody reads. Creating a destination turns collection on everywhere at once, and there's no per-deployment @@ -144,14 +161,25 @@ kind: MetricMapping metadata: name: my-engine spec: - statements: - - set(name, "modelplane_requests_waiting") - where name == "my_engine_queued_requests" + metrics: + - from: my_engine_queued_requests + to: modelplane_requests_waiting + - from: my_engine_kv_transfer_ms + fromUnit: Milliseconds + to: modelplane_request_kv_transfer_seconds ``` -`spec.statements` are OTTL, the collector's own transform language. Modelplane renders them -into every cluster's collector, so you write them once. An engine with no statements is -still collected, under its own names. +Modelplane renders every mapping into every cluster's collector, so you write one once. +`from` is the name your engine emits and `to` is what Modelplane calls it. + +Say `fromUnit` whenever the engine measures in something other than the unit the name +claims, and Modelplane converts to the base one. Skipping it is the expensive mistake here: +a series named `_seconds` that holds milliseconds reads a thousand times fast, and nothing +downstream can tell. + +Rename only where the measurements agree. Two engines' histograms under one name are worth +less than nothing if their buckets disagree, because a quantile over them is wrong rather +than approximate. One engine needs a flag. SGLang publishes `/metrics` only when it runs with `--enable-metrics`, so add it to the engine args. vLLM needs nothing. @@ -169,17 +197,19 @@ token for both. ## Migrating from a hand-written `PodMonitor` -[Collecting engine metrics]({{< ref "guides/collecting-engine-metrics.md" >}}) had you write -a `PodMonitor` and reach the in-cluster Prometheus over a `port-forward`. Both are gone, and -this page replaces that one. Three steps, and two of them fail quietly if you skip them. +Modelplane used to have you write a `PodMonitor` and reach an in-cluster Prometheus over a +`port-forward`. Both are gone. Three steps to move across, and two of them fail quietly if +you skip them. **Keep your Prometheus, and point a destination at it.** Collection becomes a push, so your store stops scraping and starts receiving. Same Prometheus, same retention, same Grafana: ```yaml spec: - exporters: - prometheusremotewrite: + sinks: + - name: prometheus + type: prometheusremotewrite + config: endpoint: http://prometheus.monitoring.svc:9090/api/v1/write ``` @@ -239,6 +269,6 @@ deleted once your panels use the new ones. A series no statement renames doesn't leave the cluster. If a panel needs an engine's own name, write a `MetricMapping` that renames it onto the `modelplane_*` surface: a -mapping for an engine Modelplane already knows adds to the built-in statements rather +mapping for an engine Modelplane already knows adds to the built-in renames rather than replacing them. diff --git a/docs/content/recipes/qwen2.5-72b.md b/docs/content/recipes/qwen2.5-72b.md index ca0fa63c8..e9bee702f 100644 --- a/docs/content/recipes/qwen2.5-72b.md +++ b/docs/content/recipes/qwen2.5-72b.md @@ -69,7 +69,7 @@ under the same model name: Both GPUs now serve the same workload, so their engine metrics give a direct performance comparison: scrape each replica's latency and throughput as in -[Collecting engine metrics]({{< ref "/guides/collecting-engine-metrics.md" >}}) and +[Monitor the Fleet]({{< ref "/platform/telemetry.md" >}}) and read the two side by side. Weights are relative, so once one platform wins, shift the 50/50 toward it - 80/20, and as far as 100/0 - without touching the deployment. diff --git a/docs/data/apigroups.yaml b/docs/data/apigroups.yaml index 9687c9303..fc508a746 100644 --- a/docs/data/apigroups.yaml +++ b/docs/data/apigroups.yaml @@ -31,6 +31,8 @@ concepts: InferenceGateway: /platform/inference-gateway InferenceClass: /platform/inference-class InferenceCluster: /platform/inference-cluster + MetricMapping: /platform/telemetry + TelemetryDestination: /platform/telemetry ModelDeployment: /models/model-deployment ModelService: /models/model-service ModelCache: /models/model-cache diff --git a/functions/compose-metric-mapping/function/fn.py b/functions/compose-metric-mapping/function/fn.py index 7729fc57d..66940dac3 100644 --- a/functions/compose-metric-mapping/function/fn.py +++ b/functions/compose-metric-mapping/function/fn.py @@ -14,10 +14,9 @@ """Compose a MetricMapping. -A MetricMapping carries OTTL statements that the collector on every -inference cluster renders into its transform processor. Modelplane does not -interpret them: what an operator writes here is the collector's own -configuration language. +A MetricMapping names the metrics one component emits and what Modelplane +calls them. compose-serving-stack compiles every mapping into the transform +processor of the collector on each inference cluster. This function composes nothing. The collector is composed by compose-serving-stack, which reads every MetricMapping. What this function @@ -35,7 +34,6 @@ CONDITION_REASON_AVAILABLE = "Available" CONDITION_REASON_WAITING_FOR_CLUSTERS = "WaitingForClusters" CONDITION_REASON_NO_CLUSTERS = "NoClusters" -CONDITION_REASON_NO_STATEMENTS = "NoStatements" class FunctionRunner(grpcv1.FunctionRunnerServiceServicer): @@ -56,7 +54,7 @@ async def RunFunction( xr = v1alpha1.MetricMapping(**resource.struct_to_dict(req.observed.composite.resource)) # Every inference cluster renders every mapping, so the count of - # clusters is the count that took these statements. + # clusters is the count that took these renames. response.require_resources( rsp, name="clusters", @@ -70,12 +68,8 @@ async def RunFunction( clusters = len(list(request.get_required_resources(req, "clusters"))) resource.update_status(rsp.desired.composite, v1alpha1.Status(clusters=clusters)) - if not xr.spec.statements: - _not_ready(rsp, CONDITION_REASON_NO_STATEMENTS, "No statements, so this mapping changes nothing") - return rsp - if clusters == 0: - _not_ready(rsp, CONDITION_REASON_NO_CLUSTERS, "No inference cluster to render these statements into") + _not_ready(rsp, CONDITION_REASON_NO_CLUSTERS, "No inference cluster to render these renames into") return rsp response.set_conditions( @@ -84,7 +78,7 @@ async def RunFunction( typ=CONDITION_TYPE_ACCEPTED, status="True", reason=CONDITION_REASON_AVAILABLE, - message=f"Rendered into {clusters} inference cluster(s)", + message=f"Renaming {len(xr.spec.metrics)} metric(s) on {clusters} inference cluster(s)", ), ) rsp.desired.composite.ready = fnv1.READY_TRUE diff --git a/functions/compose-metric-mapping/tests/test_fn.py b/functions/compose-metric-mapping/tests/test_fn.py index 6afc75fdd..5399e5dd1 100644 --- a/functions/compose-metric-mapping/tests/test_fn.py +++ b/functions/compose-metric-mapping/tests/test_fn.py @@ -52,7 +52,7 @@ async def test_compose(self) -> None: "kind": "MetricMapping", "metadata": {"name": "my-engine"}, "spec": { - "statements": ['set(name, "modelplane_requests_waiting") where name == "my_engine_queued"'], + "metrics": [{"from": "my_engine_queued", "to": "modelplane_requests_waiting"}], }, } cluster = resource.dict_to_struct( @@ -83,11 +83,9 @@ def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.R ), ) - no_statements = {**mapping, "spec": {}} - cases = [ Case( - name="ready, and says how many clusters took the statements", + name="ready, and says how many clusters took the renames", req=req(mapping, [cluster, cluster]), want=want( fnv1.READY_TRUE, @@ -96,7 +94,7 @@ def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.R type="Accepted", status=fnv1.STATUS_CONDITION_TRUE, reason="Available", - message="Rendered into 2 inference cluster(s)", + message="Renaming 1 metric(s) on 2 inference cluster(s)", ), ), ), @@ -110,21 +108,7 @@ def want(ready: fnv1.Ready, status: dict | None, cond: fnv1.Condition) -> fnv1.R type="Accepted", status=fnv1.STATUS_CONDITION_FALSE, reason="NoClusters", - message="No inference cluster to render these statements into", - ), - ), - ), - Case( - name="not ready when the mapping would change nothing", - req=req(no_statements, [cluster]), - want=want( - fnv1.READY_FALSE, - {"status": {"clusters": 1}}, - fnv1.Condition( - type="Accepted", - status=fnv1.STATUS_CONDITION_FALSE, - reason="NoStatements", - message="No statements, so this mapping changes nothing", + message="No inference cluster to render these renames into", ), ), ), diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 2842e3d38..6b3bacad2 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -29,8 +29,8 @@ from typing import Any import yaml - -from function.stacks import metrics as stacks_metrics +from models.ai.modelplane.metricmapping import v1alpha1 as mmv1alpha1 +from models.ai.modelplane.telemetrydestination import v1alpha1 as tdv1alpha1 NAMESPACE = "modelplane-system" NAME = "modelplane-collector" @@ -115,39 +115,79 @@ def _scrape_configs() -> list[dict[str, Any]]: ] -def _transform(statements: list[str]) -> dict[str, Any]: - """Value rewrites first, then every rename. +# What each source unit is worth in the base unit the target name claims. +# Written as the expression rather than a factor so nothing has to render a +# float: 1e-09 is not an OTTL literal. +_UNIT_CONVERSION = { + "Millijoules": "value_double / 1000", + "Milliseconds": "value_double / 1000", + "Nanoseconds": "value_double / 1000000000", + "Mebibytes": "value_double * 1048576", +} + + +def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], list[str]]: + """Compile the mappings to OTTL, as (datapoint, metric) statements. + + OTTL is rendered here rather than written in a MetricMapping so the kind + stays a description of what a component emits, and the collector's own + configuration language stays Modelplane's problem. It is also the only + place that knows a unit conversion has to run somewhere a rename cannot. + """ + datapoint: list[str] = [] + metric: list[str] = [] + for mapping in mappings: + for m in mapping.spec.metrics: + if m.fromUnit: + conversion = _UNIT_CONVERSION[m.fromUnit] + datapoint.append(f'set(value_double, {conversion}) where metric.name == "{m.from_}"') + metric.append(f'set(name, "{m.to}") where name == "{m.from_}"') + return datapoint, metric + + +def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: + """Unit conversions first, then every rename. Two blocks rather than one list: a statement reaching a datapoint's value can't run in the metric context, and the processor finishes a block over every datapoint before it starts the next, which is what keeps a rename - from stranding the datapoints a value rewrite hasn't reached yet. + from stranding the datapoints a conversion hasn't reached yet. """ - return { - "metric_statements": [ - {"context": "datapoint", "statements": list(stacks_metrics.DATAPOINT_STATEMENTS)}, - {"context": "metric", "statements": statements}, - ] - } + datapoint, metric = statements(mappings) + blocks = [{"context": "metric", "statements": metric}] + if datapoint: + blocks.insert(0, {"context": "datapoint", "statements": datapoint}) + return {"metric_statements": blocks} + + +def exporters(sinks: list[tdv1alpha1.Sink]) -> dict[str, Any]: + """The sinks as the collector's exporters block. + + A sink renders under `/`, which is how the collector names a + second instance of one component, so two sinks of the same type don't + collide. + """ + return {f"{sink.type}/{sink.name}": dict(sink.config) for sink in sinks} def config( cluster: str, - statements: list[str], - exporters: dict[str, Any], + mappings: list[mmv1alpha1.MetricMapping], + sinks: list[tdv1alpha1.Sink], extensions: dict[str, Any], ) -> str: """The collector's configuration, as YAML. - Only modelplane_* leaves the cluster: a series the statements did not rename - is one whose meaning Modelplane cannot vouch for across engines, and it - costs the same to carry as one that was renamed. + Only modelplane_* leaves the cluster: a series no mapping renamed is one + whose meaning Modelplane cannot vouch for across engines, and it costs the + same to carry as one that was renamed. """ + sink_exporters = exporters(sinks) processors: dict[str, Any] = { # cluster is stamped here rather than downstream: one receiver on the # control plane sees a merged stream and cannot tell senders apart. "resource/cluster": {"attributes": [{"key": "cluster", "value": cluster, "action": "upsert"}]}, - "transform/modelplane": _transform(statements), + "transform/modelplane": _transform(mappings), # A pod's identity is a resource attribute, where a metric processor # cannot reach it. Strip and merge the resources first, or the # aggregation below combines nothing. @@ -164,12 +204,18 @@ def config( ] service: dict[str, Any] = { - "pipelines": {"metrics": {"receivers": ["prometheus"], "processors": pipeline, "exporters": sorted(exporters)}}, + "pipelines": { + "metrics": { + "receivers": ["prometheus"], + "processors": pipeline, + "exporters": sorted(sink_exporters), + } + }, } cfg: dict[str, Any] = { "receivers": {"prometheus": {"config": {"scrape_configs": _scrape_configs()}}}, "processors": processors, - "exporters": exporters, + "exporters": sink_exporters, "service": service, } if extensions: @@ -187,10 +233,9 @@ def _digest(rendered: str) -> str: def objects( cluster: str, - statements: list[str], - exporters: dict[str, Any], + mappings: list[mmv1alpha1.MetricMapping], + sinks: list[tdv1alpha1.Sink], extensions: dict[str, Any], - secret_name: str | None, ) -> list[tuple[str, dict[str, Any], str | None]]: """The collector as (key, manifest, readiness CEL) triples. @@ -204,18 +249,22 @@ def objects( permissions on clusters running no collector. """ labels = {"app.kubernetes.io/name": NAME, "app.kubernetes.io/managed-by": "modelplane"} - rendered = config(cluster, statements, exporters, extensions) + rendered = config(cluster, mappings, sinks, extensions) volumes: list[dict[str, Any]] = [{"name": "config", "configMap": {"name": NAME}}] mounts: list[dict[str, Any]] = [{"name": "config", "mountPath": "/conf"}] env_from: list[dict[str, Any]] = [] - if secret_name: + for sink in sinks: + if not sink.secretRef: + continue # Mounted both ways. An environment variable is fixed for the life of a # process, so a rotated credential would need a restart to be read; a # mounted file is refreshed in place and an authenticator reading one - # picks the new credential up without one. - volumes.append({"name": "credentials", "secret": {"secretName": secret_name}}) - mounts.append({"name": "credentials", "mountPath": "/etc/modelplane/telemetry", "readOnly": True}) - env_from.append({"secretRef": {"name": secret_name}}) + # picks the new credential up without one. The file is under the sink's + # own directory, so two sinks can both hold a key called `token`. + volume = f"credentials-{sink.name}" + volumes.append({"name": volume, "secret": {"secretName": sink.secretRef.name}}) + mounts.append({"name": volume, "mountPath": f"/etc/modelplane/telemetry/{sink.name}", "readOnly": True}) + env_from.append({"secretRef": {"name": sink.secretRef.name}}) return [ ( diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index ca4b8454d..f9dbe1759 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -791,17 +791,15 @@ def compose_collector(self) -> list[str]: mappings += [ mmv1alpha1.MetricMapping.model_validate(m) for m in request.get_required_resources(self.req, "mappings") ] - statements: list[str] = [str(st.root) for mp in mappings for st in mp.spec.statements or []] pc_observed = self.provider_configs_observed() pc = _pc_name(self.xr) rendered: list[str] = [] for key, manifest, cel in collector.objects( cluster=_name(self.xr.metadata), - statements=statements, - exporters=dict(dest.spec.exporters or {}), + mappings=mappings, + sinks=list(dest.spec.sinks), extensions=dict(dest.spec.extensions or {}), - secret_name=dest.spec.secretRef.name if dest.spec.secretRef else None, ): if not (pc_observed or key in self.req.observed.resources): continue diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index 8bb49409b..237236fc7 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -74,7 +74,9 @@ "llm_d_epp_scheduler_e2e_duration_seconds": "modelplane_route_decision_seconds", } -# The GPUs, through whichever vendor's exporter the stack installed. +# The GPUs, through whichever vendor's exporter the stack installed. DCGM +# reports framebuffer memory in MiB and energy in millijoules, and both are +# renamed onto a name that claims a different unit, so both say so. _GPU = { "DCGM_FI_DEV_FB_USED": "modelplane_gpu_memory_used_bytes", "DCGM_FI_PROF_GR_ENGINE_ACTIVE": "modelplane_gpu_compute_active_ratio", @@ -82,53 +84,48 @@ "DCGM_FI_PROF_DRAM_ACTIVE": "modelplane_gpu_memory_bandwidth_ratio", "DCGM_FI_DEV_GPU_TEMP": "modelplane_gpu_temperature_celsius", "DCGM_FI_DEV_POWER_USAGE": "modelplane_gpu_power_watts", + "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION": "modelplane_energy_joules_total", } - -def _rename(pairs: dict[str, str]) -> list[str]: - return [f'set(name, "{new}") where name == "{old}"' for old, new in pairs.items()] +_GPU_UNITS = { + "DCGM_FI_DEV_FB_USED": "Mebibytes", + "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION": "Millijoules", +} -def _mapping(name: str, statements: list[str]) -> v1alpha1.MetricMapping: +def _mapping( + name: str, + pairs: dict[str, str], + units: dict[str, str] | None = None, +) -> v1alpha1.MetricMapping: + units = units or {} return v1alpha1.MetricMapping( metadata=metav1.ObjectMeta(name=name), - spec=v1alpha1.Spec(statements=[v1alpha1.Statement(st) for st in statements]), + spec=v1alpha1.Spec( + metrics=[ + v1alpha1.Metric.model_validate( + {"from": src, "to": dst} | ({"fromUnit": units[src]} if src in units else {}) + ) + for src, dst in pairs.items() + ] + ), ) def mappings() -> list[v1alpha1.MetricMapping]: """The mappings Modelplane provides, as the kind an operator would write. - Built as MetricMappings rather than as a bare list of statements so the - collector renders Modelplane's own renames through the same path as an - operator's, and a built-in that breaks breaks the path everyone uses. + Built as MetricMappings rather than as collector configuration so that + Modelplane's own renames reach the collector down the same path an + operator's do, and a built-in that breaks breaks the path everyone uses. """ return [ - _mapping("modelplane-gateway", _rename(_GATEWAY)), - _mapping("modelplane-vllm", _rename(_VLLM)), - _mapping("modelplane-sglang", _rename(_SGLANG)), - _mapping("modelplane-picker", _rename(_PICKER)), - _mapping( - "modelplane-gpu", - _rename(_GPU) + _rename({"DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION": "modelplane_energy_joules_total"}), - ), + _mapping("modelplane-gateway", _GATEWAY), + _mapping("modelplane-vllm", _VLLM), + _mapping("modelplane-sglang", _SGLANG), + _mapping("modelplane-picker", _PICKER), + _mapping("modelplane-gpu", _GPU, _GPU_UNITS), ] BUILTIN_MAPPINGS = mappings() - -# A datapoint's value is out of reach of the metric context, so the one -# statement that rewrites a value rather than a name runs in its own block -# ahead of the renames. It has to be a block of its own rather than an earlier -# line: the transform processor finishes a block over every datapoint before -# starting the next, and a rename landing first would leave every datapoint -# after the first unmatched and unscaled. -# -# DCGM reports energy in millijoules and framebuffer memory in MiB, and both -# are renamed onto a name that states a different unit. Left to a query -# instead, a name ending in _joules_total holding millijoules is the kind of -# thing nobody notices until a bill. -DATAPOINT_STATEMENTS = [ - 'set(value_double, value_double / 1000) where metric.name == "DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION"', - 'set(value_double, value_double * 1048576) where metric.name == "DCGM_FI_DEV_FB_USED"', -] diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 2f8b85cdb..a7f670f49 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -18,21 +18,35 @@ import yaml from function import collector, stacks +from models.ai.modelplane.telemetrydestination import v1alpha1 as tdv1alpha1 -_EXPORTERS = {"otlphttp": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}}} -_EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/token"}} +_EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/primary/token"}} -def _statements() -> list[str]: - return [str(st.root) for mp in stacks.BUILTIN_MAPPINGS for st in mp.spec.statements or []] +def _sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = None) -> tdv1alpha1.Sink: + return tdv1alpha1.Sink.model_validate( + { + "name": name, + "type": type_, + "config": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}}, + **({"secretRef": {"name": secret}} if secret else {}), + } + ) + + +_SINKS = [_sink()] + + +def _metric_statements() -> list[str]: + return collector.statements(list(stacks.BUILTIN_MAPPINGS))[1] def _config(*, extensions: dict | None = None) -> dict: return yaml.safe_load( collector.config( "prod-us-east", - _statements(), - _EXPORTERS, + list(stacks.BUILTIN_MAPPINGS), + _SINKS, _EXTENSIONS if extensions is None else extensions, ) ) @@ -105,7 +119,7 @@ def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: def test_sglang_latency_histograms_are_not_renamed(self) -> None: """Their buckets resolve to 100ms where vLLM's resolve to 1ms.""" - joined = " ".join(_statements()) + joined = " ".join(_metric_statements()) self.assertNotIn("sglang:time_to_first_token_seconds", joined) self.assertNotIn("sglang:inter_token_latency", joined) @@ -114,7 +128,10 @@ class TestObjects(unittest.TestCase): """The manifests this composes.""" def _objects(self, secret: str | None = None) -> dict: - return {k: m for k, m, _ in collector.objects("prod-us-east", _statements(), _EXPORTERS, _EXTENSIONS, secret)} + sinks = [_sink(secret=secret)] if secret else _SINKS + return { + k: m for k, m, _ in collector.objects("prod-us-east", list(stacks.BUILTIN_MAPPINGS), sinks, _EXTENSIONS) + } def test_config_hash_is_stable_across_processes(self) -> None: """hash() is seeded per process, so it would redeploy on every reconcile.""" @@ -126,9 +143,27 @@ def test_config_hash_is_stable_across_processes(self) -> None: def test_credentials_mount_as_a_file_and_an_environment_variable(self) -> None: """A rotated token in an environment variable needs a restart to be read.""" pod = self._objects(secret="telemetry-credentials")["collector"]["spec"]["template"]["spec"] - self.assertIn("credentials", [v["name"] for v in pod["volumes"]]) + self.assertIn("credentials-primary", [v["name"] for v in pod["volumes"]]) self.assertEqual(pod["containers"][0]["envFrom"], [{"secretRef": {"name": "telemetry-credentials"}}]) + def test_each_sink_gets_its_own_credential_directory(self) -> None: + """Two sinks can both hold a key called token, and neither reads the other's.""" + sinks = [ + _sink(name="vendor", secret="vendor-token"), + _sink(name="prometheus", type_="prometheusremotewrite", secret="prom-token"), + ] + pod = { + k: m for k, m, _ in collector.objects("prod-us-east", list(stacks.BUILTIN_MAPPINGS), sinks, _EXTENSIONS) + }["collector"]["spec"]["template"]["spec"] + mounts = {m["name"]: m["mountPath"] for m in pod["containers"][0]["volumeMounts"]} + self.assertEqual(mounts["credentials-vendor"], "/etc/modelplane/telemetry/vendor") + self.assertEqual(mounts["credentials-prometheus"], "/etc/modelplane/telemetry/prometheus") + + def test_two_sinks_of_one_type_do_not_collide(self) -> None: + """The collector names a second instance of a component /.""" + rendered = collector.exporters([_sink(name="a"), _sink(name="b")]) + self.assertEqual(sorted(rendered), ["otlphttp/a", "otlphttp/b"]) + def test_no_secret_mounts_nothing(self) -> None: pod = self._objects()["collector"]["spec"]["template"]["spec"] self.assertEqual([v["name"] for v in pod["volumes"]], ["config"]) diff --git a/functions/compose-telemetry-destination/function/fn.py b/functions/compose-telemetry-destination/function/fn.py index 309a7b912..58e667c52 100644 --- a/functions/compose-telemetry-destination/function/fn.py +++ b/functions/compose-telemetry-destination/function/fn.py @@ -14,12 +14,13 @@ """Compose a TelemetryDestination. -A TelemetryDestination carries the collector's exporters and extensions -verbatim, and compose-serving-stack renders them into the collector it -composes. Modelplane does not model what an exporter is, so there is little -here to validate and the little there is matters: an exporter naming an -authenticator that no extension defines makes a collector refuse to start, -and that failure surfaces as telemetry silently never arriving. +A TelemetryDestination names the sinks the fleet's metrics go to, each +carrying an exporter's own configuration verbatim, and compose-serving-stack +renders them into the collector it composes. Modelplane does not model what an +exporter is, so there is little here to validate and the little there is +matters: a sink naming an authenticator that no extension defines makes a +collector refuse to start, and that failure surfaces as telemetry silently +never arriving. """ import grpc @@ -30,12 +31,11 @@ CONDITION_TYPE_ACCEPTED = "Accepted" CONDITION_REASON_AVAILABLE = "Available" -CONDITION_REASON_NO_EXPORTERS = "NoExporters" CONDITION_REASON_UNKNOWN_AUTHENTICATOR = "UnknownAuthenticator" CONDITION_REASON_WAITING_FOR_SECRET = "WaitingForSecret" CONDITION_REASON_SECRET_NOT_FOUND = "SecretNotFound" -_SECRET_KEY = "secret" +_SECRET_PREFIX = "secret-" class FunctionRunner(grpcv1.FunctionRunnerServiceServicer): @@ -55,17 +55,14 @@ async def RunFunction( rsp = response.to(req) xr = v1alpha1.TelemetryDestination(**resource.struct_to_dict(req.observed.composite.resource)) - exporters = xr.spec.exporters or {} - if not exporters: - _not_ready(rsp, CONDITION_REASON_NO_EXPORTERS, "No exporters, so collected telemetry has nowhere to go") - return rsp + sinks = list(xr.spec.sinks) - # An exporter's auth block names an authenticator by extension name. The + # A sink's auth block names an authenticator by extension name. The # collector refuses to start when it names one no extension defines, and # a collector that never starts looks exactly like a fleet that produces # nothing, so it is worth catching on the object instead. extensions = set((xr.spec.extensions or {}).keys()) - missing = sorted(_authenticators(exporters) - extensions) + missing = sorted(_authenticators(sinks) - extensions) if missing: _not_ready( rsp, @@ -74,22 +71,25 @@ async def RunFunction( ) return rsp - if xr.spec.secretRef is not None: + for sink in sinks: + if sink.secretRef is None: + continue + key = f"{_SECRET_PREFIX}{sink.name}" response.require_resources( rsp, - name=_SECRET_KEY, + name=key, api_version="v1", kind="Secret", - match_name=xr.spec.secretRef.name, + match_name=sink.secretRef.name, ) - if _SECRET_KEY not in req.required_resources: + if key not in req.required_resources: _not_ready(rsp, CONDITION_REASON_WAITING_FOR_SECRET, "Waiting for the credential Secret to resolve") return rsp - if not list(request.get_required_resources(req, _SECRET_KEY)): + if not list(request.get_required_resources(req, key)): _not_ready( rsp, CONDITION_REASON_SECRET_NOT_FOUND, - f"Secret {xr.spec.secretRef.name} does not exist, so the collector has no credential to send with", + f"Secret {sink.secretRef.name} does not exist, so sink {sink.name} has no credential to send with", ) return rsp @@ -100,20 +100,18 @@ async def RunFunction( typ=CONDITION_TYPE_ACCEPTED, status="True", reason=CONDITION_REASON_AVAILABLE, - message=f"Exporting through {', '.join(sorted(exporters))}", + message=f"Exporting through {', '.join(f'{s.type}/{s.name}' for s in sinks)}", ), ) rsp.desired.composite.ready = fnv1.READY_TRUE return rsp -def _authenticators(exporters: dict) -> set[str]: - """Every authenticator an exporter references, by extension name.""" +def _authenticators(sinks: list[v1alpha1.Sink]) -> set[str]: + """Every authenticator a sink references, by extension name.""" names: set[str] = set() - for cfg in exporters.values(): - if not isinstance(cfg, dict): - continue - auth = cfg.get("auth") + for sink in sinks: + auth = sink.config.get("auth") if isinstance(auth, dict) and isinstance(auth.get("authenticator"), str): names.add(auth["authenticator"]) return names diff --git a/functions/compose-telemetry-destination/tests/test_fn.py b/functions/compose-telemetry-destination/tests/test_fn.py index d5710ba63..e15b2b0b0 100644 --- a/functions/compose-telemetry-destination/tests/test_fn.py +++ b/functions/compose-telemetry-destination/tests/test_fn.py @@ -47,9 +47,16 @@ def setUpClass(cls) -> None: async def test_compose(self) -> None: """The function reports whether a destination can actually be sent through.""" - exporters = { - "otlphttp": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}} - } + + def sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = None) -> dict: + return { + "name": name, + "type": type_, + "config": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}}, + **({"secretRef": {"name": secret}} if secret else {}), + } + + sinks = [sink()] extensions = {"bearertokenauth": {"token": "${env:OTLP_TOKEN}"}} def xr(spec: dict) -> dict: @@ -65,7 +72,7 @@ def req(spec: dict, secrets: list | None = None) -> fnv1.RunFunctionRequest: observed=fnv1.State(composite=fnv1.Resource(resource=resource.dict_to_struct(xr(spec)))), ) if secrets is not None: - r.required_resources["secret"].items.extend([fnv1.Resource(resource=s) for s in secrets]) + r.required_resources["secret-primary"].items.extend([fnv1.Resource(resource=s) for s in secrets]) return r def want( @@ -81,15 +88,15 @@ def want( context=structpb.Struct(), ) if secret is not None: - rsp.requirements.resources["secret"].api_version = "v1" - rsp.requirements.resources["secret"].kind = "Secret" - rsp.requirements.resources["secret"].match_name = secret + rsp.requirements.resources["secret-primary"].api_version = "v1" + rsp.requirements.resources["secret-primary"].kind = "Secret" + rsp.requirements.resources["secret-primary"].match_name = secret return rsp cases = [ Case( - name="ready, naming the exporters it sends through", - req=req({"exporters": exporters, "extensions": extensions}), + name="ready, naming the sinks it sends through", + req=req({"sinks": sinks, "extensions": extensions}), want=want( fnv1.READY_TRUE, {"status": {}}, @@ -97,14 +104,22 @@ def want( type="Accepted", status=fnv1.STATUS_CONDITION_TRUE, reason="Available", - message="Exporting through otlphttp", + message="Exporting through otlphttp/primary", ), ), ), Case( name="ready with an exporter that references no authenticator at all", req=req( - {"exporters": {"prometheusremotewrite": {"endpoint": "https://prom.acme.example/api/v1/write"}}} + { + "sinks": [ + { + "name": "prom", + "type": "prometheusremotewrite", + "config": {"endpoint": "https://prom.acme.example/api/v1/write"}, + } + ] + } ), want=want( fnv1.READY_TRUE, @@ -113,14 +128,14 @@ def want( type="Accepted", status=fnv1.STATUS_CONDITION_TRUE, reason="Available", - message="Exporting through prometheusremotewrite", + message="Exporting through prometheusremotewrite/prom", ), ), ), Case( name="ready once the credential Secret exists", req=req( - {"exporters": exporters, "extensions": extensions, "secretRef": {"name": "telemetry-credentials"}}, + {"sinks": [sink(secret="telemetry-credentials")], "extensions": extensions}, secrets=[ resource.dict_to_struct( {"apiVersion": "v1", "kind": "Secret", "metadata": {"name": "telemetry-credentials"}} @@ -134,7 +149,7 @@ def want( type="Accepted", status=fnv1.STATUS_CONDITION_TRUE, reason="Available", - message="Exporting through otlphttp", + message="Exporting through otlphttp/primary", ), secret="telemetry-credentials", ), @@ -142,7 +157,7 @@ def want( Case( name="waits for the credential Secret to resolve", req=req( - {"exporters": exporters, "extensions": extensions, "secretRef": {"name": "telemetry-credentials"}}, + {"sinks": [sink(secret="telemetry-credentials")], "extensions": extensions}, ), want=want( fnv1.READY_FALSE, @@ -157,8 +172,8 @@ def want( ), ), Case( - name="not ready when an exporter names an authenticator nothing defines", - req=req({"exporters": exporters}), + name="not ready when a sink names an authenticator nothing defines", + req=req({"sinks": sinks}), want=want( fnv1.READY_FALSE, None, @@ -170,38 +185,10 @@ def want( ), ), ), - Case( - name="an exporter that is not a mapping is left to the collector to reject", - req=req({"exporters": {"otlphttp": "https://otel.acme.example"}}), - want=want( - fnv1.READY_TRUE, - {"status": {}}, - fnv1.Condition( - type="Accepted", - status=fnv1.STATUS_CONDITION_TRUE, - reason="Available", - message="Exporting through otlphttp", - ), - ), - ), - Case( - name="not ready with no exporters at all", - req=req({"exporters": {}}), - want=want( - fnv1.READY_FALSE, - None, - fnv1.Condition( - type="Accepted", - status=fnv1.STATUS_CONDITION_FALSE, - reason="NoExporters", - message="No exporters, so collected telemetry has nowhere to go", - ), - ), - ), Case( name="not ready when the credential Secret is missing", req=req( - {"exporters": exporters, "extensions": extensions, "secretRef": {"name": "telemetry-credentials"}}, + {"sinks": [sink(secret="telemetry-credentials")], "extensions": extensions}, secrets=[], ), want=want( @@ -213,7 +200,7 @@ def want( reason="SecretNotFound", message=( "Secret telemetry-credentials does not exist, " - "so the collector has no credential to send with" + "so sink primary has no credential to send with" ), ), secret="telemetry-credentials", diff --git a/schemas/.lock.json b/schemas/.lock.json index 5b2f5d9c2..21a992e30 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "e4bb7953a7e73cb25d317f8beb70973a11cfbdd2bf5b75e570292e94bc9087a7", + "fs://apis": "80357e7a8407a4046262d4c3c0e13dfd4f5da938b0e1368321439319842b46bd", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index c85ad7d88..b1352fc7a 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -5,7 +5,7 @@ from typing import Literal -from pydantic import AwareDatetime, BaseModel, Field, RootModel, constr +from pydantic import AwareDatetime, BaseModel, Field, constr from ....io.k8s.apimachinery.pkg.apis.meta import v1 @@ -42,8 +42,22 @@ class Crossplane(BaseModel): resourceRefs: list[ResourceRef] | None = None -class Statement(RootModel[constr(max_length=2048)]): - root: constr(max_length=2048) +class Metric(BaseModel): + from_: constr(max_length=255) = Field(..., alias='from') + """ + The metric's name as the component emits it, matched exactly. Nothing here declares which engine a deployment runs: a name that no component emits simply matches nothing. + """ + fromUnit: ( + Literal['Millijoules', 'Mebibytes', 'Milliseconds', 'Nanoseconds'] | None + ) = None + """ + What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, and mebibytes multiplied out to bytes. + Say it whenever the source disagrees with the target, even where the factor looks obvious. A name ending in _bytes that holds mebibytes is the kind of thing nobody notices until a capacity review, and stating the source unit is what makes the conversion happen at all. + """ + to: constr(pattern=r'^modelplane_[a-z0-9_]*[a-z0-9]$', max_length=255) + """ + What Modelplane calls it. Only modelplane_* leaves a cluster, so a metric with no name here is one nobody downstream can read. + """ class Spec(BaseModel): @@ -51,10 +65,10 @@ class Spec(BaseModel): """ Configures how Crossplane will reconcile this composite resource """ - statements: list[Statement] | None = Field(None, max_length=128) + metrics: list[Metric] = Field(..., max_length=128, min_length=1) """ - OTTL statements, rendered into the collector's transform processor beside Modelplane's own. Modelplane does not interpret them: what you write here is the collector's own configuration language, documented by OpenTelemetry, and it is the same thing Modelplane writes for vLLM. - Statements select through their own where clauses, so nothing declares which engine a deployment runs. + The metrics this component emits, and what Modelplane calls them. + Rename only where the measurements agree. Two engines' histograms sharing a name are worth less than nothing if their buckets disagree, because a quantile across them is wrong rather than approximate. """ @@ -94,7 +108,7 @@ class MetricMapping(BaseModel): spec: Spec """ How one component's metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each inference cluster's collector, so a mapping is written once on the control plane and reaches the whole fleet. - A mapping naming a component Modelplane already provides statements for is additive: its statements run after the built-in ones. + A mapping naming a component Modelplane already provides renames for is additive: its renames run after the built-in ones, and a later rename of the same metric wins. """ status: Status | None = None diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index 5bad09805..1ee121e19 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -5,7 +5,7 @@ from typing import Any, Literal -from pydantic import AwareDatetime, BaseModel, constr +from pydantic import AwareDatetime, BaseModel, Field, constr from ....io.k8s.apimachinery.pkg.apis.meta import v1 @@ -49,23 +49,39 @@ class SecretRef(BaseModel): """ +class Sink(BaseModel): + config: dict[str, Any] + """ + That exporter's own configuration, passed through unread. Modelplane does not model what an exporter is, so TLS, retry, queue and compression settings all work, and a sink keeps working when the collector gains a setting Modelplane has never heard of. + """ + name: constr(pattern=r'^[a-z0-9]([-a-z0-9]*[a-z0-9])?$', max_length=63) + """ + This sink's name, unique within the destination. It names the collector's exporter instance and the directory its credential mounts at, so renaming one restarts the collector. + """ + secretRef: SecretRef | None = None + """ + A Secret holding this sink's credential. Its keys reach the collector as environment variables, for config above referring to ${env:TOKEN}, and as files under /etc/modelplane/telemetry//, for an authenticator reading one from disk. + Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. A file is refreshed in place where an environment variable is fixed for the life of the process, so an authenticator reading the file picks up a rotated credential without a restart. + """ + type: constr(max_length=63) + """ + The collector exporter to send with, by the name OpenTelemetry gives it: otlphttp, otlp, prometheusremotewrite, kafka, and every other one the collector provides. + Not an enum, because enumerating them here would mean a Modelplane release for each exporter the collector gains, and the collector already refuses to start on a name it doesn't have. + """ + + class Spec(BaseModel): crossplane: Crossplane | None = None """ Configures how Crossplane will reconcile this composite resource """ - exporters: dict[str, Any] - """ - The OpenTelemetry collector's exporters block, passed through unread. Modelplane validates that it parses and reports whether the destination accepts writes; it does not model what an exporter is. - So any exporter the collector provides works, with its TLS, retry and queue settings intact, and a destination keeps working when the collector gains an exporter Modelplane has never heard of. - """ extensions: dict[str, Any] | None = None """ - The collector's extensions block, passed through unread, for the authenticator an exporter references. Bearer token, basic auth, OIDC and SigV4 all work, because none of them is modelled here. + The collector's extensions block, passed through unread, for the authenticator a sink references. Bearer token, basic auth, OIDC and SigV4 all work, because none of them is modelled here. """ - secretRef: SecretRef | None = None + sinks: list[Sink] = Field(..., max_length=16, min_length=1) """ - A Secret whose keys Modelplane mounts into the collector as environment variables, so configuration above refers to ${env:TOKEN} and the credential itself never appears in this object or in kubectl output. + Where to send it. Every sink gets the whole stream, so two sinks is two copies of the fleet's metrics, billed twice. """ From 82a582e47cfe703dcf9970af5633871406fa17ae Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 10:45:36 -0700 Subject: [PATCH 10/42] Hold a metric name to the characters a metric name can have `from` was interpolated into an OTTL comparison with nothing constraining it, so a quote closed the string early and the rest of the value became part of the query: set(name, "modelplane_pwned") where name == "x" or true or name == "y" That renames every series the collector sees. A typo does it as readily as anything deliberate, and the result is a fleet whose metrics all arrive under one wrong name. Hold the field to the characters a Prometheus or OpenTelemetry metric name can contain, which every built-in already satisfies. Two other things found reading it back: Which TelemetryDestination wins when several exist came from the API server's list order. The collector restarts on a change to its rendered config, so an unstable choice would redeploy it on alternate reconciles. Sort by name. The guide offered four metrics nothing composes - the saturation maximum, the token counters, and the two that need per-replica GPU state - and mapped three more in its migration table. A rename carries one metric to one name, so the counters that fold several series under a label aren't part of this surface yet; say that instead of promising them, and point at the histograms that do exist. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 7 +++++++ apis/telemetrydestinations/definition.yaml | 5 ++++- docs/content/platform/telemetry.md | 20 ++++++++++--------- .../config/vocabularies/Modelplane/accept.txt | 1 + .../compose-serving-stack/function/fn.py | 13 ++++++++++-- .../tests/test_collector.py | 18 +++++++++++++++++ schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 5 ++++- .../telemetrydestination/v1alpha1.py | 2 +- 9 files changed, 58 insertions(+), 15 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index 63258af48..b90ed310a 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -58,11 +58,18 @@ spec: from: type: string maxLength: 255 + pattern: '^[a-zA-Z_:][a-zA-Z0-9_:]*$' description: >- The metric's name as the component emits it, matched exactly. Nothing here declares which engine a deployment runs: a name that no component emits simply matches nothing. + + Held to the characters a metric name can contain. The + name is matched inside the collector's own query + language, so a quote here would end the comparison + early and rename whatever the rest of the line + matched. to: type: string maxLength: 255 diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index 7bf078b02..3f1cf09c9 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -92,7 +92,10 @@ spec: Per sink rather than per destination, so two sinks with different credentials don't have to share one - Secret and tell their keys apart by prefix. A file is + Secret and tell their keys apart by prefix. The files + are per sink; the environment variables are not, so + two Secrets sharing a key name still collide there and + the file is the one to read. A file is refreshed in place where an environment variable is fixed for the life of the process, so an authenticator reading the file picks up a rotated credential without diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 82aefbfbf..ba3ae06a7 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -30,10 +30,10 @@ Every series carries `cluster`. A series about a deployment also carries `deploy | `modelplane_request_queue_seconds` | How long a request waited before the engine started | | `modelplane_requests_waiting` | Queue depth per engine | | `modelplane_kv_cache_utilization_ratio` | KV-cache occupancy, averaged over replicas | -| `modelplane_kv_cache_utilization_ratio_max` | KV-cache occupancy of the busiest replica | -| `modelplane_tokens_total` | Tokens in and out, by `direction` | -| `modelplane_replica_gpus` | GPUs a replica holds | -| `modelplane_gpu_seconds_total` | GPU-time bound to serving | +| `modelplane_request_input_tokens` | Prompt size, as a histogram | +| `modelplane_request_output_tokens` | Generated length, as a histogram | +| `modelplane_gpu_memory_used_bytes` | Framebuffer memory in use, per GPU | +| `modelplane_energy_joules_total` | Energy drawn since the driver last reloaded | Latency appears twice on purpose. The `frontend_` series are what your caller experienced, measured at the gateway. The engine's own series are what the engine spent. When the @@ -235,17 +235,19 @@ cluster, and every series now carries `cluster`. | `vllm:kv_cache_usage_perc` | `modelplane_kv_cache_utilization_ratio` | | `vllm:num_preemptions_total` | `modelplane_requests_preempted_total` | | `vllm:prefix_cache_hits_total` | `modelplane_prefix_cache_hits_total` | -| `vllm:prompt_tokens_total` | `modelplane_tokens_total{direction="input"}` | -| `vllm:generation_tokens_total` | `modelplane_tokens_total{direction="output"}` | -| `vllm:request_success_total{finished_reason}` | `modelplane_responses_total{reason}` | | `DCGM_FI_DEV_FB_USED` | `modelplane_gpu_memory_used_bytes` | | `DCGM_FI_DEV_GPU_TEMP` | `modelplane_gpu_temperature_celsius` | | `DCGM_FI_DEV_POWER_USAGE` | `modelplane_gpu_power_watts` | | `DCGM_FI_PROF_PIPE_TENSOR_ACTIVE` | `modelplane_gpu_tensor_active_ratio` | | `envoy_cluster_upstream_rq_time` | `modelplane_frontend_request_duration_seconds` | -| `envoy_cluster_upstream_rq_xx` | `modelplane_requests_total{status}` | -Two have no direct replacement. `vllm:inter_token_latency_seconds` isn't renamed, because +Some have no replacement. A rename carries one metric to one name, so the counters that +would fold several series under one label - tokens by direction, responses by reason, +requests by status - aren't part of this surface yet. Keep reading those from your engine +and your gateway directly. For prompt and output size, `modelplane_request_input_tokens` +and `modelplane_request_output_tokens` carry the same measurement as histograms. + +`vllm:inter_token_latency_seconds` isn't renamed, because SGLang publishes a metric of the same name measuring something else; use `modelplane_frontend_tpot_seconds`, which the gateway measures the same way for every engine. `DCGM_FI_DEV_GPU_UTIL` isn't renamed either, because it only tells you the card diff --git a/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt b/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt index 7fbfda46e..9486fc527 100644 --- a/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt +++ b/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt @@ -595,3 +595,4 @@ p99 OTTL per-deployment opt-out +Framebuffer diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index f9dbe1759..b1aa3aa72 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -772,10 +772,19 @@ def compose_collector(self) -> list[str]: if "destinations" not in self.req.required_resources or "mappings" not in self.req.required_resources: return [] - destinations = list(request.get_required_resources(self.req, "destinations")) + # Sorted, not whichever the API server listed first: the collector + # restarts on a change to its rendered config, so an unstable choice + # between two destinations would redeploy it on alternate reconciles. + destinations = sorted( + ( + tdv1alpha1.TelemetryDestination.model_validate(d) + for d in request.get_required_resources(self.req, "destinations") + ), + key=lambda d: _name(d.metadata), + ) if not destinations: return [] - dest = tdv1alpha1.TelemetryDestination.model_validate(destinations[0]) + dest = destinations[0] if len(destinations) > 1: # Which one wins would otherwise be whichever the API server listed # first, and a fleet would export somewhere nobody chose. diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index a7f670f49..8c9dc8c65 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -14,11 +14,14 @@ """Tests for the collector this stack composes.""" +import typing import unittest import yaml from function import collector, stacks +from models.ai.modelplane.metricmapping import v1alpha1 as mmv1alpha1 from models.ai.modelplane.telemetrydestination import v1alpha1 as tdv1alpha1 +from pydantic import ValidationError _EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/primary/token"}} @@ -111,6 +114,21 @@ def test_dcgm_units_are_converted_to_the_unit_the_name_claims(self) -> None: self.assertIn("DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION", scales) self.assertIn("DCGM_FI_DEV_FB_USED", scales) + def test_every_unit_the_api_offers_has_a_conversion(self) -> None: + """A unit the API accepts with no conversion here is a KeyError at render time.""" + annotation = mmv1alpha1.Metric.model_fields["fromUnit"].annotation + literal = next(a for a in typing.get_args(annotation) if typing.get_origin(a) is typing.Literal) + self.assertEqual(set(typing.get_args(literal)), set(collector._UNIT_CONVERSION)) + + def test_a_metric_name_cannot_end_the_comparison_early(self) -> None: + """A quote in `from` would rename whatever the rest of the line matched.""" + with self.assertRaises(ValidationError): + mmv1alpha1.Metric.model_validate({"from": 'x" or true or name == "y', "to": "modelplane_x"}) + for mapping in stacks.BUILTIN_MAPPINGS: + for m in mapping.spec.metrics: + round_tripped = mmv1alpha1.Metric.model_validate({"from": m.from_, "to": m.to}) + self.assertEqual(round_tripped.from_, m.from_) + def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: """value_double is a datapoint path; the collector refuses to start on it here.""" blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] diff --git a/schemas/.lock.json b/schemas/.lock.json index 21a992e30..453396173 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "80357e7a8407a4046262d4c3c0e13dfd4f5da938b0e1368321439319842b46bd", + "fs://apis": "934848f9e76217a886a84295c279bd9366c23d153c854f465e02ea1574c17d28", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index b1352fc7a..e108026ae 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -43,9 +43,12 @@ class Crossplane(BaseModel): class Metric(BaseModel): - from_: constr(max_length=255) = Field(..., alias='from') + from_: constr(pattern=r'^[a-zA-Z_:][a-zA-Z0-9_:]*$', max_length=255) = Field( + ..., alias='from' + ) """ The metric's name as the component emits it, matched exactly. Nothing here declares which engine a deployment runs: a name that no component emits simply matches nothing. + Held to the characters a metric name can contain. The name is matched inside the collector's own query language, so a quote here would end the comparison early and rename whatever the rest of the line matched. """ fromUnit: ( Literal['Millijoules', 'Mebibytes', 'Milliseconds', 'Nanoseconds'] | None diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index 1ee121e19..356e2fbbc 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -61,7 +61,7 @@ class Sink(BaseModel): secretRef: SecretRef | None = None """ A Secret holding this sink's credential. Its keys reach the collector as environment variables, for config above referring to ${env:TOKEN}, and as files under /etc/modelplane/telemetry//, for an authenticator reading one from disk. - Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. A file is refreshed in place where an environment variable is fixed for the life of the process, so an authenticator reading the file picks up a rotated credential without a restart. + Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. The files are per sink; the environment variables are not, so two Secrets sharing a key name still collide there and the file is the one to read. A file is refreshed in place where an environment variable is fixed for the life of the process, so an authenticator reading the file picks up a rotated credential without a restart. """ type: constr(max_length=63) """ From de140e37e1d0f14c151b3496577fdf31134bed87 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 10:57:15 -0700 Subject: [PATCH 11/42] Regenerate the schema lock after rebasing The lock records one content hash over apis/, so rebasing onto a main that changed apis/ leaves it describing neither tree. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- schemas/.lock.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/schemas/.lock.json b/schemas/.lock.json index 453396173..6b090f1e5 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "934848f9e76217a886a84295c279bd9366c23d153c854f465e02ea1574c17d28", + "fs://apis": "aef772e69a6ad58b6600ee5cbdb1759d84b258af1049aa49920408320894dfff", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", From 9912a90b25482e0fa2a206cbe3f0f2a867aadc3e Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 11:23:43 -0700 Subject: [PATCH 12/42] Type the parts of a sink that belong to Modelplane A sink was a name, an exporter type, and a block of configuration passed through whole. That left the endpoint - the one setting every exporter has and every sink needs - as an unvalidated key inside the blob, and it made authentication something an operator had to assemble by hand: the collector carries no credential on an exporter, only a reference to an extension, so bearer auth meant writing the extension, naming it, and referencing it from the sink. Lift the two out. `endpoint` is required and typed. `auth.bearerTokenKey` names the key in the sink's Secret, and Modelplane composes the bearertokenauth extension and wires the reference. Both are rendered after the operator's config, so a passed-through key can't quietly redirect a sink or unpick its credential. The rest stays pass-through, because the rest is OpenTelemetry's: TLS is sixteen keys, the retry and queue blocks a dozen more, and the schema moves on its own schedule - the Prometheus remote-write exporter is deprecating its top-level HTTP settings in this release series. Typing that would mean a Modelplane release per setting, and telling anyone who needs one we have never heard of to wait for it. A scheme Modelplane doesn't compose is unaffected: define the extension under spec.extensions and name it from the sink's config, which is what the auth block does on your behalf. The function counts both as defining an authenticator, so a sink using either is still checked against the failure that makes a collector refuse to start. Verified against `validate` in the collector image: the composed extension, its reference, and a sink carrying compression and queue settings alongside. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/telemetrydestinations/definition.yaml | 69 ++++++++++++++----- docs/content/platform/telemetry.md | 56 +++++++++------ .../function/collector.py | 43 +++++++++++- .../tests/test_collector.py | 34 +++++++-- .../function/fn.py | 17 +++-- .../tests/test_fn.py | 42 +++++++++-- schemas/.lock.json | 2 +- .../telemetrydestination/v1alpha1.py | 27 ++++++-- 8 files changed, 231 insertions(+), 59 deletions(-) diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index 3f1cf09c9..3c97e6dcd 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -47,7 +47,10 @@ spec: x-kubernetes-list-map-keys: [name] items: type: object - required: [name, type, config] + required: [name, type, endpoint] + x-kubernetes-validations: + - rule: "!has(self.auth) || has(self.secretRef)" + message: secretRef is required when auth is set, because the credential lives in it. properties: name: type: string @@ -55,7 +58,8 @@ spec: pattern: '^[a-z0-9]([-a-z0-9]*[a-z0-9])?$' description: >- This sink's name, unique within the destination. It - names the collector's exporter instance and the + names the collector's exporter instance, the + authenticator Modelplane composes for it, and the directory its credential mounts at, so renaming one restarts the collector. type: @@ -71,35 +75,68 @@ spec: a Modelplane release for each exporter the collector gains, and the collector already refuses to start on a name it doesn't have. + endpoint: + type: string + maxLength: 2048 + description: >- + Where this sink writes. Typed rather than left to the + configuration below because every exporter has one and + a destination with no endpoint is the mistake worth + catching here rather than in a collector that won't + start. + auth: + type: object + description: >- + How to authenticate, for the schemes Modelplane + composes. The collector takes no credential inline: it + authenticates through an extension an exporter names, + so setting this composes that extension and wires the + reference. + + A scheme that isn't here is still reachable. Define + the extension yourself under spec.extensions and name + it from this sink's config, which is what Modelplane + does on your behalf. + properties: + bearerTokenKey: + type: string + maxLength: 253 + description: >- + The key in this sink's Secret holding the bearer + token. Modelplane mounts it as a file and points + the authenticator at it, so a rotated token is + picked up without restarting the collector. config: type: object x-kubernetes-preserve-unknown-fields: true description: >- - That exporter's own configuration, passed through - unread. Modelplane does not model what an exporter is, - so TLS, retry, queue and compression settings all - work, and a sink keeps working when the collector - gains a setting Modelplane has never heard of. + Anything else that exporter takes, passed through + unread: TLS, retry, queueing, compression, headers. + + Modelplane does not model an exporter's configuration, + because the schema is OpenTelemetry's and versioned + separately. Typing it would mean a Modelplane release + for each setting the collector gains, and would drop + the ones this has never heard of. What is typed above + is what belongs to Modelplane: which sinks exist, what + each is called, where it writes, and which Secret it + reads. secretRef: type: object required: [name] description: >- A Secret holding this sink's credential. Its keys - reach the collector as environment variables, for - config above referring to ${env:TOKEN}, and as files - under /etc/modelplane/telemetry//, for an - authenticator reading one from disk. + reach the collector as files under + /etc/modelplane/telemetry//, and as + environment variables, for configuration above + referring to ${env:TOKEN}. Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. The files are per sink; the environment variables are not, so two Secrets sharing a key name still collide there and - the file is the one to read. A file is - refreshed in place where an environment variable is - fixed for the life of the process, so an authenticator - reading the file picks up a rotated credential without - a restart. + the file is the one to read. properties: name: type: string diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index ba3ae06a7..8c352885e 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -64,33 +64,30 @@ spec: sinks: - name: primary type: otlphttp - config: - endpoint: https://otel.example.internal + endpoint: https://otel.example.internal ``` -`type` names a collector exporter, and `config` is that exporter's own configuration, so any -exporter the collector provides works here with its usual TLS and retry settings. +`type` names a collector exporter, by the name OpenTelemetry gives it. -Put credentials in a Secret and name it with the sink's `secretRef`. Modelplane mounts its -keys as environment variables, so your config refers to `${env:OTLP_TOKEN}` and the token -never appears in `kubectl get -o yaml`: +Put the credential in a Secret, name it with the sink's `secretRef`, and say which key holds +the token. Modelplane composes the authenticator and wires it up, and the token never +appears in `kubectl get -o yaml`: ```yaml spec: - extensions: - bearertokenauth: - token: ${env:OTLP_TOKEN} sinks: - name: primary type: otlphttp + endpoint: https://otel.example.internal secretRef: name: telemetry-credentials - config: - endpoint: https://otel.example.internal - auth: - authenticator: bearertokenauth + auth: + bearerTokenKey: token ``` +It reads the token from a file rather than the environment, so rotating it doesn't need the +collector restarted. + If you run Prometheus, export to that instead and query the fleet there: ```yaml @@ -98,8 +95,7 @@ spec: sinks: - name: prometheus type: prometheusremotewrite - config: - endpoint: https://prom.example.internal/api/v1/write + endpoint: https://prom.example.internal/api/v1/write ``` Name more than one sink and every one gets the whole stream. Each carries its own @@ -110,18 +106,38 @@ spec: sinks: - name: vendor type: otlphttp + endpoint: https://otel.vendor.example secretRef: name: vendor-token - config: - endpoint: https://otel.vendor.example + auth: + bearerTokenKey: token - name: prometheus type: prometheusremotewrite - config: - endpoint: https://prom.example.internal/api/v1/write + endpoint: https://prom.example.internal/api/v1/write ``` That is two copies of the fleet's metrics, billed twice. +Anything else the exporter takes goes under `config`, passed through as you wrote it: + +```yaml + - name: vendor + type: otlphttp + endpoint: https://otel.vendor.example + config: + compression: gzip + sending_queue: + queue_size: 10000 + tls: + ca_file: /etc/ssl/certs/internal.pem +``` + +Modelplane doesn't model what an exporter is, so its TLS, retry and queue settings all work, +and a sink keeps working when the collector gains a setting Modelplane has never heard of. +An authentication scheme Modelplane doesn't compose works the same way: define the extension +under `spec.extensions` and name it from the sink's `config`, which is what the `auth` block +above does for you. + Until you create one, Modelplane composes no collectors: nothing here stores anything, so collecting with nowhere to send it would spend GPU-cluster memory on samples nobody reads. Creating a destination turns collection on everywhere at once, and there's no per-deployment diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 6b3bacad2..d368c67a5 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -34,6 +34,7 @@ NAMESPACE = "modelplane-system" NAME = "modelplane-collector" +_CREDENTIALS_DIR = "/etc/modelplane/telemetry" IMAGE = "otel/opentelemetry-collector-contrib:0.161.0" @@ -160,14 +161,49 @@ def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: return {"metric_statements": blocks} +def _credential_path(sink: tdv1alpha1.Sink, key: str) -> str: + return f"{_CREDENTIALS_DIR}/{sink.name}/{key}" + + +def authenticators(sinks: list[tdv1alpha1.Sink]) -> dict[str, Any]: + """The extensions Modelplane composes for the sinks that asked for one. + + The collector carries no credential on an exporter: it authenticates + through an extension the exporter names. A typed auth block is therefore + an extension plus a reference, not a field, and composing both is what + keeps `auth` from being something an operator has to assemble by hand. + """ + composed: dict[str, Any] = {} + for sink in sinks: + if sink.auth and sink.auth.bearerTokenKey: + composed[f"bearertokenauth/{sink.name}"] = { + # A file rather than the environment: an environment variable + # is fixed for the life of the process, so a rotated token + # would need a restart to be read. + "filename": _credential_path(sink, sink.auth.bearerTokenKey), + } + return composed + + def exporters(sinks: list[tdv1alpha1.Sink]) -> dict[str, Any]: """The sinks as the collector's exporters block. A sink renders under `/`, which is how the collector names a second instance of one component, so two sinks of the same type don't collide. + + The endpoint and the authenticator reference are Modelplane's, and go on + last: an operator's own config can carry anything the exporter takes, but + not quietly redirect the sink somewhere else or unpick its credential. """ - return {f"{sink.type}/{sink.name}": dict(sink.config) for sink in sinks} + rendered: dict[str, Any] = {} + for sink in sinks: + cfg: dict[str, Any] = dict(sink.config or {}) + cfg["endpoint"] = sink.endpoint + if sink.auth and sink.auth.bearerTokenKey: + cfg["auth"] = {"authenticator": f"bearertokenauth/{sink.name}"} + rendered[f"{sink.type}/{sink.name}"] = cfg + return rendered def config( @@ -183,6 +219,9 @@ def config( same to carry as one that was renamed. """ sink_exporters = exporters(sinks) + # Modelplane's authenticators last: an operator's extensions can define + # anything, but not replace the one composed for a sink's own auth block. + extensions = {**extensions, **authenticators(sinks)} processors: dict[str, Any] = { # cluster is stamped here rather than downstream: one receiver on the # control plane sees a merged stream and cannot tell senders apart. @@ -263,7 +302,7 @@ def objects( # own directory, so two sinks can both hold a key called `token`. volume = f"credentials-{sink.name}" volumes.append({"name": volume, "secret": {"secretName": sink.secretRef.name}}) - mounts.append({"name": volume, "mountPath": f"/etc/modelplane/telemetry/{sink.name}", "readOnly": True}) + mounts.append({"name": volume, "mountPath": f"{_CREDENTIALS_DIR}/{sink.name}", "readOnly": True}) env_from.append({"secretRef": {"name": sink.secretRef.name}}) return [ diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 8c9dc8c65..2cfefbcc3 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -23,7 +23,7 @@ from models.ai.modelplane.telemetrydestination import v1alpha1 as tdv1alpha1 from pydantic import ValidationError -_EXTENSIONS = {"bearertokenauth": {"filename": "/etc/modelplane/telemetry/primary/token"}} +_EXTENSIONS = {"oidc/acme": {"issuer_url": "https://issuer.acme.example"}} def _sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = None) -> tdv1alpha1.Sink: @@ -31,8 +31,8 @@ def _sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = N { "name": name, "type": type_, - "config": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}}, - **({"secretRef": {"name": secret}} if secret else {}), + "endpoint": "https://otel.acme.example", + **({"secretRef": {"name": secret}, "auth": {"bearerTokenKey": "token"}} if secret else {}), } ) @@ -91,7 +91,7 @@ def test_gateway_has_a_target_of_its_own(self) -> None: def test_extensions_are_declared_to_the_service(self) -> None: """An authenticator the service doesn't list is one the collector won't load.""" - self.assertEqual(_config()["service"]["extensions"], ["bearertokenauth"]) + self.assertEqual(_config()["service"]["extensions"], ["oidc/acme"]) self.assertNotIn("extensions", _config(extensions={})["service"]) def test_energy_is_scaled_before_it_is_renamed(self) -> None: @@ -182,6 +182,32 @@ def test_two_sinks_of_one_type_do_not_collide(self) -> None: rendered = collector.exporters([_sink(name="a"), _sink(name="b")]) self.assertEqual(sorted(rendered), ["otlphttp/a", "otlphttp/b"]) + def test_auth_composes_its_own_authenticator(self) -> None: + """The collector carries no credential on an exporter, only a reference.""" + sink = _sink(secret="telemetry-credentials") + self.assertEqual( + collector.authenticators([sink]), + {"bearertokenauth/primary": {"filename": "/etc/modelplane/telemetry/primary/token"}}, + ) + self.assertEqual( + collector.exporters([sink])["otlphttp/primary"]["auth"], + {"authenticator": "bearertokenauth/primary"}, + ) + + def test_a_sinks_own_config_cannot_redirect_it(self) -> None: + """The endpoint is Modelplane's, and goes on after the operator's config.""" + sink = tdv1alpha1.Sink.model_validate( + { + "name": "primary", + "type": "otlphttp", + "endpoint": "https://otel.acme.example", + "config": {"endpoint": "https://elsewhere.example", "compression": "gzip"}, + } + ) + rendered = collector.exporters([sink])["otlphttp/primary"] + self.assertEqual(rendered["endpoint"], "https://otel.acme.example") + self.assertEqual(rendered["compression"], "gzip") + def test_no_secret_mounts_nothing(self) -> None: pod = self._objects()["collector"]["spec"]["template"]["spec"] self.assertEqual([v["name"] for v in pod["volumes"]], ["config"]) diff --git a/functions/compose-telemetry-destination/function/fn.py b/functions/compose-telemetry-destination/function/fn.py index 58e667c52..2c026234b 100644 --- a/functions/compose-telemetry-destination/function/fn.py +++ b/functions/compose-telemetry-destination/function/fn.py @@ -18,9 +18,9 @@ carrying an exporter's own configuration verbatim, and compose-serving-stack renders them into the collector it composes. Modelplane does not model what an exporter is, so there is little here to validate and the little there is -matters: a sink naming an authenticator that no extension defines makes a -collector refuse to start, and that failure surfaces as telemetry silently -never arriving. +matters: a sink naming an authenticator that nothing defines makes a collector +refuse to start, and that failure surfaces as telemetry silently never +arriving. """ import grpc @@ -61,8 +61,11 @@ async def RunFunction( # collector refuses to start when it names one no extension defines, and # a collector that never starts looks exactly like a fleet that produces # nothing, so it is worth catching on the object instead. - extensions = set((xr.spec.extensions or {}).keys()) - missing = sorted(_authenticators(sinks) - extensions) + # An authenticator is defined either by the operator, under extensions, + # or by Modelplane, for a sink that set auth. Both count. + defined = set((xr.spec.extensions or {}).keys()) + defined |= {f"bearertokenauth/{s.name}" for s in sinks if s.auth and s.auth.bearerTokenKey} + missing = sorted(_authenticators(sinks) - defined) if missing: _not_ready( rsp, @@ -108,10 +111,10 @@ async def RunFunction( def _authenticators(sinks: list[v1alpha1.Sink]) -> set[str]: - """Every authenticator a sink references, by extension name.""" + """Every authenticator a sink's own config references, by extension name.""" names: set[str] = set() for sink in sinks: - auth = sink.config.get("auth") + auth = (sink.config or {}).get("auth") if isinstance(auth, dict) and isinstance(auth.get("authenticator"), str): names.add(auth["authenticator"]) return names diff --git a/functions/compose-telemetry-destination/tests/test_fn.py b/functions/compose-telemetry-destination/tests/test_fn.py index e15b2b0b0..f6c8c6ede 100644 --- a/functions/compose-telemetry-destination/tests/test_fn.py +++ b/functions/compose-telemetry-destination/tests/test_fn.py @@ -49,15 +49,17 @@ async def test_compose(self) -> None: """The function reports whether a destination can actually be sent through.""" def sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = None) -> dict: + """A sink wiring its own authenticator, which is the case worth validating.""" return { "name": name, "type": type_, - "config": {"endpoint": "https://otel.acme.example", "auth": {"authenticator": "bearertokenauth"}}, + "endpoint": "https://otel.acme.example", + "config": {"auth": {"authenticator": "oidc/acme"}}, **({"secretRef": {"name": secret}} if secret else {}), } sinks = [sink()] - extensions = {"bearertokenauth": {"token": "${env:OTLP_TOKEN}"}} + extensions = {"oidc/acme": {"issuer_url": "https://issuer.acme.example"}} def xr(spec: dict) -> dict: return { @@ -116,7 +118,7 @@ def want( { "name": "prom", "type": "prometheusremotewrite", - "config": {"endpoint": "https://prom.acme.example/api/v1/write"}, + "endpoint": "https://prom.acme.example/api/v1/write", } ] } @@ -132,6 +134,38 @@ def want( ), ), ), + Case( + name="ready with no extensions, because Modelplane composes the authenticator", + req=req( + { + "sinks": [ + { + "name": "primary", + "type": "otlphttp", + "endpoint": "https://otel.acme.example", + "secretRef": {"name": "telemetry-credentials"}, + "auth": {"bearerTokenKey": "token"}, + } + ] + }, + secrets=[ + resource.dict_to_struct( + {"apiVersion": "v1", "kind": "Secret", "metadata": {"name": "telemetry-credentials"}} + ) + ], + ), + want=want( + fnv1.READY_TRUE, + {"status": {}}, + fnv1.Condition( + type="Accepted", + status=fnv1.STATUS_CONDITION_TRUE, + reason="Available", + message="Exporting through otlphttp/primary", + ), + secret="telemetry-credentials", + ), + ), Case( name="ready once the credential Secret exists", req=req( @@ -181,7 +215,7 @@ def want( type="Accepted", status=fnv1.STATUS_CONDITION_FALSE, reason="UnknownAuthenticator", - message="No extension defines bearertokenauth, so the collector would refuse to start", + message="No extension defines oidc/acme, so the collector would refuse to start", ), ), ), diff --git a/schemas/.lock.json b/schemas/.lock.json index 6b090f1e5..281017d5c 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "aef772e69a6ad58b6600ee5cbdb1759d84b258af1049aa49920408320894dfff", + "fs://apis": "7d528c8db3656a4169462c3c41331c313310b3e87e65d61e17c8da4e1149e381", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index 356e2fbbc..f8d6778a2 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -42,6 +42,13 @@ class Crossplane(BaseModel): resourceRefs: list[ResourceRef] | None = None +class Auth(BaseModel): + bearerTokenKey: constr(max_length=253) | None = None + """ + The key in this sink's Secret holding the bearer token. Modelplane mounts it as a file and points the authenticator at it, so a rotated token is picked up without restarting the collector. + """ + + class SecretRef(BaseModel): name: constr(max_length=253) """ @@ -50,18 +57,28 @@ class SecretRef(BaseModel): class Sink(BaseModel): - config: dict[str, Any] + auth: Auth | None = None + """ + How to authenticate, for the schemes Modelplane composes. The collector takes no credential inline: it authenticates through an extension an exporter names, so setting this composes that extension and wires the reference. + A scheme that isn't here is still reachable. Define the extension yourself under spec.extensions and name it from this sink's config, which is what Modelplane does on your behalf. + """ + config: dict[str, Any] | None = None + """ + Anything else that exporter takes, passed through unread: TLS, retry, queueing, compression, headers. + Modelplane does not model an exporter's configuration, because the schema is OpenTelemetry's and versioned separately. Typing it would mean a Modelplane release for each setting the collector gains, and would drop the ones this has never heard of. What is typed above is what belongs to Modelplane: which sinks exist, what each is called, where it writes, and which Secret it reads. + """ + endpoint: constr(max_length=2048) """ - That exporter's own configuration, passed through unread. Modelplane does not model what an exporter is, so TLS, retry, queue and compression settings all work, and a sink keeps working when the collector gains a setting Modelplane has never heard of. + Where this sink writes. Typed rather than left to the configuration below because every exporter has one and a destination with no endpoint is the mistake worth catching here rather than in a collector that won't start. """ name: constr(pattern=r'^[a-z0-9]([-a-z0-9]*[a-z0-9])?$', max_length=63) """ - This sink's name, unique within the destination. It names the collector's exporter instance and the directory its credential mounts at, so renaming one restarts the collector. + This sink's name, unique within the destination. It names the collector's exporter instance, the authenticator Modelplane composes for it, and the directory its credential mounts at, so renaming one restarts the collector. """ secretRef: SecretRef | None = None """ - A Secret holding this sink's credential. Its keys reach the collector as environment variables, for config above referring to ${env:TOKEN}, and as files under /etc/modelplane/telemetry//, for an authenticator reading one from disk. - Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. The files are per sink; the environment variables are not, so two Secrets sharing a key name still collide there and the file is the one to read. A file is refreshed in place where an environment variable is fixed for the life of the process, so an authenticator reading the file picks up a rotated credential without a restart. + A Secret holding this sink's credential. Its keys reach the collector as files under /etc/modelplane/telemetry//, and as environment variables, for configuration above referring to ${env:TOKEN}. + Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. The files are per sink; the environment variables are not, so two Secrets sharing a key name still collide there and the file is the one to read. """ type: constr(max_length=63) """ From 8fc7fca355a775c2bdeb1e3c5cc9f750327c00ce Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 11:46:02 -0700 Subject: [PATCH 13/42] Let a sink address its destination the way its exporter does `endpoint` was required, on the assumption every exporter has one. Several don't: Kafka takes `brokers`, the file exporter a `path`, and the debug exporter nothing at all. The field's own neighbour advertises Kafka as a supported type, so the schema contradicted itself and made three exporters unreachable - including the one you would reach for first when nothing is arriving. Make it optional. It stays typed, because it is still the setting every destination that has one has to get right, and it is still rendered after the passed-through config so it cannot be overridden from there. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/telemetrydestinations/definition.yaml | 12 +++++++++--- .../compose-serving-stack/function/collector.py | 3 ++- .../compose-serving-stack/tests/test_collector.py | 13 +++++++++++++ schemas/.lock.json | 2 +- .../ai/modelplane/telemetrydestination/v1alpha1.py | 5 +++-- 5 files changed, 28 insertions(+), 7 deletions(-) diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index 3c97e6dcd..858034035 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -47,7 +47,7 @@ spec: x-kubernetes-list-map-keys: [name] items: type: object - required: [name, type, endpoint] + required: [name, type] x-kubernetes-validations: - rule: "!has(self.auth) || has(self.secretRef)" message: secretRef is required when auth is set, because the credential lives in it. @@ -80,10 +80,16 @@ spec: maxLength: 2048 description: >- Where this sink writes. Typed rather than left to the - configuration below because every exporter has one and - a destination with no endpoint is the mistake worth + configuration below because it is the setting every + destination has to get right, and the one worth catching here rather than in a collector that won't start. + + Optional, because not every exporter addresses its + destination this way: Kafka takes brokers, the file + exporter a path, and the debug exporter nothing at + all. Those go in the configuration below, under the + names that exporter gives them. auth: type: object description: >- diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index d368c67a5..3db168ea3 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -199,7 +199,8 @@ def exporters(sinks: list[tdv1alpha1.Sink]) -> dict[str, Any]: rendered: dict[str, Any] = {} for sink in sinks: cfg: dict[str, Any] = dict(sink.config or {}) - cfg["endpoint"] = sink.endpoint + if sink.endpoint: + cfg["endpoint"] = sink.endpoint if sink.auth and sink.auth.bearerTokenKey: cfg["auth"] = {"authenticator": f"bearertokenauth/{sink.name}"} rendered[f"{sink.type}/{sink.name}"] = cfg diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 2cfefbcc3..507d54cb8 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -182,6 +182,19 @@ def test_two_sinks_of_one_type_do_not_collide(self) -> None: rendered = collector.exporters([_sink(name="a"), _sink(name="b")]) self.assertEqual(sorted(rendered), ["otlphttp/a", "otlphttp/b"]) + def test_a_sink_that_addresses_its_destination_another_way(self) -> None: + """Kafka takes brokers, the debug exporter nothing; neither has an endpoint.""" + sinks = [ + tdv1alpha1.Sink.model_validate( + {"name": "bus", "type": "kafka", "config": {"brokers": ["kafka.acme.example:9092"]}} + ), + tdv1alpha1.Sink.model_validate({"name": "seen", "type": "debug"}), + ] + rendered = collector.exporters(sinks) + self.assertNotIn("endpoint", rendered["kafka/bus"]) + self.assertEqual(rendered["kafka/bus"]["brokers"], ["kafka.acme.example:9092"]) + self.assertEqual(rendered["debug/seen"], {}) + def test_auth_composes_its_own_authenticator(self) -> None: """The collector carries no credential on an exporter, only a reference.""" sink = _sink(secret="telemetry-credentials") diff --git a/schemas/.lock.json b/schemas/.lock.json index 281017d5c..ef60257b9 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "7d528c8db3656a4169462c3c41331c313310b3e87e65d61e17c8da4e1149e381", + "fs://apis": "2b2af1570db437b60cb3574aa86e1ed1449a410ef6869629b8ba2f9fccabacb4", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index f8d6778a2..11b278c92 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -67,9 +67,10 @@ class Sink(BaseModel): Anything else that exporter takes, passed through unread: TLS, retry, queueing, compression, headers. Modelplane does not model an exporter's configuration, because the schema is OpenTelemetry's and versioned separately. Typing it would mean a Modelplane release for each setting the collector gains, and would drop the ones this has never heard of. What is typed above is what belongs to Modelplane: which sinks exist, what each is called, where it writes, and which Secret it reads. """ - endpoint: constr(max_length=2048) + endpoint: constr(max_length=2048) | None = None """ - Where this sink writes. Typed rather than left to the configuration below because every exporter has one and a destination with no endpoint is the mistake worth catching here rather than in a collector that won't start. + Where this sink writes. Typed rather than left to the configuration below because it is the setting every destination has to get right, and the one worth catching here rather than in a collector that won't start. + Optional, because not every exporter addresses its destination this way: Kafka takes brokers, the file exporter a path, and the debug exporter nothing at all. Those go in the configuration below, under the names that exporter gives them. """ name: constr(pattern=r'^[a-z0-9]([-a-z0-9]*[a-z0-9])?$', max_length=63) """ From d878697567acb41f0fdf2a255075d8f35621284a Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 12:13:38 -0700 Subject: [PATCH 14/42] Scrape the endpoints these components actually publish Running this on a cluster showed the collector starting cleanly and collecting almost nothing. The gateway job asked Envoy for /metrics and got a 404 every interval. Envoy publishes Prometheus on its admin port and annotates the pod with the path, so take the path from the annotation. This is the whole modelplane_frontend_* group - the series an SLO is written against. The substrate job honoured prometheus.io/scrape and then ignored prometheus.io/port, so it scraped whichever port a pod declared first: the cert-manager webhook declares 10250 and annotates 9402, and the collector logged a 400 against its TLS port every interval. Honour the rest of the same convention. With the port honoured the gateways match the substrate job too, since they annotate themselves, so drop from it the pods the other two jobs already name. Three jobs over disjoint sets. Also emit OTTL paths with their context. The collector accepts bare ones, rewrites them, and logs every statement it rewrote asking the author to stop. Modelplane is the author, so this is a change nobody's stored MetricMapping has to make. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 77 ++++++++++++++++--- .../tests/test_collector.py | 4 + 2 files changed, 72 insertions(+), 9 deletions(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 3db168ea3..abe127e82 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -66,11 +66,43 @@ def _relabel_pod_identity() -> list[dict[str, Any]]: ] + [{"source_labels": ["__meta_kubernetes_namespace"], "target_label": "namespace"}] +def _annotated_path() -> list[dict[str, Any]]: + """Take the metrics path from the pod's own annotation, where it sets one.""" + return [ + { + "source_labels": ["__meta_kubernetes_pod_annotation_prometheus_io_path"], + "action": "replace", + "target_label": "__metrics_path__", + "regex": "(.+)", + } + ] + + +def _annotated_port() -> list[dict[str, Any]]: + """Take the port from the pod's own annotation, where it sets one. + + Service discovery makes a target of every declared container port, so + without this a pod is scraped on whichever it declared first. Rewriting + them all to the annotated one leaves identical targets, which discovery + then collapses to one. + """ + return [ + { + "source_labels": ["__address__", "__meta_kubernetes_pod_annotation_prometheus_io_port"], + "action": "replace", + "target_label": "__address__", + "regex": r"([^:]+)(?::\d+)?;(\d+)", + "replacement": "$1:$2", + } + ] + + def _scrape_configs() -> list[dict[str, Any]]: """What to scrape on an inference cluster. - The gateway's GenAI metrics sit on the ext-proc sidecar's admin port rather - than the proxy's, so the front door needs a target of its own. + Three jobs over disjoint sets of pods, so nothing is scraped twice: the + engines Modelplane runs, the gateways in front of them, and everything else + the serving stack installs. """ return [ { @@ -98,6 +130,11 @@ def _scrape_configs() -> list[dict[str, Any]]: "regex": ".+", }, {"source_labels": ["__meta_kubernetes_pod_container_port_name"], "action": "keep", "regex": "metrics"}, + # Envoy publishes Prometheus on its admin port, not at + # /metrics, and says so in its own annotation. Without this the + # front door 404s every interval and the modelplane_frontend_* + # series - the ones an SLO is written against - never arrive. + *_annotated_path(), ], }, { @@ -110,6 +147,26 @@ def _scrape_configs() -> list[dict[str, Any]]: "action": "keep", "regex": "true", }, + # Both of these annotate themselves for scraping, and both have + # a job above that gives them their identity. Without the drops + # they are collected twice, under two job names. + { + "source_labels": ["__meta_kubernetes_pod_label_gateway_envoyproxy_io_owning_gateway_name"], + "action": "drop", + "regex": ".+", + }, + { + "source_labels": [f"__meta_kubernetes_pod_label_{_SERVING_LABEL}"], + "action": "drop", + "regex": "true", + }, + # The rest of the same convention, not just the first line of + # it. A pod that says scrape me generally also says where: the + # cert-manager webhook declares 10250 first and annotates 9402, + # so honouring only the keep scrapes its TLS port over plain + # HTTP and logs a 400 every interval. + *_annotated_path(), + *_annotated_port(), {"source_labels": ["__meta_kubernetes_namespace"], "target_label": "namespace"}, ], }, @@ -118,12 +175,14 @@ def _scrape_configs() -> list[dict[str, Any]]: # What each source unit is worth in the base unit the target name claims. # Written as the expression rather than a factor so nothing has to render a -# float: 1e-09 is not an OTTL literal. +# float: 1e-09 is not an OTTL literal. Paths carry their context because the +# collector rewrites bare ones and asks the author to stop; Modelplane is the +# author here, so nobody's stored MetricMapping has to change. _UNIT_CONVERSION = { - "Millijoules": "value_double / 1000", - "Milliseconds": "value_double / 1000", - "Nanoseconds": "value_double / 1000000000", - "Mebibytes": "value_double * 1048576", + "Millijoules": "datapoint.value_double / 1000", + "Milliseconds": "datapoint.value_double / 1000", + "Nanoseconds": "datapoint.value_double / 1000000000", + "Mebibytes": "datapoint.value_double * 1048576", } @@ -141,8 +200,8 @@ def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], lis for m in mapping.spec.metrics: if m.fromUnit: conversion = _UNIT_CONVERSION[m.fromUnit] - datapoint.append(f'set(value_double, {conversion}) where metric.name == "{m.from_}"') - metric.append(f'set(name, "{m.to}") where name == "{m.from_}"') + datapoint.append(f'set(datapoint.value_double, {conversion}) where metric.name == "{m.from_}"') + metric.append(f'set(metric.name, "{m.to}") where metric.name == "{m.from_}"') return datapoint, metric diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 507d54cb8..d5abd80fe 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -134,6 +134,10 @@ def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] metric_block = next(b for b in blocks if b["context"] == "metric") self.assertFalse([st for st in metric_block["statements"] if "value_double" in st]) + for block in blocks: + for st in block["statements"]: + self.assertNotIn("set(name,", st) + self.assertNotIn("set(value_double,", st) def test_sglang_latency_histograms_are_not_renamed(self) -> None: """Their buckets resolve to 100ms where vLLM's resolve to 1ms.""" From aa5006612a2852599dbf328537eda8e64f34dca4 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 12:24:16 -0700 Subject: [PATCH 15/42] Label serving pods with what their metrics belong to The collector reads a pod's labels during discovery to attribute the metrics it scrapes, and it read four that nothing sets. Engine metrics were never collected at all: the job matched modelplane.ai/serving against the literal "true", where that label carries the replica's name, so it selected nothing. Relaxing the match alone would have been worse - every series would have arrived with an empty deployment, engine and role, and the merge across replicas would have collapsed every engine on the cluster into one. Stamp the identity where the pod template is built, so a backend can't compose a serving pod without it. Deployment comes off the replica, which the composite already labels; engine and role off the engine and member. Workers carry it too: a worker holds GPUs, and the GPU series are the deployment's. Match on the deployment label's presence rather than its value. No model label. A ModelReplica doesn't know which ModelService fronts it, and a model name carries a slash, which a label value can't. Deployment is finer grained anyway - a deployment serves one model, a model may have several - so it replaces model in the merge keys and in the guide. Verified on the local two-cluster e2e: a pod carrying these labels is discovered, its vllm: and DCGM_ series arrive renamed, and each carries cluster, namespace, deployment, engine and role. DCGM_FI_DEV_FB_USED at 1024 MiB arrives as modelplane_gpu_memory_used_bytes at 1073741824. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- docs/content/platform/telemetry.md | 8 +-- .../function/backends/base.py | 43 +++++++++++++++- .../function/backends/grove.py | 4 +- .../function/backends/llmd.py | 6 ++- .../function/backends/native.py | 5 +- .../tests/test_backends.py | 50 +++++++++++++++---- .../compose-model-replica/tests/test_fn.py | 3 ++ .../function/collector.py | 23 +++++++-- 8 files changed, 117 insertions(+), 25 deletions(-) diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 8c352885e..9416adf56 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -20,7 +20,7 @@ deployments change. ## What you get Every series carries `cluster`. A series about a deployment also carries `deployment`, -`namespace`, `model`, and `engine`. Some of what you can read: +`namespace`, `engine`, and `role`. Some of what you can read: | Metric | Means | | --- | --- | @@ -235,9 +235,9 @@ everything twice, under `vllm:*` and under `modelplane_*`, paying for both. One against Modelplane's Prometheus stops being read by anything, because the operator goes with the stack. -**Rewrite your dashboard queries.** Names change, and so do three labels: `model_name` -becomes `model`, pod labels are gone because replicas are summed before they leave the -cluster, and every series now carries `cluster`. +**Rewrite your dashboard queries.** Names change, and so do the labels: group by +`deployment` rather than `model_name`, pod labels are gone because replicas are summed +before they leave the cluster, and every series now carries `cluster`. | Was | Is | | --- | --- | diff --git a/functions/compose-model-replica/function/backends/base.py b/functions/compose-model-replica/function/backends/base.py index 6129039f7..66860e86a 100644 --- a/functions/compose-model-replica/function/backends/base.py +++ b/functions/compose-model-replica/function/backends/base.py @@ -241,8 +241,46 @@ def remote_namespace(replica: v1alpha1.ModelReplica) -> str: # Deployment selectors fighting over each other's pods. LABEL_WORKLOAD = "modelplane.ai/workload" - -def pod_metadata(member: v1alpha1.Member, labels: dict[str, str] | None = None) -> dict: +# Pod labels carrying the identity every metric off this pod is attributed to. +# The collector reads them off the pod during service discovery and stamps them +# on each series, so a MetricMapping needs no labels of its own and a series +# merged across replicas keeps the identity all of them share. Metrics are the +# only reason these exist; nothing selects on them. +LABEL_DEPLOYMENT = "modelplane.ai/deployment" +LABEL_ENGINE = "modelplane.ai/engine" +LABEL_ROLE = "modelplane.ai/role" + + +def telemetry_labels( + replica: v1alpha1.ModelReplica, + engine: v1alpha1.Engine, + member: v1alpha1.Member, +) -> dict[str, str]: + """What a metric off this pod is attributed to. + + The deployment comes off the replica, which the composite already labels + with it. There is no model here on purpose: a ModelReplica doesn't know + which ModelService fronts it, and a model name carries a slash, which a + label value can't. + """ + labels: dict[str, str] = {} + deployment = (replica.metadata.labels if replica.metadata else None) or {} + if name := deployment.get(LABEL_DEPLOYMENT): + labels[LABEL_DEPLOYMENT] = name + if engine.name: + labels[LABEL_ENGINE] = engine.name + if member.role: + labels[LABEL_ROLE] = member.role + return labels + + +def pod_metadata( + member: v1alpha1.Member, + labels: dict[str, str] | None = None, + *, + replica: v1alpha1.ModelReplica, + engine: v1alpha1.Engine, +) -> dict: """Pod template metadata for a member: its template.metadata plus managed labels. The member's template.metadata.labels and .annotations propagate to the pod @@ -255,6 +293,7 @@ def pod_metadata(member: v1alpha1.Member, labels: dict[str, str] | None = None) """ user = member.template.metadata merged = dict((user.labels if user else None) or {}) + merged.update(telemetry_labels(replica, engine, member)) merged.update(labels or {}) meta: dict = {} if merged: diff --git a/functions/compose-model-replica/function/backends/grove.py b/functions/compose-model-replica/function/backends/grove.py index 70a4cfad8..20925eb2c 100644 --- a/functions/compose-model-replica/function/backends/grove.py +++ b/functions/compose-model-replica/function/backends/grove.py @@ -151,6 +151,8 @@ def pod_spec(member: v1alpha1.Member, c: dict) -> dict: base.GROVE_QUEUE_LABEL: base.GROVE_QUEUE, _LABEL_CLIQUE_ROLE: "leader", }, + replica=replica, + engine=engine, ), "spec": { "roleName": base.GROVE_LEADER_CLIQUE, @@ -170,7 +172,7 @@ def pod_spec(member: v1alpha1.Member, c: dict) -> dict: # stable DNS name until it's listening. worker_clique = { "name": base.GROVE_WORKER_CLIQUE, - **base.pod_metadata(worker, {base.GROVE_QUEUE_LABEL: base.GROVE_QUEUE}), + **base.pod_metadata(worker, {base.GROVE_QUEUE_LABEL: base.GROVE_QUEUE}, replica=replica, engine=engine), "spec": { "roleName": base.GROVE_WORKER_CLIQUE, "replicas": worker_replicas, diff --git a/functions/compose-model-replica/function/backends/llmd.py b/functions/compose-model-replica/function/backends/llmd.py index d5ce24fb0..8e4a82749 100644 --- a/functions/compose-model-replica/function/backends/llmd.py +++ b/functions/compose-model-replica/function/backends/llmd.py @@ -133,7 +133,9 @@ def pod_spec(member: v1alpha1.Member, c: dict) -> dict: # serving port, and the readiness probe. The leader member's own # template.metadata merges in underneath them. leader_pod = { - "metadata": base.pod_metadata(leader, {base.LABEL_SERVING: serving_label, _LABEL_ROLE: "leader"}), + "metadata": base.pod_metadata( + leader, {base.LABEL_SERVING: serving_label, _LABEL_ROLE: "leader"}, replica=replica, engine=engine + ), "spec": pod_spec(leader, container(leader, serving=True)), } # The worker followers don't serve the OpenAI API, so they carry no @@ -145,7 +147,7 @@ def pod_spec(member: v1alpha1.Member, c: dict) -> dict: worker_pod = { "spec": pod_spec(worker, container(worker, serving=False)), } - worker_metadata = base.pod_metadata(worker) + worker_metadata = base.pod_metadata(worker, replica=replica, engine=engine) if worker_metadata: worker_pod["metadata"] = worker_metadata diff --git a/functions/compose-model-replica/function/backends/native.py b/functions/compose-model-replica/function/backends/native.py index 14ffbcc52..52b1ffcae 100644 --- a/functions/compose-model-replica/function/backends/native.py +++ b/functions/compose-model-replica/function/backends/native.py @@ -115,7 +115,10 @@ def build( "spec": { "replicas": int(engine.copies or 1), "selector": {"matchLabels": selector}, - "template": {"metadata": base.pod_metadata(member, pod_labels), "spec": pod_spec}, + "template": { + "metadata": base.pod_metadata(member, pod_labels, replica=replica, engine=engine), + "spec": pod_spec, + }, }, } diff --git a/functions/compose-model-replica/tests/test_backends.py b/functions/compose-model-replica/tests/test_backends.py index e00f7cef6..12f1420ef 100644 --- a/functions/compose-model-replica/tests/test_backends.py +++ b/functions/compose-model-replica/tests/test_backends.py @@ -35,6 +35,8 @@ from models.io.k8s.apimachinery.pkg.apis.meta import v1 as metav1 _SERVING = "modelplane.ai/serving" +_ENGINE = "modelplane.ai/engine" +_ROLE = "modelplane.ai/role" _WORKLOAD = "modelplane.ai/workload" _CLIQUE_ROLE = "modelplane.ai/clique-role" _QUEUE_LABEL = "kai.scheduler/queue" @@ -206,7 +208,9 @@ def _claim_template(count: int, *, replica: str = "r", engine: str = "main", rol "replicas": 1, "selector": {"matchLabels": {_WORKLOAD: _WORKLOAD_NAME}}, "template": { - "metadata": {"labels": {_SERVING: "r", _WORKLOAD: _WORKLOAD_NAME}}, + "metadata": { + "labels": {_ENGINE: "main", _ROLE: "Standalone", _SERVING: "r", _WORKLOAD: _WORKLOAD_NAME} + }, "spec": { "containers": [ { @@ -283,7 +287,13 @@ def pod_spec(container: dict, role: str) -> dict: "cliques": [ { "name": "leader", - "labels": {_SERVING: "r", _QUEUE_LABEL: _QUEUE, _CLIQUE_ROLE: "leader"}, + "labels": { + _ENGINE: "main", + _ROLE: "Leader", + _SERVING: "r", + _QUEUE_LABEL: _QUEUE, + _CLIQUE_ROLE: "leader", + }, "spec": { "roleName": "leader", "replicas": 1, @@ -293,7 +303,7 @@ def pod_spec(container: dict, role: str) -> dict: }, { "name": "worker", - "labels": {_QUEUE_LABEL: _QUEUE}, + "labels": {_ENGINE: "main", _ROLE: "Worker", _QUEUE_LABEL: _QUEUE}, "spec": { "roleName": "worker", "replicas": worker_replicas, @@ -481,7 +491,13 @@ def test_member_metadata_propagates_to_native_pod_template(self) -> None: meta = out["model-serving-main"].spec.forProvider.manifest["spec"]["template"]["metadata"] self.assertEqual( meta["labels"], - {"example.com/role": "standalone", _SERVING: "r", _WORKLOAD: _WORKLOAD_NAME}, + { + "example.com/role": "standalone", + _ENGINE: "main", + _ROLE: "Standalone", + _SERVING: "r", + _WORKLOAD: _WORKLOAD_NAME, + }, ) self.assertEqual(meta["annotations"], {"example.com/config": "standalone"}) @@ -502,11 +518,20 @@ def test_member_metadata_propagates_to_cliques_independently(self) -> None: leader = _clique(manifest, "leader") self.assertEqual( leader["labels"], - {"example.com/role": "leader", _SERVING: "r", _QUEUE_LABEL: _QUEUE, _CLIQUE_ROLE: "leader"}, + { + "example.com/role": "leader", + _ENGINE: "main", + _ROLE: "Leader", + _SERVING: "r", + _QUEUE_LABEL: _QUEUE, + _CLIQUE_ROLE: "leader", + }, ) self.assertEqual(leader["annotations"], {"example.com/config": "leader"}) worker = _clique(manifest, "worker") - self.assertEqual(worker["labels"], {"example.com/role": "worker", _QUEUE_LABEL: _QUEUE}) + self.assertEqual( + worker["labels"], {"example.com/role": "worker", _ENGINE: "main", _ROLE: "Worker", _QUEUE_LABEL: _QUEUE} + ) self.assertEqual(worker["annotations"], {"example.com/config": "worker"}) def test_worker_without_metadata_composes_only_managed_labels(self) -> None: @@ -517,7 +542,7 @@ def test_worker_without_metadata_composes_only_managed_labels(self) -> None: out = grove.GroveBackend().build(replica, engine, _PC, base.serving_label(replica), "Dynamo") manifest = out["model-serving-main"].spec.forProvider.manifest worker = _clique(manifest, "worker") - self.assertEqual(worker["labels"], {_QUEUE_LABEL: _QUEUE}) + self.assertEqual(worker["labels"], {_ENGINE: "main", _ROLE: "Worker", _QUEUE_LABEL: _QUEUE}) self.assertNotIn("annotations", worker) @staticmethod @@ -668,8 +693,13 @@ def test_only_leader_carries_serving_label(self) -> None: leader_labels = lwt["leaderTemplate"]["metadata"]["labels"] self.assertEqual(leader_labels[_SERVING], "r") self.assertEqual(leader_labels[self._LWS_ROLE], "leader") - # The worker followers never serve, so they carry no metadata at all. - self.assertNotIn("metadata", lwt["workerTemplate"]) + # The worker followers never serve, so they carry no serving label and + # the replica's Service can't route to them. They do carry the + # telemetry identity: a worker holds GPUs, and its metrics are the + # deployment's. + worker_labels = lwt["workerTemplate"]["metadata"]["labels"] + self.assertNotIn(_SERVING, worker_labels) + self.assertEqual(worker_labels, {_ENGINE: "main", _ROLE: "Worker"}) def test_leader_address_and_rank_env_injected(self) -> None: # Every gang container leads with the backend-neutral coordination vars @@ -1099,7 +1129,7 @@ def test_decode_can_be_a_grove_gang(self) -> None: worker_clique = _clique(manifest, "worker") worker = worker_clique["spec"]["podSpec"] self.assertEqual([c["name"] for c in worker["containers"]], ["engine"]) - self.assertEqual(worker_clique["labels"], {_QUEUE_LABEL: _QUEUE}) + self.assertEqual(worker_clique["labels"], {_ENGINE: "decode", _ROLE: "Worker", _QUEUE_LABEL: _QUEUE}) class TestUnifiedRouting(unittest.TestCase): diff --git a/functions/compose-model-replica/tests/test_fn.py b/functions/compose-model-replica/tests/test_fn.py index 74ab85151..4fec047eb 100644 --- a/functions/compose-model-replica/tests/test_fn.py +++ b/functions/compose-model-replica/tests/test_fn.py @@ -197,6 +197,9 @@ async def test_compose(self) -> None: "template": { "metadata": { "labels": { + "modelplane.ai/deployment": "my-deployment", + "modelplane.ai/engine": "main", + "modelplane.ai/role": "Standalone", "modelplane.ai/serving": "test-replica", "modelplane.ai/workload": resource.child_name( "test-replica", "main" diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index abe127e82..252b492c3 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -43,6 +43,7 @@ # default, and matching by number would find the pd-sidecar on a disaggregated # pod rather than the engine behind it. _SERVING_LABEL = "modelplane_ai_serving" +_DEPLOYMENT_LABEL = "modelplane_ai_deployment" _METRICS_PORT = "http" _SCRAPE_INTERVAL = "15s" @@ -52,14 +53,18 @@ def _relabel_pod_identity() -> list[dict[str, Any]]: """Carry Modelplane's identity from the pod's labels onto every series. - A recorded series inherits these, so a MetricMapping's statements need no - labels of their own. + A series inherits these, so a MetricMapping needs no labels of its own. + compose-model-replica stamps them; a pod carrying none is one no deployment + owns. + + No model: a ModelReplica doesn't know which ModelService fronts it, and a + model name carries a slash, which a label value can't. Deployment is finer + grained anyway - a deployment serves one model, a model may have several. """ return [ {"source_labels": [f"__meta_kubernetes_pod_label_{src}"], "target_label": dst} for src, dst in ( ("modelplane_ai_deployment", "deployment"), - ("modelplane_ai_model", "model"), ("modelplane_ai_engine", "engine"), ("modelplane_ai_role", "role"), ) @@ -110,7 +115,15 @@ def _scrape_configs() -> list[dict[str, Any]]: "scrape_interval": _SCRAPE_INTERVAL, "kubernetes_sd_configs": [{"role": "pod"}], "relabel_configs": [ - {"source_labels": [f"__meta_kubernetes_pod_label_{_SERVING_LABEL}"], "action": "keep", "regex": "true"}, + # Every serving pod carries the deployment it belongs to, + # workers included: a worker holds GPUs, and the GPU series are + # the deployment's. Matched on presence, because the value is + # the deployment's name. + { + "source_labels": [f"__meta_kubernetes_pod_label_{_DEPLOYMENT_LABEL}"], + "action": "keep", + "regex": ".+", + }, { "source_labels": ["__meta_kubernetes_pod_container_port_name"], "action": "keep", @@ -290,7 +303,7 @@ def config( # A pod's identity is a resource attribute, where a metric processor # cannot reach it. Strip and merge the resources first, or the # aggregation below combines nothing. - "groupbyattrs/replicas": {"keys": ["cluster", "namespace", "deployment", "model", "engine", "role"]}, + "groupbyattrs/replicas": {"keys": ["cluster", "namespace", "deployment", "engine", "role"]}, "filter/modelplane": {"metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}}, "batch": {"timeout": "10s"}, } From 2350058571e44f3e9cf127759c05a780de1226ef Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 13:01:43 -0700 Subject: [PATCH 16/42] Name the cluster a series belongs to after the one an operator named Every series the collector exports carries a cluster attribute, and it carried the ServingStack's own name - generated, suffixed with a hash, and matching nothing anyone would query for. Take the composite label Crossplane stamps, which is the InferenceCluster's name. Found by reading the exported series on a live cluster: they arrived labelled local-serving-stack-d4206 rather than local. While there: the engines job selects the pods to scrape by the deployment label, and the substrate job dropped them by the serving label. Workers carry the first and not the second, so a worker that annotated itself for scraping would have been collected by both jobs under two names. Both partition on the same label now. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- docs/content/platform/telemetry.md | 26 ++++---- e2e/README.md | 21 ++++++- e2e/manifests/40-model-deployment.yaml | 13 ++++ e2e/run.sh | 59 +++++++++++++++++++ .../function/backends/base.py | 5 ++ .../function/backends/grove.py | 2 +- .../function/backends/llmd.py | 2 +- .../function/backends/native.py | 2 +- .../tests/test_backends.py | 4 +- .../compose-model-replica/tests/test_fn.py | 2 +- .../function/collector.py | 19 +++--- .../compose-serving-stack/function/fn.py | 18 +++++- .../compose-serving-stack/tests/test_fn.py | 22 +++++++ 13 files changed, 167 insertions(+), 28 deletions(-) diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 9416adf56..d7e8e2789 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -163,11 +163,12 @@ export to Prometheus and write recording rules there. ## Engines -vLLM and SGLang need no configuration. +Modelplane renames vLLM's and SGLang's own metrics for you, so neither needs a mapping. +SGLang needs one flag to publish them at all, below. -Any other OpenAI-compatible engine reports its top-line numbers with no configuration -either. The gateway measures those, not the engine, so `modelplane_frontend_*` and the token -counters work for an engine Modelplane has never seen. +Any other OpenAI-compatible engine reports its top-line numbers with no configuration. The +gateway measures those, not the engine, so `modelplane_frontend_*` works for an engine +Modelplane has never seen. To normalize that engine's own metrics as well, create a `MetricMapping`: @@ -197,8 +198,8 @@ Rename only where the measurements agree. Two engines' histograms under one name less than nothing if their buckets disagree, because a quantile over them is wrong rather than approximate. -One engine needs a flag. SGLang publishes `/metrics` only when it runs with -`--enable-metrics`, so add it to the engine args. vLLM needs nothing. +SGLang publishes `/metrics` only when it runs with `--enable-metrics`, so add that to its +engine args. vLLM needs nothing. ## Why engine latency and gateway latency differ @@ -270,20 +271,19 @@ engine. `DCGM_FI_DEV_GPU_UTIL` isn't renamed either, because it only tells you t wasn't idle; use `modelplane_gpu_compute_active_ratio` and `modelplane_gpu_tensor_active_ratio`. -You can also skip the rewrite for now. Modelplane provides compatibility recording rules -that rebuild the old names from the new ones, so your existing dashboards keep working -untouched: +You can also defer the rewrite. A recording rule rebuilds an old name from a new one, so a +dashboard keeps working untouched while you migrate it: ```yaml - record: vllm:time_to_first_token_seconds_bucket expr: label_replace(modelplane_request_ttft_seconds_bucket, - "model_name", "$1", "model", "(.*)") + "model_name", "$1", "deployment", "(.*)") ``` -Load it into the Prometheus you already run and nothing on a dashboard changes. It's one +Load that into the Prometheus you already run and nothing on the dashboard changes. It's one rule evaluation per metric over series your backend already holds, so it costs far less than -collecting everything twice. It covers the names in the table above, and it's meant to be -deleted once your panels use the new ones. +collecting everything twice. Write one per name in the table above, and delete them once the +panels use the new names. A series no statement renames doesn't leave the cluster. If a panel needs an engine's own name, write a `MetricMapping` that renames it onto the `modelplane_*` surface: a diff --git a/e2e/README.md b/e2e/README.md index d8bac4ee5..43561583b 100644 --- a/e2e/README.md +++ b/e2e/README.md @@ -47,8 +47,27 @@ server exposes both, so the pod goes Ready without a real model or GPU. | `ModelDeployment` → `ModelReplica` → `ModelEndpoint` → `ModelService` wiring | Real GPU drivers / CUDA (fake DRA devices only) | | DRA `ResourceClaim` → fake device binding (the real allocation path) | Multi-node / disaggregated (`PrefillDecode`) serving | | Serving-stack install on a real (BYO) workload cluster | Cloud provisioning (EKS/GKE/Nebius) | -| `InferenceGateway` + cross-cluster routing to the replica | | +| `InferenceGateway` + cross-cluster routing to the replica | Export to a real metrics backend (debug sink only) | | Status propagation and foreground-deletion ordering | | +| Telemetry: collector composed, engine discovered, series renamed and attributed | | + +### Telemetry + +`60-telemetry.yaml` creates a `TelemetryDestination` with the collector's +**debug** exporter, which prints what reached it to the collector's own log. So +the whole path is assertable with `kubectl logs` and needs no metrics backend: +service discovery finds the engine by the labels `compose-model-replica` stamps +on serving pods, the built-in `MetricMapping`s rename its series, the unit +conversion runs, and the identity comes off the pod. + +The mock engine serves `/metrics` with two real names — `vllm:num_requests_waiting` +and `DCGM_FI_DEV_FB_USED` — so `--verify` asserts they arrive as +`modelplane_requests_waiting` and `modelplane_gpu_memory_used_bytes`, that 1024 +MiB became 1073741824 bytes, that each carries its deployment, engine, role and +cluster, and that the engine's own `vllm:` names did *not* leave the cluster. + +It costs no extra wait: the destination is applied with every other manifest, so +the collector composes while the model is still rolling out. ### Why cloud provisioning cannot be tested here diff --git a/e2e/manifests/40-model-deployment.yaml b/e2e/manifests/40-model-deployment.yaml index 2b99dc72e..e32a3508c 100644 --- a/e2e/manifests/40-model-deployment.yaml +++ b/e2e/manifests/40-model-deployment.yaml @@ -53,6 +53,19 @@ spec: def do_GET(self): if self.path == "/health": self._s({"status": "ok"}) + elif self.path == "/metrics": + # Two real vLLM names, so the collector's built-in + # renames have something to act on: one gauge and + # one the fleet converts the unit of. + b = (b'# TYPE vllm:num_requests_waiting gauge\n' + b'vllm:num_requests_waiting 3\n' + b'# TYPE DCGM_FI_DEV_FB_USED gauge\n' + b'DCGM_FI_DEV_FB_USED{gpu="0"} 1024\n') + self.send_response(200) + self.send_header("content-type", "text/plain") + self.send_header("content-length", str(len(b))) + self.end_headers() + self.wfile.write(b) elif self.path.startswith("/v1/models"): self._s({"object": "list", "data": [{"id": served, "object": "model"}]}) else: diff --git a/e2e/run.sh b/e2e/run.sh index 84662ebbb..04a050f65 100644 --- a/e2e/run.sh +++ b/e2e/run.sh @@ -494,3 +494,62 @@ log "verify (usage record): ${usage}" cleanup_verify_pods log "End to end OK: ${base} authenticates callers, serves ${model} over OpenAI and Anthropic, rewrites the model, and meters it" + +# Telemetry. The TelemetryDestination went in with the rest of the manifests, so +# the collector composed while the model rolled out and there's nothing to wait +# for beyond the first scrape. Its debug sink prints what reached it to its own +# log, which is the whole path in one assertion: discovery found the engine by +# the labels compose-model-replica stamps, the built-in mappings renamed its +# series, the unit conversion ran, and the identity came off the pod. +log "Verifying the fleet's telemetry" +kubectl --context "$WLCTX" -n modelplane-system rollout status deploy/modelplane-collector --timeout=180s || { + echo "verify: the collector never rolled out on the workload cluster" >&2 + kubectl --context "$WLCTX" -n modelplane-system describe deploy/modelplane-collector >&2 || true + exit 1 +} + +# One log read per attempt, not one per assertion. The window is generous +# because the collector's config arrives by reconcile: on a fresh install it can +# roll out once against the destination and again once the MetricMappings land, +# and the restart the config change triggers starts its log over. In steady +# state the first read already has everything. +telemetry="" +for _ in $(seq 1 30); do + telemetry="$(kubectl --context "$WLCTX" -n modelplane-system logs deploy/modelplane-collector --tail=4000 2>/dev/null || true)" + case "$telemetry" in *modelplane_gpu_memory_used_bytes*) break ;; esac + sleep 10 +done + +missing="" +for want in \ + 'Name: modelplane_requests_waiting' \ + 'Name: modelplane_gpu_memory_used_bytes' \ + 'deployment: Str(mock-demo)' \ + 'engine: Str(mock)' \ + 'role: Str(Standalone)' \ + 'cluster: Str(local)'; do + case "$telemetry" in + *"$want"*) ;; + *) missing="$missing [$want]" ;; + esac +done + +# 1024 MiB as bytes. DCGM reports the framebuffer in MiB and the name says +# bytes, so a mapping that forgot the unit reads 1024 here instead. +case "$telemetry" in +*"Value: 1073741824"*) ;; +*) missing="$missing [DCGM_FI_DEV_FB_USED converted from MiB to bytes]" ;; +esac + +# Nothing the mappings didn't rename leaves a cluster, so the engine's own +# names must not appear downstream. +case "$telemetry" in +*"Name: vllm:num_requests_waiting"*) missing="$missing [vllm: names should not leave the cluster]" ;; +esac + +[ -z "$missing" ] || { + echo "verify: telemetry missing:$missing" >&2 + echo "$telemetry" | grep -E "Name: |-> (cluster|deployment|engine|role): |Value: " | tail -40 >&2 || true + exit 1 +} +log "Telemetry OK: the engine's series arrive renamed, converted, and attributed to its deployment" diff --git a/functions/compose-model-replica/function/backends/base.py b/functions/compose-model-replica/function/backends/base.py index 66860e86a..27f5a1784 100644 --- a/functions/compose-model-replica/function/backends/base.py +++ b/functions/compose-model-replica/function/backends/base.py @@ -228,6 +228,11 @@ def remote_namespace(replica: v1alpha1.ModelReplica) -> str: # the ModelEndpoint URLs, so it must not diverge between backends. ENGINE_PORT = 8000 +# The name given to that port. Named because the collector's engine scrape job +# selects on it: matching by number instead would find the pd-sidecar's port on +# a disaggregated pod rather than the engine behind it. +ENGINE_PORT_NAME = "http" + # Pod label carrying the serving identity (the replica name). The replica's one # shared Service selects on it, so every engine's serving pods carry it - a # Standalone pod, or a gang's leader (a LeaderWorkerSet leader or a Grove leader diff --git a/functions/compose-model-replica/function/backends/grove.py b/functions/compose-model-replica/function/backends/grove.py index 20925eb2c..500e2ecbe 100644 --- a/functions/compose-model-replica/function/backends/grove.py +++ b/functions/compose-model-replica/function/backends/grove.py @@ -106,7 +106,7 @@ def container(member: v1alpha1.Member, *, serving: bool) -> dict: if security_context: c["securityContext"] = security_context if serving: - c["ports"] = [{"containerPort": base.ENGINE_PORT}] + c["ports"] = [{"name": base.ENGINE_PORT_NAME, "containerPort": base.ENGINE_PORT}] c["readinessProbe"] = { "httpGet": {"path": "/health", "port": base.ENGINE_PORT}, "initialDelaySeconds": 30, diff --git a/functions/compose-model-replica/function/backends/llmd.py b/functions/compose-model-replica/function/backends/llmd.py index 8e4a82749..b1ad11348 100644 --- a/functions/compose-model-replica/function/backends/llmd.py +++ b/functions/compose-model-replica/function/backends/llmd.py @@ -100,7 +100,7 @@ def container(member: v1alpha1.Member, *, serving: bool) -> dict: env.extend(e.model_dump(exclude_none=True) for e in engine_container.env) c["env"] = env if serving: - c["ports"] = [{"containerPort": base.ENGINE_PORT}] + c["ports"] = [{"name": base.ENGINE_PORT_NAME, "containerPort": base.ENGINE_PORT}] c["readinessProbe"] = { "httpGet": {"path": "/health", "port": base.ENGINE_PORT}, "initialDelaySeconds": 30, diff --git a/functions/compose-model-replica/function/backends/native.py b/functions/compose-model-replica/function/backends/native.py index 52b1ffcae..9568d1e3e 100644 --- a/functions/compose-model-replica/function/backends/native.py +++ b/functions/compose-model-replica/function/backends/native.py @@ -61,7 +61,7 @@ def build( "name": "engine", "image": engine_container.image, "args": list(engine_container.args or []), - "ports": [{"containerPort": base.ENGINE_PORT}], + "ports": [{"name": base.ENGINE_PORT_NAME, "containerPort": base.ENGINE_PORT}], # vLLM tensor parallelism needs a large /dev/shm. "volumeMounts": [{"name": "dshm", "mountPath": "/dev/shm"}, *cache_volume_mounts], "readinessProbe": { diff --git a/functions/compose-model-replica/tests/test_backends.py b/functions/compose-model-replica/tests/test_backends.py index 12f1420ef..01fe61269 100644 --- a/functions/compose-model-replica/tests/test_backends.py +++ b/functions/compose-model-replica/tests/test_backends.py @@ -217,7 +217,7 @@ def _claim_template(count: int, *, replica: str = "r", engine: str = "main", rol "name": "engine", "image": "vllm/vllm-openai:latest", "args": ["--model=Qwen/Qwen3-0.6B"], - "ports": [{"containerPort": 8000}], + "ports": [{"name": "http", "containerPort": 8000}], "resources": {"claims": [{"name": "devices"}]}, "volumeMounts": [{"name": "dshm", "mountPath": "/dev/shm"}], "readinessProbe": { @@ -346,7 +346,7 @@ def _engine( if env is not None: c["env"] = env if serving: - c["ports"] = [{"containerPort": 8000}] + c["ports"] = [{"name": "http", "containerPort": 8000}] c["readinessProbe"] = { "httpGet": {"path": "/health", "port": 8000}, "initialDelaySeconds": 30, diff --git a/functions/compose-model-replica/tests/test_fn.py b/functions/compose-model-replica/tests/test_fn.py index 4fec047eb..ad6d524ec 100644 --- a/functions/compose-model-replica/tests/test_fn.py +++ b/functions/compose-model-replica/tests/test_fn.py @@ -212,7 +212,7 @@ async def test_compose(self) -> None: "name": "engine", "image": "vllm/vllm-openai:latest", "args": ["--model=Qwen/Qwen3-0.6B"], - "ports": [{"containerPort": 8000}], + "ports": [{"name": "http", "containerPort": 8000}], "resources": {"claims": [{"name": "devices"}]}, "volumeMounts": [ {"name": "dshm", "mountPath": "/dev/shm"}, diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 252b492c3..5d02c079d 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -38,12 +38,17 @@ IMAGE = "otel/opentelemetry-collector-contrib:0.161.0" -# The label Modelplane stamps on every serving pod, and the port name it gives -# the engine's metrics. Both matter: an engine container's port is unnamed by -# default, and matching by number would find the pd-sidecar on a disaggregated -# pod rather than the engine behind it. -_SERVING_LABEL = "modelplane_ai_serving" +# Pod labels as Kubernetes service discovery spells them: modelplane.ai/x +# arrives as __meta_kubernetes_pod_label_modelplane_ai_x. +# +# On every serving pod, workers included. The engines job selects on it and the +# substrate job drops on it, so the two partition the same set rather than +# leaving a worker to be collected by both. _DEPLOYMENT_LABEL = "modelplane_ai_deployment" + +# The name compose-model-replica gives the engine's port. Selecting by name +# rather than number is what keeps this off the pd-sidecar's port on a +# disaggregated pod, which would answer and serve the wrong thing. _METRICS_PORT = "http" _SCRAPE_INTERVAL = "15s" @@ -169,9 +174,9 @@ def _scrape_configs() -> list[dict[str, Any]]: "regex": ".+", }, { - "source_labels": [f"__meta_kubernetes_pod_label_{_SERVING_LABEL}"], + "source_labels": [f"__meta_kubernetes_pod_label_{_DEPLOYMENT_LABEL}"], "action": "drop", - "regex": "true", + "regex": ".+", }, # The rest of the same convention, not just the first line of # it. A pod that says scrape me generally also says where: the diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index b1aa3aa72..af2be852d 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -269,6 +269,22 @@ def _ensure_trailing_newline(cert: str) -> str: return cert if cert.endswith("\n") else cert + "\n" +# The label Crossplane stamps on a composed resource, naming the composite that +# claimed it. A ServingStack's own name is generated and carries a suffix, so +# this is what an operator calls the cluster. +_LABEL_COMPOSITE = "crossplane.io/composite" + + +def _cluster_name(xr: v1alpha1.ServingStack) -> str: + """The InferenceCluster this stack serves, as its operator named it. + + Every series the collector exports is stamped with this, and a metric + labelled with a generated name matches nothing an operator would query for. + """ + labels = (xr.metadata.labels if xr.metadata else None) or {} + return labels.get(_LABEL_COMPOSITE) or _name(xr.metadata) + + def _pc_name(xr: v1alpha1.ServingStack) -> str: """Derive the ProviderConfig name from the XR.""" return resource.child_name(_name(xr.metadata), "cluster") @@ -805,7 +821,7 @@ def compose_collector(self) -> list[str]: pc = _pc_name(self.xr) rendered: list[str] = [] for key, manifest, cel in collector.objects( - cluster=_name(self.xr.metadata), + cluster=_cluster_name(self.xr), mappings=mappings, sinks=list(dest.spec.sinks), extensions=dict(dest.spec.extensions or {}), diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index a57eb673d..bf4a46801 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -93,6 +93,28 @@ def _crds(filename: str) -> list[dict]: ] +class TestClusterName(unittest.TestCase): + """The name every exported series is stamped with.""" + + def _stack(self, labels: dict[str, str] | None) -> v1alpha1.ServingStack: + return v1alpha1.ServingStack( + metadata=metav1.ObjectMeta(name="local-serving-stack-d4206", labels=labels), + spec=v1alpha1.Spec( + cloud="Existing", + secrets=[v1alpha1.Secret(type="Kubeconfig", name="kube-secret", key="kubeconfig")], + gateway=v1alpha1.Gateway(hostname=_GATEWAY_HOSTNAME), + ), + ) + + def test_it_is_the_composite_an_operator_named(self) -> None: + """A ServingStack's own name is generated and carries a suffix.""" + self.assertEqual(fn._cluster_name(self._stack({"crossplane.io/composite": "local"})), "local") + + def test_it_falls_back_to_the_stack(self) -> None: + """Better a generated name on the series than none at all.""" + self.assertEqual(fn._cluster_name(self._stack(None)), "local-serving-stack-d4206") + + def _request(cloud: str, stack: str, observed: dict | None = None) -> fnv1.RunFunctionRequest: """Build a RunFunctionRequest for a test-backend ServingStack.""" return fnv1.RunFunctionRequest( From 4f011233d561fc8bd39e8e4c6213883e9de628d4 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 13:53:14 -0700 Subject: [PATCH 17/42] Assert the fleet's telemetry in the local end-to-end test Six bugs in the metrics path reached a green CI: the engine scrape job matched a label value that never existed, the gateway job asked Envoy for a path it 404s on, the substrate job ignored the port annotation beside the one it read, a value rewrite ran in a context that cannot reach values, framebuffer memory was renamed to bytes while holding MiB, and every series was labelled with a generated name rather than the cluster's. Unit tests, crossplane render and the collector's own config validation all passed throughout, because none of them scrapes anything. Assert it where something does. The destination goes in with the other manifests, so the collector composes while the model rolls out and --verify waits for nothing extra. Its debug sink prints what reached it to its own log, so the assertion is a log read and no metrics backend has to exist. The mock engine serves two real names, one of which needs a unit conversion, and --verify checks they arrive renamed, converted, carrying the identity off the pod, and that the engine's own names did not leave the cluster. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- e2e/manifests/60-telemetry.yaml | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) create mode 100644 e2e/manifests/60-telemetry.yaml diff --git a/e2e/manifests/60-telemetry.yaml b/e2e/manifests/60-telemetry.yaml new file mode 100644 index 000000000..97c076afd --- /dev/null +++ b/e2e/manifests/60-telemetry.yaml @@ -0,0 +1,16 @@ +# Telemetry for the fleet. Applied with everything else so the collector +# composes while the model rolls out, which costs --verify no extra wait. +# +# The debug exporter prints what reached it to the collector's own log, so the +# whole path - discovery, scrape, rename, unit conversion, identity - is +# assertable with kubectl logs and needs no metrics backend to run. +apiVersion: modelplane.ai/v1alpha1 +kind: TelemetryDestination +metadata: + name: default +spec: + sinks: + - name: debug + type: debug + config: + verbosity: detailed From 2d7265cbfc4d000645e2a56fdd592ba590d7c8ef Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 14:55:10 -0700 Subject: [PATCH 18/42] Scrape the GPU exporter the cluster came with Every modelplane_gpu_* series and the energy counter were empty on GKE. The substrate job finds a component by its prometheus.io/scrape annotation, and nothing annotates DCGM: GKE runs a managed exporter in gke-managed-system that carries only GKE's own component labels, and the NVIDIA GPU operator installs none at all there. Modelplane doesn't install one either - it reads whatever the cluster provides - so nothing was reading it. Give it a job of its own, keyed on the exporter's name rather than an annotation it doesn't carry. The two packagings spell it differently, gke-managed-dcgm-exporter and dcgm-exporter, and put it on different labels depending on who packaged it, so both labels are read and the match is on the name they share. Verified against a real L4 on GKE: all seven series arrive renamed, carrying the node and the card's UUID, and the conversions hold on real values - 7060 J from DCGM's millijoules, 16.89 W, 42 C. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 30 +++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 5d02c079d..e47d5367f 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -155,6 +155,36 @@ def _scrape_configs() -> list[dict[str, Any]]: *_annotated_path(), ], }, + { + "job_name": "modelplane-gpu", + "scrape_interval": _SCRAPE_INTERVAL, + "kubernetes_sd_configs": [{"role": "pod"}], + "relabel_configs": [ + # A job of its own because nobody annotates DCGM for scraping + # and Modelplane doesn't install it: GKE runs a managed one in + # gke-managed-system, the NVIDIA GPU operator installs its own + # elsewhere, and the substrate job sees neither. Without this + # every modelplane_gpu_* series is empty on a cloud that + # provides its own - which is every cloud. + # + # Matched on the name rather than an exact label, because the + # two spell it differently: gke-managed-dcgm-exporter and + # dcgm-exporter. Both labels are read, since which one carries + # the name depends on who packaged it. + { + "source_labels": [ + "__meta_kubernetes_pod_label_app_kubernetes_io_name", + "__meta_kubernetes_pod_label_app", + ], + "action": "keep", + "regex": ".*dcgm.*", + }, + {"source_labels": ["__meta_kubernetes_pod_container_port_name"], "action": "keep", "regex": "metrics"}, + # The node, because a GPU series belongs to hardware rather + # than to a deployment. DCGM names the card itself. + {"source_labels": ["__meta_kubernetes_pod_node_name"], "target_label": "node"}, + ], + }, { "job_name": "modelplane-substrate", "scrape_interval": _SUBSTRATE_INTERVAL, From 9c006ec024bc7e9abd79a4b804dc8da4fcc68dff Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 15:51:56 -0700 Subject: [PATCH 19/42] Keep a series' identity when the backend flattens it A Prometheus backend received every series stripped of everything that says what it measures: no cluster, no deployment, no engine, no role. Only job and instance survived. Modelplane's identity is held as resource attributes, because that is what the merge across replicas groups on, and the Prometheus remote-write exporter drops resource attributes unless asked not to. The debug exporter prints them, so every check up to this point showed them arriving. It took a Grafana panel legend rendering as "/" to notice that the one backend most people use sees none of it. Turn the conversion on for the exporters that flatten, under an operator's own config rather than over it, so a sink that sets it wins. Also share the DCGM selector between the job that keeps those pods and the job that has to leave them alone. A GPU operator's exporter usually does annotate itself for scraping, so the two drifting apart would collect it twice under two job names - which is the bug the gateway and engine drops already exist to prevent. A test now holds every such pair together. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 57 +++++++++++++++---- .../tests/test_collector.py | 18 ++++++ 2 files changed, 63 insertions(+), 12 deletions(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index e47d5367f..473547968 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -76,6 +76,25 @@ def _relabel_pod_identity() -> list[dict[str, Any]]: ] + [{"source_labels": ["__meta_kubernetes_namespace"], "target_label": "namespace"}] +def _dcgm_selector(action: str) -> dict[str, Any]: + """Match a GPU exporter, however it was packaged. + + The two spell it differently - gke-managed-dcgm-exporter and dcgm-exporter - + and put the name on different labels depending on who packaged it, so both + are read and the match is on what they share. One definition, used by the + job that keeps them and the job that has to leave them alone, because the + two drifting apart is how a series gets collected twice. + """ + return { + "source_labels": [ + "__meta_kubernetes_pod_label_app_kubernetes_io_name", + "__meta_kubernetes_pod_label_app", + ], + "action": action, + "regex": ".*dcgm.*", + } + + def _annotated_path() -> list[dict[str, Any]]: """Take the metrics path from the pod's own annotation, where it sets one.""" return [ @@ -171,14 +190,7 @@ def _scrape_configs() -> list[dict[str, Any]]: # two spell it differently: gke-managed-dcgm-exporter and # dcgm-exporter. Both labels are read, since which one carries # the name depends on who packaged it. - { - "source_labels": [ - "__meta_kubernetes_pod_label_app_kubernetes_io_name", - "__meta_kubernetes_pod_label_app", - ], - "action": "keep", - "regex": ".*dcgm.*", - }, + _dcgm_selector("keep"), {"source_labels": ["__meta_kubernetes_pod_container_port_name"], "action": "keep", "regex": "metrics"}, # The node, because a GPU series belongs to hardware rather # than to a deployment. DCGM names the card itself. @@ -195,9 +207,10 @@ def _scrape_configs() -> list[dict[str, Any]]: "action": "keep", "regex": "true", }, - # Both of these annotate themselves for scraping, and both have - # a job above that gives them their identity. Without the drops - # they are collected twice, under two job names. + # Everything a job above already names. Each of these can + # annotate itself for scraping - the gateway does, and a + # GPU operator's DCGM usually does - and collecting one here + # as well would carry it twice under two job names. { "source_labels": ["__meta_kubernetes_pod_label_gateway_envoyproxy_io_owning_gateway_name"], "action": "drop", @@ -208,6 +221,7 @@ def _scrape_configs() -> list[dict[str, Any]]: "action": "drop", "regex": ".+", }, + _dcgm_selector("drop"), # The rest of the same convention, not just the first line of # it. A pod that says scrape me generally also says where: the # cert-manager webhook declares 10250 first and annotates 9402, @@ -268,6 +282,25 @@ def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: return {"metric_statements": blocks} +# Exporters that flatten a series into labels, losing anything held as a +# resource attribute unless told otherwise. Modelplane's identity - the +# cluster, deployment, engine and role a series belongs to - is all held there, +# because that is what the merge across replicas groups on, so without this a +# Prometheus backend receives every series stripped of everything that says +# what it measures. +# +# Applied under an operator's own config rather than over it: this is a default, +# and a sink that sets it wins. +_SINK_DEFAULTS = { + "prometheusremotewrite": {"resource_to_telemetry_conversion": {"enabled": True}}, + "prometheus": {"resource_to_telemetry_conversion": {"enabled": True}}, +} + + +def _sink_defaults(exporter: str) -> dict[str, Any]: + return {k: dict(v) for k, v in _SINK_DEFAULTS.get(exporter, {}).items()} + + def _credential_path(sink: tdv1alpha1.Sink, key: str) -> str: return f"{_CREDENTIALS_DIR}/{sink.name}/{key}" @@ -305,7 +338,7 @@ def exporters(sinks: list[tdv1alpha1.Sink]) -> dict[str, Any]: """ rendered: dict[str, Any] = {} for sink in sinks: - cfg: dict[str, Any] = dict(sink.config or {}) + cfg: dict[str, Any] = _sink_defaults(sink.type) | dict(sink.config or {}) if sink.endpoint: cfg["endpoint"] = sink.endpoint if sink.auth and sink.auth.bearerTokenKey: diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index d5abd80fe..c250b82c0 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -78,6 +78,24 @@ def test_cluster_is_stamped_here(self) -> None: attrs = _config()["processors"]["resource/cluster"]["attributes"] self.assertEqual(attrs, [{"key": "cluster", "value": "prod-us-east", "action": "upsert"}]) + def test_the_jobs_cover_disjoint_pods(self) -> None: + """A pod two jobs both collect arrives twice, under two job names.""" + jobs = { + j["job_name"]: j["relabel_configs"] + for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"] + } + substrate = jobs["modelplane-substrate"] + + def predicate(rules: list[dict], action: str) -> set[tuple]: + return {(tuple(r["source_labels"]), r["regex"]) for r in rules if r.get("action") == action} + + # Everything another job keeps, the substrate job drops on the same terms. + for job in ("modelplane-engines", "modelplane-gateway", "modelplane-gpu"): + for kept in predicate(jobs[job], "keep"): + if kept[0] == ("__meta_kubernetes_pod_container_port_name",): + continue # a port filter, not a pod filter + self.assertIn(kept, predicate(substrate, "drop"), f"{job} keeps {kept}, substrate does not drop it") + def test_engine_scrape_selects_the_port_by_name(self) -> None: """Matching by number would find the pd-sidecar on a disaggregated pod.""" jobs = {j["job_name"]: j for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"]} From dced0f8172846cdee079aa74b9b8f09358b29032 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Wed, 30 Sep 2026 22:01:47 -0700 Subject: [PATCH 20/42] Carry only the identity a series is attributed to Service discovery attaches the pod's name, uid and replicaset to every scrape, and the scrape its address, and all of it lands on the resource. Keeping it does two kinds of damage. The merge across replicas groups on the resource, so a pod name there holds every replica in a resource of its own and nothing merges. And an exporter that flattens resources into labels then publishes a pod label - the one the design rules out by name, because a rolling update mints a fresh one on every deploy and a billing backend counts it active for half an hour after it dies. Keep an allowlist instead, ahead of the merge: the cluster, namespace, deployment, engine and role a series belongs to, plus the node for the GPU series, which belong to hardware rather than to a deployment. An allowlist rather than a list of what to drop, because what discovery attaches grows. Verified on GKE with two replicas on two nodes: one series, carrying those attributes and nothing else, where before there was one per pod carrying k8s_pod_name, k8s_pod_uid, k8s_replicaset_name and the pod's address. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 35 ++++++++++++++++++- .../tests/test_collector.py | 11 ++++++ 2 files changed, 45 insertions(+), 1 deletion(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 473547968..2ffd613f2 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -51,6 +51,19 @@ # disaggregated pod, which would answer and serve the wrong thing. _METRICS_PORT = "http" +# What a series is attributed to, and the only resource attributes that survive +# to the exporter. The merge across replicas groups on exactly these, so a +# series differing in nothing else is one series. +# +# node is here for the GPU job, whose series belong to hardware rather than to +# a deployment; an engine's series carry no node, which is what lets replicas on +# different nodes merge. +_IDENTITY = ("cluster", "namespace", "deployment", "engine", "role", "node") + +# OTTL quotes strings with double quotes; a Python list renders single ones and +# the collector refuses to start on it. +_IDENTITY_OTTL = ", ".join(f'"{k}"' for k in sorted(_IDENTITY)) + _SCRAPE_INTERVAL = "15s" _SUBSTRATE_INTERVAL = "30s" @@ -367,16 +380,36 @@ def config( # cluster is stamped here rather than downstream: one receiver on the # control plane sees a merged stream and cannot tell senders apart. "resource/cluster": {"attributes": [{"key": "cluster", "value": cluster, "action": "upsert"}]}, + # Everything a replica carries that is the replica's rather than the + # deployment's. Discovery attaches the pod's name, uid and replicaset, + # and the scrape its address, and all of it lands on the resource - + # where it does two kinds of damage. + # + # The merge below groups on the resource, so a pod name on it keeps + # every replica in a resource of its own and nothing merges. And an + # exporter that flattens resources into labels then publishes a pod + # label, which a rolling update mints afresh on every deploy and a + # billing backend counts as active for half an hour after it dies. + # + # An allowlist rather than a list of what to drop: what discovery + # attaches grows, and a series carrying something nobody chose is the + # failure this prevents. + "transform/identity": { + "metric_statements": [ + {"context": "resource", "statements": [f"keep_keys(resource.attributes, [{_IDENTITY_OTTL}])"]} + ] + }, "transform/modelplane": _transform(mappings), # A pod's identity is a resource attribute, where a metric processor # cannot reach it. Strip and merge the resources first, or the # aggregation below combines nothing. - "groupbyattrs/replicas": {"keys": ["cluster", "namespace", "deployment", "engine", "role"]}, + "groupbyattrs/replicas": {"keys": list(_IDENTITY)}, "filter/modelplane": {"metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}}, "batch": {"timeout": "10s"}, } pipeline = [ "resource/cluster", + "transform/identity", "transform/modelplane", "groupbyattrs/replicas", "filter/modelplane", diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index c250b82c0..ac27eb2eb 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -96,6 +96,17 @@ def predicate(rules: list[dict], action: str) -> set[tuple]: continue # a port filter, not a pod filter self.assertIn(kept, predicate(substrate, "drop"), f"{job} keeps {kept}, substrate does not drop it") + def test_only_the_identity_survives_to_the_exporter(self) -> None: + """Discovery attaches the pod's name and uid; neither is the deployment's.""" + blocks = _config()["processors"]["transform/identity"]["metric_statements"] + statement = blocks[0]["statements"][0] + # OTTL quotes with double quotes. A Python list renders single ones and + # the collector refuses to start, which a unit test on shape won't catch. + self.assertNotIn("'", statement) + self.assertIn('keep_keys(resource.attributes, ["cluster"', statement) + pipeline = _config()["service"]["pipelines"]["metrics"]["processors"] + self.assertLess(pipeline.index("transform/identity"), pipeline.index("groupbyattrs/replicas")) + def test_engine_scrape_selects_the_port_by_name(self) -> None: """Matching by number would find the pd-sidecar on a disaggregated pod.""" jobs = {j["job_name"]: j for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"]} From 14818784ccdfd311444d926d565546e969a0707e Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 09:04:52 -0700 Subject: [PATCH 21/42] Tell a deployment's replicas apart Two replicas of one deployment reported one number, and it was one replica's. Measured on GKE: two engines holding 300 and 4640 prefix-cache lookups arrived as a single series reading 300, then 4640, never 4940. No duplicate-sample error, no gap in the graph - just a number six per cent of the truth that looks entirely plausible on a dashboard. Their series were identical, so they collided. The design says replicas are merged before they leave the cluster, and groupbyattrs does not merge them: a scrape of one replica is one batch, so there is never more than one replica present to merge. Nor could a collector sum them honestly - two cumulative readings taken at different moments add up to more traffic than happened. Carry the replica instead, and let the backend combine them in the query, where the arithmetic is right. The index rather than the pod: bounded by the replica count, and it survives a restart and a rolling update, where a pod name is minted afresh each time and a billing backend counts it live for half an hour after it dies. acrossReplicas stays, and now says what it does - which combination is the right one for this metric - rather than naming an aggregation nothing performs. Also scrape the AI gateway's ext-proc. The GenAI metrics, which are the ones a caller's experience is measured by, come from a sidecar on its own admin port rather than from the proxy, so the gateway needs a second job at a second path. The sidecar is an initContainer with restartPolicy Always, which is why it reads as absent from a pod's containers. Verified on GKE: two series, replica=0 at 4640 and replica=1 at 300, summing to the 4940 the engines actually served. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 28 +++++++++- docs/content/platform/telemetry.md | 18 +++++- .../function/backends/base.py | 16 +++++- .../function/collector.py | 55 ++++++++++++++++--- .../function/stacks/metrics.py | 22 +++++++- .../tests/test_collector.py | 8 ++- schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 6 ++ 8 files changed, 138 insertions(+), 17 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index b90ed310a..78594b8d1 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -53,7 +53,7 @@ spec: wrong rather than approximate. items: type: object - required: [from, to] + required: [from, to, acrossReplicas] properties: from: type: string @@ -78,6 +78,32 @@ spec: What Modelplane calls it. Only modelplane_* leaves a cluster, so a metric with no name here is one nobody downstream can read. + acrossReplicas: + type: string + enum: [Sum, Mean, Max] + description: >- + How this metric combines over a deployment's replicas. + + Each replica publishes its own series, told apart by + the replica label, and a query over a deployment + combines them. This says which combination is the right + one: Sum for anything counted - requests, tokens, + joules, a queue's depth. Mean for a ratio, where + summing reads two replicas at half capacity as one at + full. Max for a saturation figure an alert fires on, + where a mean hides the replica in trouble. + + Modelplane does not combine them in the collector. A + scrape of one replica is one batch, so a collector that + added them up would be adding readings taken at + different moments, and two readings of one cumulative + counter sum to twice the traffic that happened. The + backend holds every replica's series and combines them + at query time, where the arithmetic is right. + + Required, with no default, because the wrong + combination is silent: a deployment reports a number + that looks entirely plausible. fromUnit: type: string enum: [Millijoules, Mebibytes, Milliseconds, Nanoseconds] diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index d7e8e2789..f1467ba59 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -20,7 +20,18 @@ deployments change. ## What you get Every series carries `cluster`. A series about a deployment also carries `deployment`, -`namespace`, `engine`, and `role`. Some of what you can read: +`replica`, `namespace`, `engine`, and `role`. + +Each replica publishes its own series. Combine them in the query, the way the metric's +`acrossReplicas` says: `sum by (deployment)` for anything counted, `avg by (deployment)` +for a ratio, `max by (deployment)` for a saturation figure an alert fires on. The +collector doesn't add them up for you, because a scrape of one replica is one batch, and +adding readings taken at different moments is not the traffic that happened. + +The replica is an index, not a pod. It's bounded by the replica count and it survives a +restart and a rolling update, so the series count doesn't grow every time you deploy. + +Some of what you can read: | Metric | Means | | --- | --- | @@ -237,8 +248,9 @@ against Modelplane's Prometheus stops being read by anything, because the operat the stack. **Rewrite your dashboard queries.** Names change, and so do the labels: group by -`deployment` rather than `model_name`, pod labels are gone because replicas are summed -before they leave the cluster, and every series now carries `cluster`. +`deployment` rather than `model_name`, there's no pod label, and every series carries +`cluster` and `replica`. A panel that showed one engine now shows one replica, so wrap it +in `sum by (deployment)` or the aggregation that metric's `acrossReplicas` names. | Was | Is | | --- | --- | diff --git a/functions/compose-model-replica/function/backends/base.py b/functions/compose-model-replica/function/backends/base.py index 27f5a1784..88d19a47c 100644 --- a/functions/compose-model-replica/function/backends/base.py +++ b/functions/compose-model-replica/function/backends/base.py @@ -252,10 +252,15 @@ def remote_namespace(replica: v1alpha1.ModelReplica) -> str: # merged across replicas keeps the identity all of them share. Metrics are the # only reason these exist; nothing selects on them. LABEL_DEPLOYMENT = "modelplane.ai/deployment" +LABEL_REPLICA = "modelplane.ai/replica" LABEL_ENGINE = "modelplane.ai/engine" LABEL_ROLE = "modelplane.ai/role" +# Set on the ModelReplica by the composite that scheduled it. +_LABEL_REPLICA_INDEX = "modelplane.ai/replica-index" + + def telemetry_labels( replica: v1alpha1.ModelReplica, engine: v1alpha1.Engine, @@ -269,9 +274,16 @@ def telemetry_labels( label value can't. """ labels: dict[str, str] = {} - deployment = (replica.metadata.labels if replica.metadata else None) or {} - if name := deployment.get(LABEL_DEPLOYMENT): + own = (replica.metadata.labels if replica.metadata else None) or {} + if name := own.get(LABEL_DEPLOYMENT): labels[LABEL_DEPLOYMENT] = name + # Which replica of that deployment. Two replicas reporting the same metric + # need something to tell them apart or they are one series downstream and + # one of them is simply lost. The index rather than the pod: it is bounded + # by the replica count, and it survives a restart and a rolling update, + # where a pod name is minted afresh each time. + if index := own.get(_LABEL_REPLICA_INDEX): + labels[LABEL_REPLICA] = index if engine.name: labels[LABEL_ENGINE] = engine.name if member.role: diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 2ffd613f2..49e9ce097 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -51,6 +51,11 @@ # disaggregated pod, which would answer and serve the wrong thing. _METRICS_PORT = "http" +# The port the AI gateway's ext-proc sidecar serves its own metrics on, which +# is where the GenAI semantic-convention series live. Envoy's admin port +# carries Envoy's own statistics and none of these. +_GENAI_PORT = "aigw-admin" + # What a series is attributed to, and the only resource attributes that survive # to the exporter. The merge across replicas groups on exactly these, so a # series differing in nothing else is one series. @@ -58,7 +63,7 @@ # node is here for the GPU job, whose series belong to hardware rather than to # a deployment; an engine's series carry no node, which is what lets replicas on # different nodes merge. -_IDENTITY = ("cluster", "namespace", "deployment", "engine", "role", "node") +_IDENTITY = ("cluster", "namespace", "deployment", "replica", "engine", "role", "node") # OTTL quotes strings with double quotes; a Python list renders single ones and # the collector refuses to start on it. @@ -83,6 +88,7 @@ def _relabel_pod_identity() -> list[dict[str, Any]]: {"source_labels": [f"__meta_kubernetes_pod_label_{src}"], "target_label": dst} for src, dst in ( ("modelplane_ai_deployment", "deployment"), + ("modelplane_ai_replica", "replica"), ("modelplane_ai_engine", "engine"), ("modelplane_ai_role", "role"), ) @@ -142,9 +148,13 @@ def _annotated_port() -> list[dict[str, Any]]: def _scrape_configs() -> list[dict[str, Any]]: """What to scrape on an inference cluster. - Three jobs over disjoint sets of pods, so nothing is scraped twice: the - engines Modelplane runs, the gateways in front of them, and everything else - the serving stack installs. + Jobs over disjoint sets of pods, so nothing is scraped twice: the engines + Modelplane runs, the gateways in front of them, the GPU exporter the + cluster came with, and everything else the serving stack installs. + + The front door needs two of them. Envoy publishes its own statistics on its + admin port, and the GenAI metrics a caller's experience is measured by come + from the AI gateway's ext-proc on a different port, at a different path. """ return [ { @@ -187,6 +197,32 @@ def _scrape_configs() -> list[dict[str, Any]]: *_annotated_path(), ], }, + { + # The GenAI metrics, which are the SLO ones: what a caller waited, + # measured the same way whatever engine served it. + # + # They come from the AI gateway's ext-proc, not from the proxy. It + # runs as a native sidecar - an initContainer with restartPolicy + # Always - so it is easy to miss when reading the pod, and it + # serves its own admin port rather than Envoy's. A job of its own + # because the two ports want different paths: the proxy publishes + # at the path its annotation names, the ext-proc at /metrics. + "job_name": "modelplane-gateway-genai", + "scrape_interval": _SCRAPE_INTERVAL, + "kubernetes_sd_configs": [{"role": "pod"}], + "relabel_configs": [ + { + "source_labels": ["__meta_kubernetes_pod_label_gateway_envoyproxy_io_owning_gateway_name"], + "action": "keep", + "regex": ".+", + }, + { + "source_labels": ["__meta_kubernetes_pod_container_port_name"], + "action": "keep", + "regex": _GENAI_PORT, + }, + ], + }, { "job_name": "modelplane-gpu", "scrape_interval": _SCRAPE_INTERVAL, @@ -400,9 +436,14 @@ def config( ] }, "transform/modelplane": _transform(mappings), - # A pod's identity is a resource attribute, where a metric processor - # cannot reach it. Strip and merge the resources first, or the - # aggregation below combines nothing. + # Lifts the identity onto the resource, where an exporter that flattens + # a series into labels will find it. It does not merge a deployment's + # replicas: each is scraped separately, so each is its own batch, and + # there is never more than one replica here to merge. They stay + # separate series, told apart by the replica label, and a query over + # the deployment combines them - which is the only place the + # arithmetic can be right, because adding two cumulative readings + # taken at different moments is not the traffic that happened. "groupbyattrs/replicas": {"keys": list(_IDENTITY)}, "filter/modelplane": {"metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}}, "batch": {"timeout": "10s"}, diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index 237236fc7..f856cced1 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -93,6 +93,25 @@ } +# What describes a piece of hardware or a replica rather than a fleet. A ratio +# summed reads two replicas at half their KV cache as one at capacity, and a +# temperature summed is not a temperature at all. +# +# Everything else here is counted - requests, tokens, joules, watts drawn, +# queue depth - and a histogram can only be summed, which merges its buckets. +_MEAN = { + "modelplane_kv_cache_utilization_ratio", + "modelplane_gpu_compute_active_ratio", + "modelplane_gpu_tensor_active_ratio", + "modelplane_gpu_memory_bandwidth_ratio", + "modelplane_gpu_temperature_celsius", +} + + +def _across_replicas(target: str) -> str: + return "Mean" if target in _MEAN else "Sum" + + def _mapping( name: str, pairs: dict[str, str], @@ -104,7 +123,8 @@ def _mapping( spec=v1alpha1.Spec( metrics=[ v1alpha1.Metric.model_validate( - {"from": src, "to": dst} | ({"fromUnit": units[src]} if src in units else {}) + {"from": src, "to": dst, "acrossReplicas": _across_replicas(dst)} + | ({"fromUnit": units[src]} if src in units else {}) ) for src, dst in pairs.items() ] diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index ac27eb2eb..b3f61103a 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -152,10 +152,14 @@ def test_every_unit_the_api_offers_has_a_conversion(self) -> None: def test_a_metric_name_cannot_end_the_comparison_early(self) -> None: """A quote in `from` would rename whatever the rest of the line matched.""" with self.assertRaises(ValidationError): - mmv1alpha1.Metric.model_validate({"from": 'x" or true or name == "y', "to": "modelplane_x"}) + mmv1alpha1.Metric.model_validate( + {"from": 'x" or true or name == "y', "to": "modelplane_x", "acrossReplicas": "Sum"} + ) for mapping in stacks.BUILTIN_MAPPINGS: for m in mapping.spec.metrics: - round_tripped = mmv1alpha1.Metric.model_validate({"from": m.from_, "to": m.to}) + round_tripped = mmv1alpha1.Metric.model_validate( + {"from": m.from_, "to": m.to, "acrossReplicas": m.acrossReplicas} + ) self.assertEqual(round_tripped.from_, m.from_) def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: diff --git a/schemas/.lock.json b/schemas/.lock.json index ef60257b9..3dadd3b04 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "2b2af1570db437b60cb3574aa86e1ed1449a410ef6869629b8ba2f9fccabacb4", + "fs://apis": "6af56fb94787e27c4d192b4186d553704b01692e39bf593019a6cfb10f6f0234", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index e108026ae..21e7a9670 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -43,6 +43,12 @@ class Crossplane(BaseModel): class Metric(BaseModel): + acrossReplicas: Literal['Sum', 'Mean', 'Max'] + """ + How a deployment's replicas combine into one series. Every replica reports this metric for itself, and what leaves the cluster is one series for the deployment, so something has to say what the deployment's value is. + Sum for anything counted: requests, tokens, joules, a queue's depth. Mean for a ratio, where summing would read two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where the mean hides the replica that is actually in trouble. + Required, with no default, because the wrong answer here is silent: a deployment reports a number that looks entirely plausible and is one replica's. A histogram can only be summed, which merges its buckets. + """ from_: constr(pattern=r'^[a-zA-Z_:][a-zA-Z0-9_:]*$', max_length=255) = Field( ..., alias='from' ) From 5400e3d327697a10519c65a971b442fc248777bc Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 09:26:27 -0700 Subject: [PATCH 22/42] Fold several metrics into one, and take a histogram's count A rename carries one metric to one name, which leaves out the series a fleet most wants: tokens in and out under one name told apart by direction, responses told apart by the reason they ended, requests told apart by status. Each needs several of a component's metrics to become one of Modelplane's, or a label the component already writes to be carried over and its vocabulary translated. `labels` does both. A fixed `value` is what tells two folded mappings apart - each renames its own source and stamps its own value. A `from` carries a label the component already emits, and `values` puts its vocabulary into Modelplane's, leaving anything it doesn't name as the component wrote it. The component's own label is dropped afterwards, or the series says the same thing twice and costs twice the cardinality. `part` takes a histogram's count or sum as a counter beside it, which is how a request count comes out of a duration histogram. The histogram carries on unchanged. The statements run before the rename, while the series still answers to the name the mapping selected on: after it, two folded mappings share one name and nothing could tell them apart. Neither is used by a built-in yet. Both are what the metric surface in the design needs next, and the shape is better settled against the kind than alongside it. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 64 +++++++++++++++++++ .../function/collector.py | 40 +++++++++++- schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 39 ++++++++++- 4 files changed, 139 insertions(+), 6 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index 78594b8d1..8130230c9 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -104,6 +104,70 @@ spec: Required, with no default, because the wrong combination is silent: a deployment reports a number that looks entirely plausible. + part: + type: string + enum: [Count, Sum] + description: >- + Take a part of a histogram as a counter of its own, + rather than the histogram itself. Count is how many + observations it holds, which is a request count where + the histogram measures request duration. Sum is their + total. + + The histogram carries on unchanged under its own name. + This adds a series beside it. + labels: + type: array + maxItems: 16 + description: >- + Labels to set on the series, for folding several + metrics into one that a label tells apart - tokens in + and out under one name with a direction, responses + under one name with the reason they ended. + + Two mappings writing the same `to` with a different + fixed value is how the fold is expressed: each renames + its own source and stamps its own value. + items: + type: object + required: [name] + x-kubernetes-validations: + - rule: "has(self.value) != has(self.from)" + message: set either value, for a fixed label, or from, to carry one the component already emits. + properties: + name: + type: string + maxLength: 63 + pattern: '^[a-zA-Z_][a-zA-Z0-9_]*$' + description: The label to set. + value: + type: string + maxLength: 253 + description: >- + A fixed value, the same on every series this + mapping produces. This is what tells two folded + metrics apart. + from: + type: string + maxLength: 63 + pattern: '^[a-zA-Z_][a-zA-Z0-9_]*$' + description: >- + A label the component already emits, carried onto + the new name and dropped from the series under + its old one. + values: + type: object + maxProperties: 32 + additionalProperties: + type: string + maxLength: 253 + description: >- + What each of that label's values becomes, for + putting an engine's own vocabulary into + Modelplane's. A value with no entry here is left + as the component wrote it. + + Only meaningful alongside `from`. fromUnit: type: string enum: [Millijoules, Mebibytes, Milliseconds, Nanoseconds] diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 49e9ce097..274b3474d 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -289,6 +289,11 @@ def _scrape_configs() -> list[dict[str, Any]]: # float: 1e-09 is not an OTTL literal. Paths carry their context because the # collector rewrites bare ones and asks the author to stop; Modelplane is the # author here, so nobody's stored MetricMapping has to change. +# Taking a part of a histogram is a function that mints a new metric beside it, +# named for the part. The rename then applies to that. +_PART_FUNCTION = {"Count": "extract_count_metric(true)", "Sum": "extract_sum_metric(true)"} +_PART_SUFFIX = {"Count": "_count", "Sum": "_sum"} + _UNIT_CONVERSION = { "Millijoules": "datapoint.value_double / 1000", "Milliseconds": "datapoint.value_double / 1000", @@ -309,13 +314,44 @@ def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], lis metric: list[str] = [] for mapping in mappings: for m in mapping.spec.metrics: + source = m.from_ + if m.part: + # Lift the part out first, under the name the extraction gives + # it, and rename that. The histogram carries on untouched. + suffix = _PART_SUFFIX[m.part] + metric.append(f'{_PART_FUNCTION[m.part]} where metric.name == "{source}"') + source = f"{source}{suffix}" if m.fromUnit: conversion = _UNIT_CONVERSION[m.fromUnit] - datapoint.append(f'set(datapoint.value_double, {conversion}) where metric.name == "{m.from_}"') - metric.append(f'set(metric.name, "{m.to}") where metric.name == "{m.from_}"') + datapoint.append(f'set(datapoint.value_double, {conversion}) where metric.name == "{source}"') + # Labels before the rename, while the series still answers to the + # name this mapping selected on. After it, two folded mappings share + # one name and a statement could no longer tell them apart. + datapoint.extend(_label_statements(source, m.labels or [])) + metric.append(f'set(metric.name, "{m.to}") where metric.name == "{source}"') return datapoint, metric +def _label_statements(source: str, labels: list[Any]) -> list[str]: + """Set this mapping's labels on the datapoints of one source metric.""" + out: list[str] = [] + for label in labels: + if label.value is not None: + out.append(f'set(datapoint.attributes["{label.name}"], "{label.value}") where metric.name == "{source}"') + continue + carried = f'datapoint.attributes["{label.from_}"]' + out.append(f'set(datapoint.attributes["{label.name}"], {carried}) where metric.name == "{source}"') + for old, new in sorted((label.values or {}).items()): + out.append( + f'set(datapoint.attributes["{label.name}"], "{new}") ' + f'where metric.name == "{source}" and {carried} == "{old}"' + ) + # The component's own name for it goes, or the series carries the same + # fact twice under two labels and costs twice the cardinality. + out.append(f'delete_key(datapoint.attributes, "{label.from_}") where metric.name == "{source}"') + return out + + def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: """Unit conversions first, then every rename. diff --git a/schemas/.lock.json b/schemas/.lock.json index 3dadd3b04..ff98a6809 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "6af56fb94787e27c4d192b4186d553704b01692e39bf593019a6cfb10f6f0234", + "fs://apis": "c0d16165e05888fb82f0d605a67b5993d178c4b2305143a84d988782d257a53c", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index 21e7a9670..8dd7442df 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -42,12 +42,35 @@ class Crossplane(BaseModel): resourceRefs: list[ResourceRef] | None = None +class Label(BaseModel): + from_: constr(pattern=r'^[a-zA-Z_][a-zA-Z0-9_]*$', max_length=63) | None = Field( + None, alias='from' + ) + """ + A label the component already emits, carried onto the new name and dropped from the series under its old one. + """ + name: constr(pattern=r'^[a-zA-Z_][a-zA-Z0-9_]*$', max_length=63) + """ + The label to set. + """ + value: constr(max_length=253) | None = None + """ + A fixed value, the same on every series this mapping produces. This is what tells two folded metrics apart. + """ + values: dict[str, constr(max_length=253)] | None = Field(None, max_length=32) + """ + What each of that label's values becomes, for putting an engine's own vocabulary into Modelplane's. A value with no entry here is left as the component wrote it. + Only meaningful alongside `from`. + """ + + class Metric(BaseModel): acrossReplicas: Literal['Sum', 'Mean', 'Max'] """ - How a deployment's replicas combine into one series. Every replica reports this metric for itself, and what leaves the cluster is one series for the deployment, so something has to say what the deployment's value is. - Sum for anything counted: requests, tokens, joules, a queue's depth. Mean for a ratio, where summing would read two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where the mean hides the replica that is actually in trouble. - Required, with no default, because the wrong answer here is silent: a deployment reports a number that looks entirely plausible and is one replica's. A histogram can only be summed, which merges its buckets. + How this metric combines over a deployment's replicas. + Each replica publishes its own series, told apart by the replica label, and a query over a deployment combines them. This says which combination is the right one: Sum for anything counted - requests, tokens, joules, a queue's depth. Mean for a ratio, where summing reads two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where a mean hides the replica in trouble. + Modelplane does not combine them in the collector. A scrape of one replica is one batch, so a collector that added them up would be adding readings taken at different moments, and two readings of one cumulative counter sum to twice the traffic that happened. The backend holds every replica's series and combines them at query time, where the arithmetic is right. + Required, with no default, because the wrong combination is silent: a deployment reports a number that looks entirely plausible. """ from_: constr(pattern=r'^[a-zA-Z_:][a-zA-Z0-9_:]*$', max_length=255) = Field( ..., alias='from' @@ -63,6 +86,16 @@ class Metric(BaseModel): What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, and mebibytes multiplied out to bytes. Say it whenever the source disagrees with the target, even where the factor looks obvious. A name ending in _bytes that holds mebibytes is the kind of thing nobody notices until a capacity review, and stating the source unit is what makes the conversion happen at all. """ + labels: list[Label] | None = Field(None, max_length=16) + """ + Labels to set on the series, for folding several metrics into one that a label tells apart - tokens in and out under one name with a direction, responses under one name with the reason they ended. + Two mappings writing the same `to` with a different fixed value is how the fold is expressed: each renames its own source and stamps its own value. + """ + part: Literal['Count', 'Sum'] | None = None + """ + Take a part of a histogram as a counter of its own, rather than the histogram itself. Count is how many observations it holds, which is a request count where the histogram measures request duration. Sum is their total. + The histogram carries on unchanged under its own name. This adds a series beside it. + """ to: constr(pattern=r'^modelplane_[a-z0-9_]*[a-z0-9]$', max_length=255) """ What Modelplane calls it. Only modelplane_* leaves a cluster, so a metric with no name here is one nobody downstream can read. From 00755044bf7143115e2229c1632282f7c7c6f61a Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 09:26:27 -0700 Subject: [PATCH 23/42] Resolve a cluster gateway's name where kube-dns is the resolver Every request through an InferenceGateway to a model on a GKE cluster was refused with "no healthy upstream". The gateway's Envoy held zero healthy members for the backend and zero TLS errors, which together mean it never had an address to try. A cluster gateway's name is published as a selectorless headless Service and an EndpointSlice. kube-dns builds its records from the Endpoints API and was never taught the EndpointSlice one, so the name answers NXDOMAIN there - and kube-dns is what GKE runs, still, by default. CoreDNS reads the slice, which is why kind serves this path and CI has never seen it. Publish the Endpoints the slice supersedes alongside it. Deprecated since 1.33 and still the only thing kube-dns reads; one more object against every cross-cluster request failing. Verified on GKE: the name went from NXDOMAIN to the gateway's address, and a request through the fleet gateway returned 200 on the first attempt. That also unblocked the metric measuring what a caller waited - gen_ai_server_request_duration_seconds exists only once a request reaches the AI gateway's ext-proc, so with every request refused there was nothing to scrape. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-inference-gateway/function/fn.py | 28 ++++++++++++++++++- .../tests/test_fn.py | 18 ++++++++++-- 2 files changed, 43 insertions(+), 3 deletions(-) diff --git a/functions/compose-inference-gateway/function/fn.py b/functions/compose-inference-gateway/function/fn.py index dbf36caae..33257c1e7 100644 --- a/functions/compose-inference-gateway/function/fn.py +++ b/functions/compose-inference-gateway/function/fn.py @@ -836,7 +836,7 @@ def compose_cluster_name(self, hostname: str, address: str) -> None: ), ) return - # Selectorless and headless: cluster DNS answers with the EndpointSlice's + # Selectorless and headless: cluster DNS answers with the endpoint's # address directly, so Envoy connects to the load balancer rather than # hairpinning through a ClusterIP. resource.update( @@ -869,6 +869,32 @@ def compose_cluster_name(self, hostname: str, address: str) -> None: }, ), ) + # The same address again, as the Endpoints this supersedes. kube-dns + # reads only that API and was never taught the EndpointSlice one, so on + # a cluster running it - which is every GKE cluster, where it is still + # the default - the name answers NXDOMAIN with the slice alone. Envoy + # then resolves no address for the backend and every cross-cluster + # request is refused with no healthy upstream. + # + # Deprecated since 1.33 and still the only thing kube-dns reads. It + # costs one object; getting it wrong costs every request to the cluster. + resource.update( + self.rsp.desired.resources[f"cluster-name-endpoints-{label}"], + _k8s_object( + self.pc, + { + "apiVersion": "v1", + "kind": "Endpoints", + "metadata": {"name": label, "namespace": REMOTE_NAMESPACE}, + "subsets": [ + { + "addresses": [{"ip": address}], + "ports": [{"name": "https", "port": _CLUSTER_GATEWAY_PORT}], + } + ], + }, + ), + ) def compose_caller_auth(self) -> None: """A SecurityPolicy authenticating callers against the selected Secrets. diff --git a/functions/compose-inference-gateway/tests/test_fn.py b/functions/compose-inference-gateway/tests/test_fn.py index c19ce6379..f5b9b50cb 100644 --- a/functions/compose-inference-gateway/tests/test_fn.py +++ b/functions/compose-inference-gateway/tests/test_fn.py @@ -773,8 +773,10 @@ async def test_resolves_each_cluster_gateway_name(self) -> None: compose-inference-cluster derived, and this gateway's Envoy resolves it, so its cluster needs a Service of that name. An IP is served by a headless Service and an EndpointSlice; a hostname, which is how a cloud - load balancer names itself, by an ExternalName Service. A cluster that - hasn't published both an address and a name gets neither. + load balancer names itself, by an ExternalName Service. An IP also gets + the Endpoints the slice supersedes, because kube-dns reads only that and + is what GKE runs. A cluster that hasn't published both an address and a + name gets neither. """ ipv4 = "prod-ipv4-gateway-aaaaa.modelplane-system.svc.cluster.local" ipv6 = "prod-ipv6-gateway-bbbbb.modelplane-system.svc.cluster.local" @@ -814,6 +816,12 @@ async def test_resolves_each_cluster_gateway_name(self) -> None: "metadata": {"name": "prod-ipv4-gateway-aaaaa", "namespace": fn.REMOTE_NAMESPACE}, "spec": {"clusterIP": "None", "ports": [{"name": "https", "port": 443}]}, }, + "cluster-name-endpoints-prod-ipv4-gateway-aaaaa": { + "apiVersion": "v1", + "kind": "Endpoints", + "metadata": {"name": "prod-ipv4-gateway-aaaaa", "namespace": fn.REMOTE_NAMESPACE}, + "subsets": [{"addresses": [{"ip": "203.0.113.7"}], "ports": [{"name": "https", "port": 443}]}], + }, "cluster-name-slice-prod-ipv4-gateway-aaaaa": { "apiVersion": "discovery.k8s.io/v1", "kind": "EndpointSlice", @@ -832,6 +840,12 @@ async def test_resolves_each_cluster_gateway_name(self) -> None: "metadata": {"name": "prod-ipv6-gateway-bbbbb", "namespace": fn.REMOTE_NAMESPACE}, "spec": {"clusterIP": "None", "ports": [{"name": "https", "port": 443}]}, }, + "cluster-name-endpoints-prod-ipv6-gateway-bbbbb": { + "apiVersion": "v1", + "kind": "Endpoints", + "metadata": {"name": "prod-ipv6-gateway-bbbbb", "namespace": fn.REMOTE_NAMESPACE}, + "subsets": [{"addresses": [{"ip": "2001:db8::1"}], "ports": [{"name": "https", "port": 443}]}], + }, "cluster-name-slice-prod-ipv6-gateway-bbbbb": { "apiVersion": "discovery.k8s.io/v1", "kind": "EndpointSlice", From 96b9c80602338df34a0c5ab18ef3eccbe811039f Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 10:11:25 -0700 Subject: [PATCH 24/42] Set the required aggregation in the metric-mapping fixture acrossReplicas is required, and this fixture predates it. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- functions/compose-metric-mapping/tests/test_fn.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/functions/compose-metric-mapping/tests/test_fn.py b/functions/compose-metric-mapping/tests/test_fn.py index 5399e5dd1..51e87bcea 100644 --- a/functions/compose-metric-mapping/tests/test_fn.py +++ b/functions/compose-metric-mapping/tests/test_fn.py @@ -52,7 +52,7 @@ async def test_compose(self) -> None: "kind": "MetricMapping", "metadata": {"name": "my-engine"}, "spec": { - "metrics": [{"from": "my_engine_queued", "to": "modelplane_requests_waiting"}], + "metrics": [{"from": "my_engine_queued", "to": "modelplane_requests_waiting", "acrossReplicas": "Sum"}], }, } cluster = resource.dict_to_struct( From 4a3566046ecb88d61bed38d898976639cd8aee8c Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 10:31:22 -0700 Subject: [PATCH 25/42] Regenerate the schema lock after rebasing The lock records one content hash over apis/, so rebasing onto a main that changed apis/ leaves it describing neither tree. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- schemas/.lock.json | 13 +++++++------ 1 file changed, 7 insertions(+), 6 deletions(-) diff --git a/schemas/.lock.json b/schemas/.lock.json index ff98a6809..5b8f42bc3 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,14 +1,15 @@ { "packages": { - "fs://apis": "c0d16165e05888fb82f0d605a67b5993d178c4b2305143a84d988782d257a53c", + "fs://apis": "c0d3914bd348e896c589e1f5122671eba1e95f68d6489a06e503b556a55e94e8", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", - "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.6.0": "sha256:acc26c8d2710e0306185b6c626a2f8c8fe0fdf89874e85efe7944a3668322865", - "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.6.0": "sha256:00f1bbbb3c0f1948b6dd45c841a083d63e41850f5a559a15580915311f404911", - "xpkg://xpkg.upbound.io/upbound/provider-aws-eks:v2.6.0": "sha256:5d144b19e188cb96c918aa7e4ccbc6759b8733ff09bd8ce412723c669aa3f763", - "xpkg://xpkg.upbound.io/upbound/provider-aws-iam:v2.6.0": "sha256:dbc5288589ccb302d527565680477f08477c280fc5c616dda95dfd558108a038", + "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.8.1": "sha256:ca2e9e3b2e3a8b6ca44a9700d5abf7abd733cfa388d1afe9bb2bf7c847cd39ef", + "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.8.1": "sha256:ebb1bcd8dc9a7e60e97a324652fbd9d1b0609ddbb1107db518ce999cbc1113bd", + "xpkg://xpkg.upbound.io/upbound/provider-aws-eks:v2.8.1": "sha256:45ba27a14d6f8c4b9acdced0c48bec5650aab22bfc275af8bf1f5c770da3fa24", + "xpkg://xpkg.upbound.io/upbound/provider-aws-iam:v2.8.1": "sha256:ef802f20c76dc7d513d811beb7a9f8b4f5360eaf0d133c71bda289d02f766472", "xpkg://xpkg.upbound.io/upbound/provider-azure-containerservice:v2.6.0": "sha256:7d8a9bb3eb168e6eef0694253a23321fa98acad8b897ade168a4f0bb89985fed", "xpkg://xpkg.upbound.io/upbound/provider-azure-network:v2.6.0": "sha256:0d83bc4964488e5602b56dd7680e4bae0fbd5fa94b3a9af403d8431ded45635f", - "xpkg://xpkg.upbound.io/upbound/provider-family-aws:v2.6.0": "sha256:9fbe222866e9b763dae2db4de1599a52273125a75e8fe6328abbdc827695a74c", + "xpkg://xpkg.upbound.io/upbound/provider-civo:v1.0.0": "sha256:9cb7a795930d0b467a44d6b083560c48d977d802fb6be73e0ad909d9700874b7", + "xpkg://xpkg.upbound.io/upbound/provider-family-aws:v2.8.1": "sha256:b9c8e06e52c30c664a70997cc881e8e6b7823d7806a75319364930fcaf54df46", "xpkg://xpkg.upbound.io/upbound/provider-family-azure:v2.6.0": "sha256:1f2f597d5702ccb241429f0d8942f3c21e0ea55fffcef4ff9dca76b32bdc5dff", "xpkg://xpkg.upbound.io/upbound/provider-family-gcp:v2.6.0": "sha256:2e33bfd0f501155e7be63f13471d2f87f00519f99b916b7cd63ff40faa517145", "xpkg://xpkg.upbound.io/upbound/provider-gcp-cloudplatform:v2.6.0": "sha256:f1fe8bc55c474464642303e6fa8608c83e369b42ff12bb8a60a3e2d77339a52b", From 2397158327efc6bda1693654f6f213a9594560f0 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 10:42:57 -0700 Subject: [PATCH 26/42] Name the processor that lifts the identity for what it does It is called groupbyattrs/replicas and it does not merge replicas - that it cannot, since each replica is scraped separately and is therefore its own batch. What it does is move the identity discovery wrote onto each datapoint up onto the resource, which is where an exporter that flattens a series into labels reads it. The old name describes an intent the pipeline abandoned when replicas got a label of their own, and it reads like a no-op worth deleting. It isn't: running the collector without it on a cluster, the resource came back carrying the cluster and nothing else - no deployment, replica, engine or role. A test holds that now, so the next person to read it as dead weight finds out here rather than from a dashboard. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 24 +++++++++++-------- .../tests/test_collector.py | 24 +++++++++++++++---- 2 files changed, 33 insertions(+), 15 deletions(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 274b3474d..53e36ffd3 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -472,15 +472,19 @@ def config( ] }, "transform/modelplane": _transform(mappings), - # Lifts the identity onto the resource, where an exporter that flattens - # a series into labels will find it. It does not merge a deployment's - # replicas: each is scraped separately, so each is its own batch, and - # there is never more than one replica here to merge. They stay - # separate series, told apart by the replica label, and a query over - # the deployment combines them - which is the only place the - # arithmetic can be right, because adding two cumulative readings - # taken at different moments is not the traffic that happened. - "groupbyattrs/replicas": {"keys": list(_IDENTITY)}, + # Discovery writes the identity onto each datapoint; this lifts it onto + # the resource, which is where an exporter that flattens a series into + # labels looks for it. Without it a series arrives carrying only the + # cluster. + # + # It does not merge a deployment's replicas, whatever its name + # suggests. Each replica is scraped separately, so each is its own + # batch, and there is never a second replica here to merge with. They + # stay separate series, told apart by the replica label, and a query + # over the deployment combines them - which is the only place the + # arithmetic can be right, because adding two cumulative readings taken + # at different moments is not the traffic that happened. + "groupbyattrs/identity": {"keys": list(_IDENTITY)}, "filter/modelplane": {"metrics": {"metric": ['not IsMatch(name, "^modelplane_.*")']}}, "batch": {"timeout": "10s"}, } @@ -488,7 +492,7 @@ def config( "resource/cluster", "transform/identity", "transform/modelplane", - "groupbyattrs/replicas", + "groupbyattrs/identity", "filter/modelplane", "batch", ] diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index b3f61103a..9552b5c11 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -59,13 +59,14 @@ class TestConfig(unittest.TestCase): """The collector configuration this renders.""" def test_pipeline_order(self) -> None: - """groupbyattrs runs before the merge, or the merge combines nothing. + """The rename runs before the identity is lifted onto the resource. - A pod's identity is a resource attribute, which a metric processor - can't see, so the resources have to be stripped and merged first. + Discovery writes the identity onto each datapoint and groupbyattrs + lifts it; a statement matching on a metric's name has to run while the + datapoints are still where the rename can reach them. """ procs = _config()["service"]["pipelines"]["metrics"]["processors"] - self.assertLess(procs.index("transform/modelplane"), procs.index("groupbyattrs/replicas")) + self.assertLess(procs.index("transform/modelplane"), procs.index("groupbyattrs/identity")) self.assertEqual(procs[-1], "batch") def test_only_modelplane_leaves_the_cluster(self) -> None: @@ -105,7 +106,20 @@ def test_only_the_identity_survives_to_the_exporter(self) -> None: self.assertNotIn("'", statement) self.assertIn('keep_keys(resource.attributes, ["cluster"', statement) pipeline = _config()["service"]["pipelines"]["metrics"]["processors"] - self.assertLess(pipeline.index("transform/identity"), pipeline.index("groupbyattrs/replicas")) + self.assertLess(pipeline.index("transform/identity"), pipeline.index("groupbyattrs/identity")) + + def test_the_identity_is_lifted_onto_the_resource(self) -> None: + """Without this a series arrives carrying only the cluster. + + Discovery writes the identity onto each datapoint. An exporter that + flattens a series into labels reads the resource, so something has to + move it, and this is the only processor that does. Removing it as a + no-op strips every series of what says who it belongs to - verified on + a cluster, where the resource came back carrying `cluster` alone. + """ + cfg = _config() + self.assertEqual(cfg["processors"]["groupbyattrs/identity"]["keys"], list(collector._IDENTITY)) + self.assertIn("groupbyattrs/identity", cfg["service"]["pipelines"]["metrics"]["processors"]) def test_engine_scrape_selects_the_port_by_name(self) -> None: """Matching by number would find the pd-sidecar on a disaggregated pod.""" From 505835ac1d93655b0d57975e725b3b6de7c0ae23 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 11:12:55 -0700 Subject: [PATCH 27/42] Keep every series unique to whatever produced it Four things @haarchri found reviewing the collector, three of them the same bug: a series that two producers write is one series, and one of them is lost. The identity names an engine and nothing else. Two gateway pods carry none of it and publish identical modelplane_frontend_* series. Two replicas of a substrate controller share only their namespace. And a ModelReplica with copies greater than one runs several pods under one replica index, so the label meant to tell replicas apart doesn't tell those apart. Each is the bug the replica label was added to fix, in a place the label doesn't reach. Keep what the scrape came from - service.name and service.instance.id, which an exporter renders as job and instance - so a series is unique to its target wherever the modelplane identity doesn't reach. It is the pod's address, so it churns on a rolling update; aggregate it away in the query, which is what the identity and acrossReplicas are for. A part is now extracted in a block of its own ahead of the datapoint statements. Extraction mints _count, and a label or a unit conversion for it selects on that name, which did not exist until the extraction ran - so labels and fromUnit combined with part silently did nothing. The port rewrite's host pattern could not match a bracketed IPv6 address, so an IPv6 pod kept whichever port discovery picked rather than the one it annotated. And a memory limiter, first in the pipeline. Nothing bounds what one interval brings off a fleet of engines, and the kill that follows a spike takes the cluster's telemetry with it. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 63 ++++++++++++++---- .../tests/test_collector.py | 66 ++++++++++++++++++- 2 files changed, 116 insertions(+), 13 deletions(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 53e36ffd3..80d6f144c 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -63,7 +63,31 @@ # node is here for the GPU job, whose series belong to hardware rather than to # a deployment; an engine's series carry no node, which is what lets replicas on # different nodes merge. -_IDENTITY = ("cluster", "namespace", "deployment", "replica", "engine", "role", "node") +_IDENTITY = ( + "cluster", + "namespace", + "deployment", + "replica", + "engine", + "role", + "node", + # What the scrape came from, as the Prometheus receiver names it, which an + # exporter renders as job and instance. + # + # Kept because a series has to be unique to whatever produced it or one + # producer's numbers silently replace another's. The identity above covers + # an engine; it covers nothing else. Two gateway pods carry no identity at + # all, two replicas of a substrate controller share a namespace, and a + # ModelReplica with copies greater than one runs several pods under one + # replica index. Each of those is a collision, and a collision is a wrong + # number that looks right. + # + # It is the pod's address, so it does churn on a rolling update, which is + # the cost. Aggregate it away in the query: the identity above is what to + # group by, and `acrossReplicas` names how. + "service.name", + "service.instance.id", +) # OTTL quotes strings with double quotes; a Python list renders single ones and # the collector refuses to start on it. @@ -139,7 +163,10 @@ def _annotated_port() -> list[dict[str, Any]]: "source_labels": ["__address__", "__meta_kubernetes_pod_annotation_prometheus_io_port"], "action": "replace", "target_label": "__address__", - "regex": r"([^:]+)(?::\d+)?;(\d+)", + # A bracketed IPv6 host or a bare one: [^:]+ alone never matches + # [2001:db8::1]:9090, so an IPv6 pod would keep whichever port + # discovery happened to pick. + "regex": r"(\[.+\]|[^:]+)(?::\d+)?;(\d+)", "replacement": "$1:$2", } ] @@ -302,7 +329,7 @@ def _scrape_configs() -> list[dict[str, Any]]: } -def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], list[str]]: +def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], list[str], list[str]]: """Compile the mappings to OTTL, as (datapoint, metric) statements. OTTL is rendered here rather than written in a MetricMapping so the kind @@ -310,17 +337,20 @@ def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], lis configuration language stays Modelplane's problem. It is also the only place that knows a unit conversion has to run somewhere a rename cannot. """ + extract: list[str] = [] datapoint: list[str] = [] metric: list[str] = [] for mapping in mappings: for m in mapping.spec.metrics: source = m.from_ if m.part: - # Lift the part out first, under the name the extraction gives - # it, and rename that. The histogram carries on untouched. - suffix = _PART_SUFFIX[m.part] - metric.append(f'{_PART_FUNCTION[m.part]} where metric.name == "{source}"') - source = f"{source}{suffix}" + # In a block of its own, ahead of the datapoint statements. The + # extraction mints a new metric, and a label or a unit + # conversion for it selects on the name that extraction gives + # it - a name that does not exist until the extraction has run. + # The histogram carries on untouched. + extract.append(f'{_PART_FUNCTION[m.part]} where metric.name == "{source}"') + source = f"{source}{_PART_SUFFIX[m.part]}" if m.fromUnit: conversion = _UNIT_CONVERSION[m.fromUnit] datapoint.append(f'set(datapoint.value_double, {conversion}) where metric.name == "{source}"') @@ -329,7 +359,7 @@ def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], lis # one name and a statement could no longer tell them apart. datapoint.extend(_label_statements(source, m.labels or [])) metric.append(f'set(metric.name, "{m.to}") where metric.name == "{source}"') - return datapoint, metric + return extract, datapoint, metric def _label_statements(source: str, labels: list[Any]) -> list[str]: @@ -360,10 +390,13 @@ def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: every datapoint before it starts the next, which is what keeps a rename from stranding the datapoints a conversion hasn't reached yet. """ - datapoint, metric = statements(mappings) - blocks = [{"context": "metric", "statements": metric}] + extract, datapoint, metric = statements(mappings) + blocks: list[dict[str, Any]] = [] + if extract: + blocks.append({"context": "metric", "statements": extract}) if datapoint: - blocks.insert(0, {"context": "datapoint", "statements": datapoint}) + blocks.append({"context": "datapoint", "statements": datapoint}) + blocks.append({"context": "metric", "statements": metric}) return {"metric_statements": blocks} @@ -449,6 +482,11 @@ def config( # anything, but not replace the one composed for a sink's own auth block. extensions = {**extensions, **authenticators(sinks)} processors: dict[str, Any] = { + # First in the pipeline, so it refuses work before anything allocates + # for it. Nothing bounds what one interval brings off a fleet of + # engines, and the kill that follows a spike takes the whole cluster's + # telemetry with it until the pod is back. + "memory_limiter": {"check_interval": "1s", "limit_percentage": 80, "spike_limit_percentage": 25}, # cluster is stamped here rather than downstream: one receiver on the # control plane sees a merged stream and cannot tell senders apart. "resource/cluster": {"attributes": [{"key": "cluster", "value": cluster, "action": "upsert"}]}, @@ -489,6 +527,7 @@ def config( "batch": {"timeout": "10s"}, } pipeline = [ + "memory_limiter", "resource/cluster", "transform/identity", "transform/modelplane", diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 9552b5c11..59f318a48 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -14,6 +14,7 @@ """Tests for the collector this stack composes.""" +import re import typing import unittest @@ -41,7 +42,7 @@ def _sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = N def _metric_statements() -> list[str]: - return collector.statements(list(stacks.BUILTIN_MAPPINGS))[1] + return collector.statements(list(stacks.BUILTIN_MAPPINGS))[2] def _config(*, extensions: dict | None = None) -> dict: @@ -121,6 +122,69 @@ def test_the_identity_is_lifted_onto_the_resource(self) -> None: self.assertEqual(cfg["processors"]["groupbyattrs/identity"]["keys"], list(collector._IDENTITY)) self.assertIn("groupbyattrs/identity", cfg["service"]["pipelines"]["metrics"]["processors"]) + def test_a_part_is_extracted_before_anything_selects_on_it(self) -> None: + """A label or a unit for an extracted part names a metric that must exist. + + The extraction mints `_count`, and a datapoint statement for it + selects on that name. Run the datapoint block first and it matches + nothing, silently. + """ + mapping = mmv1alpha1.MetricMapping.model_validate( + { + "spec": { + "metrics": [ + { + "from": "my_engine_duration_ms", + "to": "modelplane_requests_total", + "acrossReplicas": "Sum", + "part": "Count", + "fromUnit": "Milliseconds", + "labels": [{"name": "status", "value": "ok"}], + } + ] + } + } + ) + blocks = collector._transform([mapping])["metric_statements"] + contexts = [b["context"] for b in blocks] + self.assertEqual(contexts, ["metric", "datapoint", "metric"]) + self.assertIn("extract_count_metric", blocks[0]["statements"][0]) + # Everything selecting on the extracted name comes after the extraction. + for statement in blocks[1]["statements"]: + self.assertIn("my_engine_duration_ms_count", statement) + + def test_every_job_carries_something_unique_to_its_target(self) -> None: + """Two producers whose series are identical are one series, and one is lost. + + The modelplane identity names an engine and nothing else: a gateway pod + carries none of it, two replicas of a substrate controller share a + namespace, and a ModelReplica with copies > 1 runs several pods under + one replica index. + """ + self.assertIn("service.instance.id", collector._IDENTITY) + statement = _config()["processors"]["transform/identity"]["metric_statements"][0]["statements"][0] + self.assertIn('"service.instance.id"', statement) + + def test_a_scrape_spike_cannot_take_the_collector_down(self) -> None: + """Nothing bounds what one interval brings off a fleet of engines.""" + cfg = _config() + self.assertIn("memory_limiter", cfg["processors"]) + self.assertEqual(cfg["service"]["pipelines"]["metrics"]["processors"][0], "memory_limiter") + + def test_the_port_rewrite_matches_an_ipv6_pod(self) -> None: + """__address__ is [2001:db8::1]:9090 there, which [^:]+ never matches.""" + rule = next( + r + for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"] + if j["job_name"] == "modelplane-substrate" + for r in j["relabel_configs"] + if r.get("target_label") == "__address__" + ) + for address in ("10.1.0.5:8000", "[2001:db8::1]:9090"): + matched = re.fullmatch(rule["regex"], f"{address};9402") + assert matched is not None, address + self.assertTrue(matched.expand(r"\1:\2").endswith(":9402")) + def test_engine_scrape_selects_the_port_by_name(self) -> None: """Matching by number would find the pd-sidecar on a disaggregated pod.""" jobs = {j["job_name"]: j for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"]} From 47bf5698ec2e53b5b898af9259198bfd9e3b50e1 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 11:21:12 -0700 Subject: [PATCH 28/42] Say what a series actually carries Keeping the scrape target's identity so two pods can't collide made three statements in the docs false, and they were the reassuring kind. The guide said the series count doesn't grow when you deploy. It does now: instance is the pod's address and turns over on a rolling update. It said there is no pod label. instance is one, in all but name. And both the guide and the MetricMapping said a replica publishes its own series, when what publishes one is a pod - which is the whole reason instance is there, since a deployment can run several pods per replica. Say what is true instead, and say what to do about it: group by the replica, which is an index and stable, and aggregate the instance away. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 6 +++--- docs/content/platform/telemetry.md | 20 +++++++++++++------- 2 files changed, 16 insertions(+), 10 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index 8130230c9..e6ef9756a 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -84,9 +84,9 @@ spec: description: >- How this metric combines over a deployment's replicas. - Each replica publishes its own series, told apart by - the replica label, and a query over a deployment - combines them. This says which combination is the right + Every pod publishes its own series, told apart by the + replica it belongs to and the instance it was scraped + from, and a query over a deployment combines them. This says which combination is the right one: Sum for anything counted - requests, tokens, joules, a queue's depth. Mean for a ratio, where summing reads two replicas at half capacity as one at diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index f1467ba59..1a5895475 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -19,8 +19,9 @@ deployments change. ## What you get -Every series carries `cluster`. A series about a deployment also carries `deployment`, -`replica`, `namespace`, `engine`, and `role`. +Every series carries `cluster`, and `job` and `instance` naming the target it was scraped +from. A series about a deployment also carries `deployment`, `replica`, `namespace`, +`engine`, and `role`. Each replica publishes its own series. Combine them in the query, the way the metric's `acrossReplicas` says: `sum by (deployment)` for anything counted, `avg by (deployment)` @@ -28,8 +29,13 @@ for a ratio, `max by (deployment)` for a saturation figure an alert fires on. Th collector doesn't add them up for you, because a scrape of one replica is one batch, and adding readings taken at different moments is not the traffic that happened. -The replica is an index, not a pod. It's bounded by the replica count and it survives a -restart and a rolling update, so the series count doesn't grow every time you deploy. +The replica is an index rather than a pod, so it is bounded by the replica count and +survives a restart and a rolling update. Group by it, not by `instance`. + +`instance` is the pod's address, and it is there because two pods writing one series is one +series with one of them lost - a deployment running several pods per replica, or two +gateway pods, have nothing else to tell them apart. It does turn over on a rolling update, +so a query that groups by it grows a series every time you deploy. Aggregate it away. Some of what you can read: @@ -248,9 +254,9 @@ against Modelplane's Prometheus stops being read by anything, because the operat the stack. **Rewrite your dashboard queries.** Names change, and so do the labels: group by -`deployment` rather than `model_name`, there's no pod label, and every series carries -`cluster` and `replica`. A panel that showed one engine now shows one replica, so wrap it -in `sum by (deployment)` or the aggregation that metric's `acrossReplicas` names. +`deployment` rather than `model_name`, and every series carries `cluster`, `replica`, and +the `instance` it was scraped from. A panel that showed one engine now shows one pod, so +wrap it in `sum by (deployment)` or the aggregation that metric's `acrossReplicas` names. | Was | Is | | --- | --- | From 313006ee32825350b203a57a11ce416ffa8b2af7 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 12:09:26 -0700 Subject: [PATCH 29/42] Let a component say it counts in percent fromUnit could convert millijoules, mebibytes, milliseconds and nanoseconds, and not the one an operator is most likely to meet: a saturation counted from nought to a hundred, mapped onto a name that says a ratio. There was no way to say so, and the series lands a hundred times high - which reads on a dashboard as a cache at ten thousand per cent. Neither built-in needs it. vLLM divides used by total, and so does SGLang, so both are already fractions. But vLLM calls its metric kv_cache_usage_perc, and a name that says percent over a value that is a ratio is exactly why the field has to be asked rather than inferred. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 12 ++++++-- .../function/collector.py | 1 + .../tests/test_collector.py | 30 +++++++++++++++++++ schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 8 +++-- 5 files changed, 46 insertions(+), 7 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index e6ef9756a..563476827 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -170,13 +170,19 @@ spec: Only meaningful alongside `from`. fromUnit: type: string - enum: [Millijoules, Mebibytes, Milliseconds, Nanoseconds] + enum: [Millijoules, Mebibytes, Milliseconds, Nanoseconds, Percent] description: >- What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by - a thousand, nanoseconds by a billion, and mebibytes - multiplied out to bytes. + a thousand, nanoseconds by a billion, percent by a + hundred, and mebibytes multiplied out to bytes. + + Percent is for a component that counts a saturation + from nought to a hundred where the name says a ratio. + Check rather than assume: vLLM publishes + kv_cache_usage_perc and the value is a fraction, so a + name is no guide. Say it whenever the source disagrees with the target, even where the factor looks obvious. A name ending in diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 80d6f144c..61dc5ff5f 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -326,6 +326,7 @@ def _scrape_configs() -> list[dict[str, Any]]: "Milliseconds": "datapoint.value_double / 1000", "Nanoseconds": "datapoint.value_double / 1000000000", "Mebibytes": "datapoint.value_double * 1048576", + "Percent": "datapoint.value_double / 100", } diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 59f318a48..c8ade6483 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -227,6 +227,36 @@ def test_every_unit_the_api_offers_has_a_conversion(self) -> None: literal = next(a for a in typing.get_args(annotation) if typing.get_origin(a) is typing.Literal) self.assertEqual(set(typing.get_args(literal)), set(collector._UNIT_CONVERSION)) + def test_a_percentage_is_divided_into_a_ratio(self) -> None: + """A component counting 0 to 100 under a name that says a ratio is 100x out. + + vLLM and SGLang both publish a fraction, so no built-in needs this, but + vLLM's is called kv_cache_usage_perc - the name is no guide, and an + engine that means it has to be able to say so. + """ + mapping = mmv1alpha1.MetricMapping.model_validate( + { + "spec": { + "metrics": [ + { + "from": "my_engine_cache_percent", + "to": "modelplane_kv_cache_utilization_ratio", + "acrossReplicas": "Mean", + "fromUnit": "Percent", + } + ] + } + } + ) + _, datapoint, _ = collector.statements([mapping]) + self.assertEqual( + datapoint, + [ + "set(datapoint.value_double, datapoint.value_double / 100) " + 'where metric.name == "my_engine_cache_percent"' + ], + ) + def test_a_metric_name_cannot_end_the_comparison_early(self) -> None: """A quote in `from` would rename whatever the rest of the line matched.""" with self.assertRaises(ValidationError): diff --git a/schemas/.lock.json b/schemas/.lock.json index 5b8f42bc3..1e1c56c9c 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "c0d3914bd348e896c589e1f5122671eba1e95f68d6489a06e503b556a55e94e8", + "fs://apis": "7d6ebf1a8ad6d6797cff6bf790ff01378ba3ead27240e0ddb20f32536c9a66ea", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.8.1": "sha256:ca2e9e3b2e3a8b6ca44a9700d5abf7abd733cfa388d1afe9bb2bf7c847cd39ef", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.8.1": "sha256:ebb1bcd8dc9a7e60e97a324652fbd9d1b0609ddbb1107db518ce999cbc1113bd", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index 8dd7442df..23d5a30d7 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -68,7 +68,7 @@ class Metric(BaseModel): acrossReplicas: Literal['Sum', 'Mean', 'Max'] """ How this metric combines over a deployment's replicas. - Each replica publishes its own series, told apart by the replica label, and a query over a deployment combines them. This says which combination is the right one: Sum for anything counted - requests, tokens, joules, a queue's depth. Mean for a ratio, where summing reads two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where a mean hides the replica in trouble. + Every pod publishes its own series, told apart by the replica it belongs to and the instance it was scraped from, and a query over a deployment combines them. This says which combination is the right one: Sum for anything counted - requests, tokens, joules, a queue's depth. Mean for a ratio, where summing reads two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where a mean hides the replica in trouble. Modelplane does not combine them in the collector. A scrape of one replica is one batch, so a collector that added them up would be adding readings taken at different moments, and two readings of one cumulative counter sum to twice the traffic that happened. The backend holds every replica's series and combines them at query time, where the arithmetic is right. Required, with no default, because the wrong combination is silent: a deployment reports a number that looks entirely plausible. """ @@ -80,10 +80,12 @@ class Metric(BaseModel): Held to the characters a metric name can contain. The name is matched inside the collector's own query language, so a quote here would end the comparison early and rename whatever the rest of the line matched. """ fromUnit: ( - Literal['Millijoules', 'Mebibytes', 'Milliseconds', 'Nanoseconds'] | None + Literal['Millijoules', 'Mebibytes', 'Milliseconds', 'Nanoseconds', 'Percent'] + | None ) = None """ - What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, and mebibytes multiplied out to bytes. + What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, percent by a hundred, and mebibytes multiplied out to bytes. + Percent is for a component that counts a saturation from nought to a hundred where the name says a ratio. Check rather than assume: vLLM publishes kv_cache_usage_perc and the value is a fraction, so a name is no guide. Say it whenever the source disagrees with the target, even where the factor looks obvious. A name ending in _bytes that holds mebibytes is the kind of thing nobody notices until a capacity review, and stating the source unit is what makes the conversion happen at all. """ labels: list[Label] | None = Field(None, max_length=16) From f93cbcd4f546318569f6671d2052ca8f9635b1b3 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:51:37 -0700 Subject: [PATCH 30/42] Compile a mapping's unit conversion to scale_metric A conversion was a datapoint statement setting value_double, which a histogram has none of: its measurements live in its sum, its minimum and maximum, and every bucket boundary. A histogram with fromUnit was renamed to seconds with its buckets still in milliseconds, putting every quantile a thousand times out with nothing to say so. Reading value_double also read an integer datapoint as nought. scale_metric carries all of them, so the conversion moves to the metric context, in a block of its own ahead of the renames. It ignores its own errors: it refuses an exponential histogram, and under the default error mode that one refusal would drop the whole scrape rather than the metric it could not convert. Two other things a mapping could put into OTTL unescaped. A label's value and the keys and values of a values remap are free text - they carry whatever vocabulary the component emits - so a quote in one closed the literal early and left the rest as OTTL. And a label carried onto its own name, which is how a mapping rewrites values in place, was deleted again by the statement that stops a carried label costing twice the cardinality. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/collector.py | 110 +++++++++---- .../tests/test_collector.py | 146 +++++++++++++++--- 2 files changed, 203 insertions(+), 53 deletions(-) diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index 61dc5ff5f..de81414cc 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -311,27 +311,36 @@ def _scrape_configs() -> list[dict[str, Any]]: ] -# What each source unit is worth in the base unit the target name claims. -# Written as the expression rather than a factor so nothing has to render a -# float: 1e-09 is not an OTTL literal. Paths carry their context because the -# collector rewrites bare ones and asks the author to stop; Modelplane is the -# author here, so nobody's stored MetricMapping has to change. # Taking a part of a histogram is a function that mints a new metric beside it, # named for the part. The rename then applies to that. _PART_FUNCTION = {"Count": "extract_count_metric(true)", "Sum": "extract_sum_metric(true)"} _PART_SUFFIX = {"Count": "_count", "Sum": "_sum"} -_UNIT_CONVERSION = { - "Millijoules": "datapoint.value_double / 1000", - "Milliseconds": "datapoint.value_double / 1000", - "Nanoseconds": "datapoint.value_double / 1000000000", - "Mebibytes": "datapoint.value_double * 1048576", - "Percent": "datapoint.value_double / 100", +# What one of the source unit is worth in the base unit the target name claims. +# +# Applied with scale_metric in the metric context rather than by setting a +# datapoint's value, because a histogram has no single value to set: its sum, +# its minimum and maximum, and every one of its bucket boundaries are all in +# the source unit, and a conversion that reached only the value would rename a +# histogram to seconds with its buckets still in milliseconds. scale_metric +# carries all of them, and an integer datapoint as well, which reading +# value_double would have read as nought. +# +# Written out rather than as a Python float so nothing has to render one: +# 1e-09 is not an OTTL literal. Every one carries a decimal point, because +# scale_metric takes a float and the collector refuses to start on an integer +# literal in that position. +_UNIT_FACTOR = { + "Millijoules": "0.001", + "Milliseconds": "0.001", + "Nanoseconds": "0.000000001", + "Mebibytes": "1048576.0", + "Percent": "0.01", } -def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], list[str], list[str]]: - """Compile the mappings to OTTL, as (datapoint, metric) statements. +def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], list[str], list[str], list[str]]: + """Compile the mappings to OTTL, as (extract, scale, datapoint, metric) statements. OTTL is rendered here rather than written in a MetricMapping so the kind stays a description of what a component emits, and the collector's own @@ -339,6 +348,7 @@ def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], lis place that knows a unit conversion has to run somewhere a rename cannot. """ extract: list[str] = [] + scale: list[str] = [] datapoint: list[str] = [] metric: list[str] = [] for mapping in mappings: @@ -349,18 +359,32 @@ def statements(mappings: list[mmv1alpha1.MetricMapping]) -> tuple[list[str], lis # extraction mints a new metric, and a label or a unit # conversion for it selects on the name that extraction gives # it - a name that does not exist until the extraction has run. - # The histogram carries on untouched. + # The extraction leaves the histogram itself alone, though the + # filter downstream drops it unless a mapping renames it too. extract.append(f'{_PART_FUNCTION[m.part]} where metric.name == "{source}"') source = f"{source}{_PART_SUFFIX[m.part]}" if m.fromUnit: - conversion = _UNIT_CONVERSION[m.fromUnit] - datapoint.append(f'set(datapoint.value_double, {conversion}) where metric.name == "{source}"') + scale.append(f'scale_metric({_UNIT_FACTOR[m.fromUnit]}) where metric.name == "{source}"') # Labels before the rename, while the series still answers to the # name this mapping selected on. After it, two folded mappings share # one name and a statement could no longer tell them apart. datapoint.extend(_label_statements(source, m.labels or [])) metric.append(f'set(metric.name, "{m.to}") where metric.name == "{source}"') - return extract, datapoint, metric + return extract, scale, datapoint, metric + + +def _quote(value: str) -> str: + """One OTTL string literal, with anything that would end it escaped. + + A metric name and a label name are pattern-constrained by the XRD, but a + label's value is free text, and so is every key and value of a `values` + remap - they have to be, because they carry whatever vocabulary the + component already emits. An unescaped quote in one of them would close the + literal early and leave the rest of it as OTTL, which at best stops the + collector from starting and at worst runs. + """ + escaped = value.replace("\\", "\\\\").replace('"', '\\"') + return f'"{escaped}"' def _label_statements(source: str, labels: list[Any]) -> list[str]: @@ -368,33 +392,49 @@ def _label_statements(source: str, labels: list[Any]) -> list[str]: out: list[str] = [] for label in labels: if label.value is not None: - out.append(f'set(datapoint.attributes["{label.name}"], "{label.value}") where metric.name == "{source}"') + out.append( + f'set(datapoint.attributes["{label.name}"], {_quote(label.value)}) where metric.name == "{source}"' + ) continue carried = f'datapoint.attributes["{label.from_}"]' out.append(f'set(datapoint.attributes["{label.name}"], {carried}) where metric.name == "{source}"') for old, new in sorted((label.values or {}).items()): out.append( - f'set(datapoint.attributes["{label.name}"], "{new}") ' - f'where metric.name == "{source}" and {carried} == "{old}"' + f'set(datapoint.attributes["{label.name}"], {_quote(new)}) ' + f'where metric.name == "{source}" and {carried} == {_quote(old)}' ) # The component's own name for it goes, or the series carries the same - # fact twice under two labels and costs twice the cardinality. - out.append(f'delete_key(datapoint.attributes, "{label.from_}") where metric.name == "{source}"') + # fact twice under two labels and costs twice the cardinality. Unless + # the mapping carried the label onto its own name, which is how a pure + # value remap is written: there the delete would take the label the + # statements above just set. + if label.from_ != label.name: + out.append(f'delete_key(datapoint.attributes, "{label.from_}") where metric.name == "{source}"') return out def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: - """Unit conversions first, then every rename. - - Two blocks rather than one list: a statement reaching a datapoint's value - can't run in the metric context, and the processor finishes a block over - every datapoint before it starts the next, which is what keeps a rename - from stranding the datapoints a conversion hasn't reached yet. + """Parts extracted, then units converted, then labels, then every rename. + + Four blocks rather than one list. Each block finishes over every metric + before the next one starts, and each step here depends on the one before + having finished everywhere: an extraction mints the metric a conversion + scales, a conversion has to reach a datapoint the rename would otherwise + have stranded in the source unit, and a label has to be set while the + series still answers to the name its mapping selected on. A statement + reaching a datapoint's attributes also can't run in the metric context. + + The conversions ignore their own errors. scale_metric refuses an + exponential histogram, and under the default error mode that one refusal + would drop the whole batch - every metric from every pod in the scrape, + not just the one it could not convert. """ - extract, datapoint, metric = statements(mappings) + extract, scale, datapoint, metric = statements(mappings) blocks: list[dict[str, Any]] = [] if extract: blocks.append({"context": "metric", "statements": extract}) + if scale: + blocks.append({"context": "metric", "error_mode": "ignore", "statements": scale}) if datapoint: blocks.append({"context": "datapoint", "statements": datapoint}) blocks.append({"context": "metric", "statements": metric}) @@ -410,9 +450,17 @@ def _transform(mappings: list[mmv1alpha1.MetricMapping]) -> dict[str, Any]: # # Applied under an operator's own config rather than over it: this is a default, # and a sink that sets it wins. +# +# Keyed under both spellings of the remote-write exporter. The component +# registers as prometheus_remote_write and still answers to the older +# prometheusremotewrite, so a sink writing the name the collector's own +# documentation gives would otherwise match nothing here and export every +# series stripped of its identity - a silent loss, since the sink itself works. +_RESOURCE_TO_LABELS = {"resource_to_telemetry_conversion": {"enabled": True}} _SINK_DEFAULTS = { - "prometheusremotewrite": {"resource_to_telemetry_conversion": {"enabled": True}}, - "prometheus": {"resource_to_telemetry_conversion": {"enabled": True}}, + "prometheus_remote_write": _RESOURCE_TO_LABELS, + "prometheusremotewrite": _RESOURCE_TO_LABELS, + "prometheus": _RESOURCE_TO_LABELS, } diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index c8ade6483..98398878f 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -24,7 +24,9 @@ from models.ai.modelplane.telemetrydestination import v1alpha1 as tdv1alpha1 from pydantic import ValidationError -_EXTENSIONS = {"oidc/acme": {"issuer_url": "https://issuer.acme.example"}} +# A client authenticator: an exporter needs one of those, not the oidc +# extension, which authenticates callers of a receiver. +_EXTENSIONS = {"oauth2client/acme": {"token_url": "https://issuer.acme.example/token"}} def _sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = None) -> tdv1alpha1.Sink: @@ -42,15 +44,16 @@ def _sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = N def _metric_statements() -> list[str]: - return collector.statements(list(stacks.BUILTIN_MAPPINGS))[2] + """The rename statements, which are the last of the four blocks.""" + return collector.statements(list(stacks.BUILTIN_MAPPINGS))[-1] -def _config(*, extensions: dict | None = None) -> dict: +def _config(*, extensions: dict | None = None, sinks: list | None = None) -> dict: return yaml.safe_load( collector.config( "prod-us-east", list(stacks.BUILTIN_MAPPINGS), - _SINKS, + _SINKS if sinks is None else sinks, _EXTENSIONS if extensions is None else extensions, ) ) @@ -147,11 +150,12 @@ def test_a_part_is_extracted_before_anything_selects_on_it(self) -> None: ) blocks = collector._transform([mapping])["metric_statements"] contexts = [b["context"] for b in blocks] - self.assertEqual(contexts, ["metric", "datapoint", "metric"]) + self.assertEqual(contexts, ["metric", "metric", "datapoint", "metric"]) self.assertIn("extract_count_metric", blocks[0]["statements"][0]) # Everything selecting on the extracted name comes after the extraction. - for statement in blocks[1]["statements"]: - self.assertIn("my_engine_duration_ms_count", statement) + for block in blocks[1:]: + for statement in block["statements"]: + self.assertIn("my_engine_duration_ms_count", statement) def test_every_job_carries_something_unique_to_its_target(self) -> None: """Two producers whose series are identical are one series, and one is lost. @@ -196,23 +200,76 @@ def test_gateway_has_a_target_of_its_own(self) -> None: jobs = [j["job_name"] for j in _config()["receivers"]["prometheus"]["config"]["scrape_configs"]] self.assertIn("modelplane-gateway", jobs) + def test_both_spellings_of_remote_write_keep_their_identity(self) -> None: + """The exporter registers as prometheus_remote_write in 0.161.0. + + prometheusremotewrite is the older name it still answers to. A sink + writing the one the collector's own documentation gives would + otherwise match no default here and export every series stripped of + the cluster, deployment, engine and role it belongs to - silently, + because the sink itself works. + """ + for type_ in ("prometheus_remote_write", "prometheusremotewrite", "prometheus"): + with self.subTest(type=type_): + exporters = _config(sinks=[_sink(type_=type_)])["exporters"] + exporter = next(v for k, v in exporters.items() if k.startswith(f"{type_}/")) + self.assertTrue(exporter["resource_to_telemetry_conversion"]["enabled"]) + def test_extensions_are_declared_to_the_service(self) -> None: """An authenticator the service doesn't list is one the collector won't load.""" - self.assertEqual(_config()["service"]["extensions"], ["oidc/acme"]) + self.assertEqual(_config()["service"]["extensions"], ["oauth2client/acme"]) self.assertNotIn("extensions", _config(extensions={})["service"]) def test_energy_is_scaled_before_it_is_renamed(self) -> None: """DCGM counts millijoules, and the name says joules. The scale is a block ahead of the renames, not a line ahead. The - processor finishes a block over every datapoint before the next one - starts, so a rename sharing the block would strand every datapoint + processor finishes a block over every metric before the next one + starts, so a rename sharing the block would strand every metric after the first at millijoules. """ blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] - self.assertEqual([b["context"] for b in blocks], ["datapoint", "metric"]) - self.assertTrue(any("value_double / 1000" in st for st in blocks[0]["statements"])) - self.assertTrue(any("modelplane_energy_joules_total" in st for st in blocks[1]["statements"])) + self.assertEqual([b["context"] for b in blocks], ["metric", "metric"]) + self.assertTrue(any("scale_metric(0.001)" in st for st in blocks[0]["statements"])) + self.assertTrue(any("modelplane_energy_joules_total" in st for st in blocks[-1]["statements"])) + + def test_a_conversion_reaches_a_histogram_bucket(self) -> None: + """Setting value_double converts a gauge and leaves a histogram lying. + + A histogram holds its measurements in its sum, its minimum and maximum + and every bucket boundary, none of which is value_double. Renaming one + to seconds with its buckets still at milliseconds puts every quantile + a thousand times out, and nothing says so. + """ + blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] + for block in blocks: + for statement in block["statements"]: + self.assertNotIn("value_double", statement) + scales = next(b for b in blocks if any("scale_metric" in st for st in b["statements"])) + self.assertEqual(scales["context"], "metric") + + def test_every_conversion_factor_is_a_float_literal(self) -> None: + """scale_metric takes a float, and 1048576 is an integer to OTTL. + + The collector refuses to start on it - "must be a float" - which + takes the whole cluster's telemetry down, and nothing short of + running the collector catches it. + """ + for unit, factor in collector._UNIT_FACTOR.items(): + with self.subTest(unit=unit): + self.assertIn(".", factor, "an OTTL float literal needs a decimal point") + float(factor) + + def test_a_conversion_cannot_drop_the_batch_it_rides_in(self) -> None: + """scale_metric refuses an exponential histogram. + + Under the default error mode that one refusal fails the whole batch: + every metric from every pod in the scrape is lost, not the one it + could not convert. + """ + blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] + scales = next(b for b in blocks if any("scale_metric" in st for st in b["statements"])) + self.assertEqual(scales["error_mode"], "ignore") def test_dcgm_units_are_converted_to_the_unit_the_name_claims(self) -> None: """DCGM reports mJ and MiB; the names say joules and bytes.""" @@ -225,7 +282,7 @@ def test_every_unit_the_api_offers_has_a_conversion(self) -> None: """A unit the API accepts with no conversion here is a KeyError at render time.""" annotation = mmv1alpha1.Metric.model_fields["fromUnit"].annotation literal = next(a for a in typing.get_args(annotation) if typing.get_origin(a) is typing.Literal) - self.assertEqual(set(typing.get_args(literal)), set(collector._UNIT_CONVERSION)) + self.assertEqual(set(typing.get_args(literal)), set(collector._UNIT_FACTOR)) def test_a_percentage_is_divided_into_a_ratio(self) -> None: """A component counting 0 to 100 under a name that says a ratio is 100x out. @@ -248,14 +305,8 @@ def test_a_percentage_is_divided_into_a_ratio(self) -> None: } } ) - _, datapoint, _ = collector.statements([mapping]) - self.assertEqual( - datapoint, - [ - "set(datapoint.value_double, datapoint.value_double / 100) " - 'where metric.name == "my_engine_cache_percent"' - ], - ) + _, scale, _, _ = collector.statements([mapping]) + self.assertEqual(scale, ['scale_metric(0.01) where metric.name == "my_engine_cache_percent"']) def test_a_metric_name_cannot_end_the_comparison_early(self) -> None: """A quote in `from` would rename whatever the rest of the line matched.""" @@ -270,6 +321,57 @@ def test_a_metric_name_cannot_end_the_comparison_early(self) -> None: ) self.assertEqual(round_tripped.from_, m.from_) + def test_a_label_value_cannot_end_the_string_it_sits_in(self) -> None: + """`from` is pattern-constrained; a label's value cannot be. + + A value and a `values` remap carry whatever vocabulary the component + already writes, so the schema has to take free text. A quote in one + would close the OTTL literal early and leave the remainder of the + value as OTTL - at best the collector refuses to start. + """ + mapping = mmv1alpha1.MetricMapping.model_validate( + { + "spec": { + "metrics": [ + { + "from": "my_engine_finish", + "to": "modelplane_requests_total", + "acrossReplicas": "Sum", + "labels": [{"name": "reason", "from": "finish", "values": {'ab"c': 'x"y'}}], + } + ] + } + } + ) + _, _, datapoint, _ = collector.statements([mapping]) + joined = " ".join(datapoint) + self.assertIn(r'"ab\"c"', joined) + self.assertIn(r'"x\"y"', joined) + + def test_carrying_a_label_onto_itself_keeps_it(self) -> None: + """`from` equal to `name` is how a mapping remaps values in place. + + The delete that stops a carried label costing twice the cardinality + would otherwise take the label the statements before it just set, and + the series would lose the label entirely. + """ + mapping = mmv1alpha1.MetricMapping.model_validate( + { + "spec": { + "metrics": [ + { + "from": "my_engine_finish", + "to": "modelplane_requests_total", + "acrossReplicas": "Sum", + "labels": [{"name": "reason", "from": "reason", "values": {"eos": "stop"}}], + } + ] + } + } + ) + _, _, datapoint, _ = collector.statements([mapping]) + self.assertFalse([st for st in datapoint if st.startswith("delete_key")]) + def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: """value_double is a datapoint path; the collector refuses to start on it here.""" blocks = _config()["processors"]["transform/modelplane"]["metric_statements"] From ef3313eb21229acee50d74a5fce64d38c8d9bcc6 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:51:45 -0700 Subject: [PATCH 31/42] Scrape the decode engine and the endpoint picker Two producers the collector could never reach, each making a mapping that matches nothing. The engine scrape job keeps a pod on a container port named http, because matching by number would find the pd-sidecar rather than the engine behind it. Fronting a decode engine with that sidecar replaced the engine's ports with an unnamed one, and the sidecar takes its port unnamed too, so a disaggregated decode pod carried no named port at all. A PrefillDecode deployment reported half its engines, and the missing half looked like idle capacity. Nothing scraped the endpoint picker either: no identity labels, no metrics port. It now carries both, so the same job discovers it and a series off it joins the deployment it routes for. Its metrics endpoint authenticates callers by TokenReview by default, which needs a ClusterRole its namespaced ServiceAccount cannot hold, so every scrape would have been rejected; --metrics-endpoint-auth=false leaves it plain HTTP on the pod network, the way every engine already serves /metrics. --secure-serving is untouched - that one is the TLS of the ext-proc gRPC server Envoy calls. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../function/backends/base.py | 17 +++++++ .../compose-model-replica/function/routing.py | 48 +++++++++++++++-- .../tests/test_backends.py | 51 +++++++++++++++++++ 3 files changed, 111 insertions(+), 5 deletions(-) diff --git a/functions/compose-model-replica/function/backends/base.py b/functions/compose-model-replica/function/backends/base.py index 88d19a47c..ebe9cd1cb 100644 --- a/functions/compose-model-replica/function/backends/base.py +++ b/functions/compose-model-replica/function/backends/base.py @@ -291,6 +291,23 @@ def telemetry_labels( return labels +def fleet_labels(replica: v1alpha1.ModelReplica, role: str) -> dict[str, str]: + """What a metric off a non-serving pod of this replica is attributed to. + + The picker is one of these: it belongs to a replica of a deployment and + carries no engine, because it serves no model. Same labels as a serving + pod otherwise, so the collector reads it off the pod with the rules it + already has and a series off the picker joins the deployment's. + """ + labels: dict[str, str] = {LABEL_ROLE: role} + own = (replica.metadata.labels if replica.metadata else None) or {} + if name := own.get(LABEL_DEPLOYMENT): + labels[LABEL_DEPLOYMENT] = name + if index := own.get(_LABEL_REPLICA_INDEX): + labels[LABEL_REPLICA] = index + return labels + + def pod_metadata( member: v1alpha1.Member, labels: dict[str, str] | None = None, diff --git a/functions/compose-model-replica/function/routing.py b/functions/compose-model-replica/function/routing.py index 68bdd05a5..2be71ad6b 100644 --- a/functions/compose-model-replica/function/routing.py +++ b/functions/compose-model-replica/function/routing.py @@ -66,6 +66,12 @@ def _namespace(meta: metav1.ObjectMeta | None) -> str: # config, so the mount path never needs to encode which one. _EPP_CONFIG_FILE = "epp-config.yaml" +# The picker's Prometheus endpoint, and what a series off it is attributed to. +# 9090 is the picker's own default; it is named here because the port has to be +# declared on the pod for the collector to discover it either way. +_EPP_METRICS_PORT = 9090 +_EPP_ROLE = "picker" + # The pd-sidecar takes ENGINE_PORT (8000), so the decode engine listens here. _DECODE_ENGINE_PORT = 8001 @@ -279,7 +285,7 @@ def _unified( ns = base.remote_namespace(replica) out["inference-pool"] = base.wrap_object(provider_config, _inference_pool(name, ns, selector)) out[base.ROUTE_KEY] = base.wrap_object(provider_config, _http_route(replica, name)) - out.update(_epp_objects(name, ns, provider_config, _unified_epp_config_yaml(block_size))) + out.update(_epp_objects(name, ns, provider_config, _unified_epp_config_yaml(block_size), replica)) return out @@ -328,7 +334,7 @@ def _disaggregated( ns = base.remote_namespace(replica) out["inference-pool"] = base.wrap_object(provider_config, _inference_pool(name, ns, selector)) out[base.ROUTE_KEY] = base.wrap_object(provider_config, _http_route(replica, name)) - out.update(_epp_objects(name, ns, provider_config, _disaggregated_epp_config_yaml(block_size))) + out.update(_epp_objects(name, ns, provider_config, _disaggregated_epp_config_yaml(block_size), replica)) return out @@ -392,7 +398,12 @@ def _add_sidecar_to_decode(obj: k8sobjv1alpha1.Object) -> None: containers = tmpl["spec"]["containers"] engine = next(c for c in containers if c["name"] == "engine") port = _decode_port(engine) - engine["ports"] = [{"containerPort": port}] + # Named, because the collector's engine scrape job selects on the port + # name: dropping it here would leave a disaggregated decode pod with no + # named port at all, and nothing would scrape the engine behind the + # sidecar. The sidecar's own port stays unnamed for the same reason - + # it serves inference, not /metrics. + engine["ports"] = [{"name": base.ENGINE_PORT_NAME, "containerPort": port}] engine["readinessProbe"] = { "httpGet": {"path": "/health", "port": port}, "initialDelaySeconds": 30, @@ -490,7 +501,13 @@ def _http_route(replica: v1alpha1.ModelReplica, name: str) -> dict: } -def _epp_objects(name: str, ns: str, provider_config: str, config_yaml: str) -> dict[str, k8sobjv1alpha1.Object]: +def _epp_objects( + name: str, + ns: str, + provider_config: str, + config_yaml: str, + replica: v1alpha1.ModelReplica, +) -> dict[str, k8sobjv1alpha1.Object]: """The endpoint picker: ServiceAccount, RBAC, ConfigMap, Deployment, Service. config_yaml is the rendered EndpointPickerConfig the picker runs with; it @@ -545,7 +562,11 @@ def _epp_objects(name: str, ns: str, provider_config: str, config_yaml: str) -> "selector": {"matchLabels": {"app": epp}}, "template": { "metadata": { - "labels": {"app": epp}, + # The identity labels beside the selector label: they are + # how the collector attributes a series off this pod to the + # deployment and replica it routes for. Nothing selects on + # them, so they can't collide with the selector above. + "labels": {"app": epp, **base.fleet_labels(replica, _EPP_ROLE)}, "annotations": {"modelplane.ai/epp-config-checksum": config_checksum}, }, "spec": { @@ -560,10 +581,27 @@ def _epp_objects(name: str, ns: str, provider_config: str, config_yaml: str) -> "--pool-group=inference.networking.k8s.io", f"--config-file=/config/{_EPP_CONFIG_FILE}", "--grpc-port=9002", + f"--metrics-port={_EPP_METRICS_PORT}", + # The picker's metrics endpoint authenticates + # its callers by TokenReview by default, which + # needs the system:auth-delegator ClusterRole + # its ServiceAccount does not hold and cannot + # be given namespace-scoped. Left on, every + # scrape is rejected. Off, the endpoint is + # plain HTTP on the pod network, which is how + # every engine already serves /metrics. + # + # --secure-serving is a different thing and + # stays on: it is the TLS of the ext-proc gRPC + # server Envoy calls, not of this. + "--metrics-endpoint-auth=false", ], "ports": [ {"name": "grpc", "containerPort": 9002}, {"name": "grpc-health", "containerPort": 9003}, + # Named for the collector's engine scrape job, + # which keeps a pod on this name. + {"name": base.ENGINE_PORT_NAME, "containerPort": _EPP_METRICS_PORT}, ], "volumeMounts": [{"name": "config", "mountPath": "/config"}], } diff --git a/functions/compose-model-replica/tests/test_backends.py b/functions/compose-model-replica/tests/test_backends.py index 01fe61269..79728fa0b 100644 --- a/functions/compose-model-replica/tests/test_backends.py +++ b/functions/compose-model-replica/tests/test_backends.py @@ -955,6 +955,44 @@ def test_replaces_unified_service_with_pool_and_epp(self) -> None: self.assertEqual(pool["kind"], "InferencePool") self.assertEqual(pool["spec"]["endpointPickerRef"]["name"], "r-epp") + def test_the_picker_is_scrapeable_and_attributed(self) -> None: + """A built-in MetricMapping renames the picker's scheduling latency. + + Nothing can match it unless something scrapes the picker, and the + collector's engine job keeps a pod on two things: the deployment + label, and a container port named `http`. Without both, the mapping + is config that matches nothing for the life of the fleet. + + Its metrics endpoint authenticates callers by TokenReview by default, + which needs a ClusterRole the picker's namespaced ServiceAccount + cannot hold, so every scrape would be rejected. --secure-serving is + left alone: that one is the ext-proc gRPC server Envoy calls. + """ + replica = _replica() + replica.metadata = metav1.ObjectMeta( + name="r", + namespace="ml-team", + labels={base.LABEL_DEPLOYMENT: "qwen3-8b", "modelplane.ai/replica-index": "2"}, + ) + replica.spec.serving = v1alpha1.Serving(mode="Unified") + composed = {} + for engine in replica.spec.engines: + composed.update(native.NativeBackend().build(replica, engine, _PC, base.serving_label(replica), "Standard")) + template = routing.apply(composed, replica, _PC)["epp"].spec.forProvider.manifest["spec"]["template"] + + labels = template["metadata"]["labels"] + self.assertEqual(labels[base.LABEL_DEPLOYMENT], "qwen3-8b") + self.assertEqual(labels[base.LABEL_REPLICA], "2") + self.assertEqual(labels[base.LABEL_ROLE], "picker") + # Still the Deployment's own selector label, which must not move. + self.assertEqual(labels["app"], "r-epp") + + container = next(c for c in template["spec"]["containers"] if c["name"] == "epp") + self.assertIn({"name": base.ENGINE_PORT_NAME, "containerPort": 9090}, container["ports"]) + self.assertIn("--metrics-port=9090", container["args"]) + self.assertIn("--metrics-endpoint-auth=false", container["args"]) + self.assertNotIn("--secure-serving=false", container["args"]) + def test_injects_nixl_plumbing(self) -> None: """Both disagg engines get the NIXL plumbing the schema can't express: a Memory /dev/shm and VLLM_NIXL_SIDE_CHANNEL_HOST = pod IP.""" @@ -1059,6 +1097,19 @@ def test_decode_gets_sidecar_and_moves_engine_port(self) -> None: self.assertEqual(sidecar["readinessProbe"]["timeoutSeconds"], 5) self.assertIn("--secure-proxy=false", sidecar["args"]) + def test_a_decode_engine_keeps_a_scrapeable_port(self) -> None: + """The collector's engine job keeps a pod on the port named `http`. + + Moving the decode engine off 8000 for the sidecar drops the name with + it, and the sidecar takes the port unnamed because it serves inference + rather than /metrics. A decode pod with no named port anywhere is one + nothing scrapes, so a disaggregated deployment reports half its + engines and the shortfall looks like idle capacity. + """ + containers = self._serving_pod(self._apply(), "decode")["spec"]["containers"] + engine = next(c for c in containers if c["name"] == "engine") + self.assertEqual(engine["ports"], [{"name": base.ENGINE_PORT_NAME, "containerPort": 8001}]) + def test_prefill_has_no_sidecar(self) -> None: containers = self._serving_pod(self._apply(), "prefill")["spec"]["containers"] self.assertEqual([c["name"] for c in containers], ["engine"]) From 9a368aa6ce8c30b6daa775472564986819117251 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:52:36 -0700 Subject: [PATCH 32/42] Stop the collector gating the serving stack's readiness The collector counted towards the ServingStack the same way every serving component does, so a collector that could not start - a sink naming an exporter that does not exist, an endpoint that has gone away - took the ServingStack and then the InferenceCluster out of Ready, and the scheduler stopped placing replicas there. One bad TelemetryDestination is fleet-wide, so that is every cluster at once: a telemetry mistake taking down serving. The collector observes the fleet and nothing serving depends on it, so it is Ready on arrival. A collector that cannot start still says so on its own objects and on the TelemetryDestination, which is where that belongs. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-serving-stack/function/fn.py | 21 +++++--- .../compose-serving-stack/tests/test_fn.py | 50 +++++++++++++++++++ 2 files changed, 63 insertions(+), 8 deletions(-) diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index af2be852d..b69a69f48 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -338,7 +338,7 @@ def compose(self) -> None: rendered = self.compose_components(components) rendered += self.compose_gateway() rendered += self.compose_gateway_pki() - rendered += self.compose_collector() + self.compose_collector() self.compose_component_usages(components) self.compose_gateway_usages() self.write_status() @@ -763,7 +763,7 @@ def compose_gateway_pki(self) -> list[str]: rendered.append("gateway-client-auth") return rendered - def compose_collector(self) -> list[str]: + def compose_collector(self) -> None: """Compose the collector that gathers this cluster's telemetry. Nothing until a TelemetryDestination exists. Neither collector stores @@ -771,7 +771,14 @@ def compose_collector(self) -> list[str]: CPU spent on samples nobody will ever read; a fleet that has not said where its telemetry goes gets none composed. - Returns the composed-resource keys it rendered, for readiness. + Ready on arrival, unlike everything else the stack composes. The + collector observes the fleet; nothing serving depends on it. Gating the + stack on it would put the fleet's ability to place a replica behind its + ability to export a metric, so one TelemetryDestination naming an + endpoint that has gone away would leave every InferenceCluster in the + fleet not Ready and stop the scheduler placing anything, anywhere. A + collector that cannot start reports it on its own objects and on the + TelemetryDestination, which is where that failure belongs. """ response.require_resources( self.rsp, @@ -786,7 +793,7 @@ def compose_collector(self) -> list[str]: kind="MetricMapping", ) if "destinations" not in self.req.required_resources or "mappings" not in self.req.required_resources: - return [] + return # Sorted, not whichever the API server listed first: the collector # restarts on a change to its rendered config, so an unstable choice @@ -799,7 +806,7 @@ def compose_collector(self) -> list[str]: key=lambda d: _name(d.metadata), ) if not destinations: - return [] + return dest = destinations[0] if len(destinations) > 1: # Which one wins would otherwise be whichever the API server listed @@ -819,7 +826,6 @@ def compose_collector(self) -> list[str]: pc_observed = self.provider_configs_observed() pc = _pc_name(self.xr) - rendered: list[str] = [] for key, manifest, cel in collector.objects( cluster=_cluster_name(self.xr), mappings=mappings, @@ -837,8 +843,7 @@ def compose_collector(self) -> list[str]: ready_when=cel, ), ) - rendered.append(key) - return rendered + self.rsp.desired.resources[key].ready = fnv1.READY_TRUE def compose_gateway_usages(self) -> None: """Compose Usages ordering the hand-rendered gateway teardown. diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index bf4a46801..e0de06c8a 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -1490,3 +1490,53 @@ async def test_composed_resource_keys(self) -> None: ) got = await self.runner.RunFunction(_request(cloud, stack, observed=observed), None) self.assertEqual(expected, set(got.desired.resources.keys())) + + +class TestCollectorReadiness(unittest.IsolatedAsyncioTestCase): + """The collector is composed, but the stack never waits on it.""" + + maxDiff = None + + @classmethod + def setUpClass(cls) -> None: + cls.runner = fn.FunctionRunner() + + @staticmethod + def _with_destination(req: fnv1.RunFunctionRequest) -> fnv1.RunFunctionRequest: + req.required_resources["destinations"].items.append( + fnv1.Resource( + resource=resource.dict_to_struct( + { + "apiVersion": "modelplane.ai/v1alpha1", + "kind": "TelemetryDestination", + "metadata": {"name": "acme"}, + "spec": { + "sinks": [{"name": "primary", "type": "otlphttp", "endpoint": "https://otel.acme.example"}] + }, + } + ) + ) + ) + req.required_resources["mappings"].items.extend([]) + return req + + async def test_the_collector_does_not_gate_the_stack(self) -> None: + """A collector nothing has observed yet is still Ready. + + Everything else the stack composes is Ready only once its observed + Ready condition says so, because the fleet cannot serve without it. + The collector only watches, so gating on it would put placing a + replica behind exporting a metric: one destination pointing at an + endpoint that has gone away would take every InferenceCluster in the + fleet out of Ready and stop the scheduler. + """ + req = self._with_destination(_request("GKE", "Standard", observed=_observed_pcs())) + got = await self.runner.RunFunction(req, None) + collector_keys = [k for k in got.desired.resources if k == "collector" or k.startswith("collector-")] + # The Deployment, which is the one with a readiness CEL of its own and + # so the one that would have gated the stack. + self.assertIn("collector", collector_keys) + for key in collector_keys: + with self.subTest(key=key): + self.assertNotIn(key, req.observed.resources) + self.assertEqual(got.desired.resources[key].ready, fnv1.READY_TRUE) From ea13a087ba9c37ec33939ccb9720c5d04623910f Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:53:09 -0700 Subject: [PATCH 33/42] Copy a sink's credential to the clusters that mount it A TelemetryDestination is cluster-scoped on the control plane and its sinks name one Secret. The collector runs on every workload cluster, and its Deployment mounts that Secret by name from modelplane-system there - a namespace nothing copied it into. Every destination with a credential composed a pod that could never start, which is every destination that reaches a real backend. The destination reported Ready throughout, because it had resolved the Secret it could see. The serving stack now resolves it and composes a copy onto each cluster, the way compose-model-cache already propagates a ModelCache's token. Same name and namespace out there, so the mount needs nothing rewritten, and the base64 data is copied verbatim rather than decoded and re-encoded. A Secret that resolves to nothing still composes the collector. The destination reports it, and the pod waits the way any pod waits for a Secret, which recovers the moment the operator creates one. The destination's own check asked for the Secret with no namespace, so it would have accepted one of that name in any namespace while the one it meant was absent. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-serving-stack/function/fn.py | 51 +++++++++++++++++ .../compose-serving-stack/tests/test_fn.py | 56 +++++++++++++++++-- .../function/fn.py | 7 +++ .../tests/test_fn.py | 13 +++-- 4 files changed, 118 insertions(+), 9 deletions(-) diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index b69a69f48..7b0169d70 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -817,6 +817,35 @@ def compose_collector(self) -> None: f"{_name(dest.metadata)}. Telemetry has one destination per fleet.", ) + # The collector mounts each sink's credential as a Secret in its own + # namespace on this cluster, and the operator wrote one Secret, on the + # control plane. Resolve it here so it can be composed out there: + # without this the Deployment mounts a Secret nobody creates and the + # pod never starts, which is the one failure mode a destination with a + # credential would always hit. + secrets: dict[str, tuple[str, dict]] = {} + for sink in dest.spec.sinks: + if sink.secretRef is None: + continue + key = f"collector-secret-{sink.name}" + response.require_resources( + self.rsp, + name=key, + api_version="v1", + kind="Secret", + match_name=sink.secretRef.name, + namespace=_CLUSTER_RESOURCE_NAMESPACE, + ) + if key not in self.req.required_resources: + return + found = request.get_required_resource(self.req, key) + # Resolved and absent is not a reason to withhold the collector. + # The TelemetryDestination reports a Secret that isn't there, and + # the pod waits for it the way any pod waits for a Secret, which is + # recoverable the moment the operator creates it. + if found and found.get("data"): + secrets[key] = (sink.secretRef.name, found["data"]) + # Modelplane's own mappings first, then the operator's, which add to # them rather than replacing them. mappings = list(stacks.BUILTIN_MAPPINGS) @@ -845,6 +874,28 @@ def compose_collector(self) -> None: ) self.rsp.desired.resources[key].ready = fnv1.READY_TRUE + for key, (secret_name, data) in secrets.items(): + if not (pc_observed or key in self.req.observed.resources): + continue + # Same name, same namespace, other cluster, so the Deployment's + # mount needs nothing rewritten. The base64 `data` is copied + # verbatim: decoding and re-encoding would corrupt a credential + # that isn't text. + resource.update( + self.rsp.desired.resources[key], + _k8s_object( + pc, + { + "apiVersion": "v1", + "kind": "Secret", + "metadata": {"name": secret_name, "namespace": collector.NAMESPACE}, + "data": data, + }, + metadata=metav1.ObjectMeta(labels={_LABEL_RESOURCE: key}), + ), + ) + self.rsp.desired.resources[key].ready = fnv1.READY_TRUE + def compose_gateway_usages(self) -> None: """Compose Usages ordering the hand-rendered gateway teardown. diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index e0de06c8a..85946d8a1 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -1502,7 +1502,10 @@ def setUpClass(cls) -> None: cls.runner = fn.FunctionRunner() @staticmethod - def _with_destination(req: fnv1.RunFunctionRequest) -> fnv1.RunFunctionRequest: + def _with_destination(req: fnv1.RunFunctionRequest, *, secret: str | None = None) -> fnv1.RunFunctionRequest: + sink: dict = {"name": "primary", "type": "otlphttp", "endpoint": "https://otel.acme.example"} + if secret: + sink |= {"secretRef": {"name": secret}, "auth": {"bearerTokenKey": "token"}} req.required_resources["destinations"].items.append( fnv1.Resource( resource=resource.dict_to_struct( @@ -1510,14 +1513,25 @@ def _with_destination(req: fnv1.RunFunctionRequest) -> fnv1.RunFunctionRequest: "apiVersion": "modelplane.ai/v1alpha1", "kind": "TelemetryDestination", "metadata": {"name": "acme"}, - "spec": { - "sinks": [{"name": "primary", "type": "otlphttp", "endpoint": "https://otel.acme.example"}] - }, + "spec": {"sinks": [sink]}, } ) ) ) req.required_resources["mappings"].items.extend([]) + if secret: + req.required_resources["collector-secret-primary"].items.append( + fnv1.Resource( + resource=resource.dict_to_struct( + { + "apiVersion": "v1", + "kind": "Secret", + "metadata": {"name": secret, "namespace": "modelplane-system"}, + "data": {"token": "c2hoaGg="}, + } + ) + ) + ) return req async def test_the_collector_does_not_gate_the_stack(self) -> None: @@ -1540,3 +1554,37 @@ async def test_the_collector_does_not_gate_the_stack(self) -> None: with self.subTest(key=key): self.assertNotIn(key, req.observed.resources) self.assertEqual(got.desired.resources[key].ready, fnv1.READY_TRUE) + + async def test_the_credential_reaches_the_cluster_that_mounts_it(self) -> None: + """The operator writes one Secret; the collector mounts it elsewhere. + + A TelemetryDestination is cluster-scoped on the control plane and the + collector runs on every workload cluster in the fleet. Resolving the + Secret and stopping there leaves the Deployment mounting a name + nothing out there creates, so the pod never starts and the fleet + exports nothing - the failure every destination with a credential + would hit, which is every destination that reaches a real backend. + """ + req = self._with_destination( + _request("GKE", "Standard", observed=_observed_pcs()), secret="telemetry-credentials" + ) + got = await self.runner.RunFunction(req, None) + self.assertIn("collector-secret-primary", got.desired.resources) + composed = resource.struct_to_dict(got.desired.resources["collector-secret-primary"].resource) + manifest = composed["spec"]["forProvider"]["manifest"] + self.assertEqual(manifest["kind"], "Secret") + self.assertEqual(manifest["metadata"]["name"], "telemetry-credentials") + self.assertEqual(manifest["metadata"]["namespace"], "modelplane-system") + # Copied verbatim: re-encoding would corrupt a credential that is not + # text, and the mount reads the same key the sink's auth names. + self.assertEqual(manifest["data"], {"token": "c2hoaGg="}) + + async def test_a_destination_asks_for_its_credential_in_one_namespace(self) -> None: + """Unqualified, the requirement matches a Secret of that name anywhere.""" + req = self._with_destination( + _request("GKE", "Standard", observed=_observed_pcs()), secret="telemetry-credentials" + ) + got = await self.runner.RunFunction(req, None) + selector = got.requirements.resources["collector-secret-primary"] + self.assertEqual(selector.match_name, "telemetry-credentials") + self.assertEqual(selector.namespace, "modelplane-system") diff --git a/functions/compose-telemetry-destination/function/fn.py b/functions/compose-telemetry-destination/function/fn.py index 2c026234b..89d237822 100644 --- a/functions/compose-telemetry-destination/function/fn.py +++ b/functions/compose-telemetry-destination/function/fn.py @@ -37,6 +37,12 @@ _SECRET_PREFIX = "secret-" +# Where a sink's credential lives on the control plane. Unqualified, the +# requirement resolves a Secret of that name in any namespace, so a +# destination would accept a credential that happens to share a name with one +# in some unrelated namespace while the one it meant is absent. +_NAMESPACE = "modelplane-system" + class FunctionRunner(grpcv1.FunctionRunnerServiceServicer): """A FunctionRunner handles gRPC RunFunctionRequests.""" @@ -84,6 +90,7 @@ async def RunFunction( api_version="v1", kind="Secret", match_name=sink.secretRef.name, + namespace=_NAMESPACE, ) if key not in req.required_resources: _not_ready(rsp, CONDITION_REASON_WAITING_FOR_SECRET, "Waiting for the credential Secret to resolve") diff --git a/functions/compose-telemetry-destination/tests/test_fn.py b/functions/compose-telemetry-destination/tests/test_fn.py index f6c8c6ede..02367c924 100644 --- a/functions/compose-telemetry-destination/tests/test_fn.py +++ b/functions/compose-telemetry-destination/tests/test_fn.py @@ -12,7 +12,7 @@ # See the License for the specific language governing permissions and # limitations under the License. -"""Tests for the compose-metric-mapping function.""" +"""Tests for the compose-telemetry-destination function.""" import dataclasses import unittest @@ -27,7 +27,7 @@ @dataclasses.dataclass class Case: - """A test case for compose-metric-mapping.""" + """A test case for compose-telemetry-destination.""" name: str req: fnv1.RunFunctionRequest @@ -54,12 +54,12 @@ def sink(name: str = "primary", type_: str = "otlphttp", secret: str | None = No "name": name, "type": type_, "endpoint": "https://otel.acme.example", - "config": {"auth": {"authenticator": "oidc/acme"}}, + "config": {"auth": {"authenticator": "oauth2client/acme"}}, **({"secretRef": {"name": secret}} if secret else {}), } sinks = [sink()] - extensions = {"oidc/acme": {"issuer_url": "https://issuer.acme.example"}} + extensions = {"oauth2client/acme": {"token_url": "https://issuer.acme.example/token"}} def xr(spec: dict) -> dict: return { @@ -93,6 +93,9 @@ def want( rsp.requirements.resources["secret-primary"].api_version = "v1" rsp.requirements.resources["secret-primary"].kind = "Secret" rsp.requirements.resources["secret-primary"].match_name = secret + # Qualified: unqualified it would resolve a Secret of that + # name in any namespace, and accept the wrong credential. + rsp.requirements.resources["secret-primary"].namespace = "modelplane-system" return rsp cases = [ @@ -215,7 +218,7 @@ def want( type="Accepted", status=fnv1.STATUS_CONDITION_FALSE, reason="UnknownAuthenticator", - message="No extension defines oidc/acme, so the collector would refuse to start", + message="No extension defines oauth2client/acme, so the collector would refuse to start", ), ), ), From a30b21a1fa934cbc6688eccd266708b30641fa03 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:53:20 -0700 Subject: [PATCH 34/42] Build the collector's config from every TelemetryDestination The fleet exported through one destination - whichever sorted first - and warned about the rest. Adding a second backend therefore meant editing an object another team owned, and creating your own got a warning and silence. Their sinks and extensions are concatenated instead, so a destination is an additive object rather than a singleton. A sink's name is what the collector calls its exporter, so two destinations cannot share one: the destination sorting first keeps the name and the other is dropped with a warning, because taking the whole fleet's telemetry down over a name collision is the worse failure. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-serving-stack/function/fn.py | 61 ++++++++++++---- .../compose-serving-stack/tests/test_fn.py | 72 +++++++++++++++++++ 2 files changed, 119 insertions(+), 14 deletions(-) diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index 7b0169d70..9303c0dd5 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -34,6 +34,8 @@ ahead of the Envoy Gateway release. """ +from typing import Any + import grpc from crossplane.function import logging, request, resource, response from crossplane.function.proto.v1 import run_function_pb2 as fnv1 @@ -796,8 +798,8 @@ def compose_collector(self) -> None: return # Sorted, not whichever the API server listed first: the collector - # restarts on a change to its rendered config, so an unstable choice - # between two destinations would redeploy it on alternate reconciles. + # restarts on a change to its rendered config, so an unstable order + # would redeploy it on alternate reconciles. destinations = sorted( ( tdv1alpha1.TelemetryDestination.model_validate(d) @@ -807,15 +809,7 @@ def compose_collector(self) -> None: ) if not destinations: return - dest = destinations[0] - if len(destinations) > 1: - # Which one wins would otherwise be whichever the API server listed - # first, and a fleet would export somewhere nobody chose. - response.warning( - self.rsp, - f"{len(destinations)} TelemetryDestinations exist; using " - f"{_name(dest.metadata)}. Telemetry has one destination per fleet.", - ) + sinks, extensions = self.merge_destinations(destinations) # The collector mounts each sink's credential as a Secret in its own # namespace on this cluster, and the operator wrote one Secret, on the @@ -824,7 +818,7 @@ def compose_collector(self) -> None: # pod never starts, which is the one failure mode a destination with a # credential would always hit. secrets: dict[str, tuple[str, dict]] = {} - for sink in dest.spec.sinks: + for sink in sinks: if sink.secretRef is None: continue key = f"collector-secret-{sink.name}" @@ -858,8 +852,8 @@ def compose_collector(self) -> None: for key, manifest, cel in collector.objects( cluster=_cluster_name(self.xr), mappings=mappings, - sinks=list(dest.spec.sinks), - extensions=dict(dest.spec.extensions or {}), + sinks=sinks, + extensions=extensions, ): if not (pc_observed or key in self.req.observed.resources): continue @@ -896,6 +890,45 @@ def compose_collector(self) -> None: ) self.rsp.desired.resources[key].ready = fnv1.READY_TRUE + def merge_destinations( + self, destinations: list[tdv1alpha1.TelemetryDestination] + ) -> tuple[list[tdv1alpha1.Sink], dict[str, Any]]: + """Every destination's sinks and extensions, as one collector config. + + Concatenated rather than one of them chosen, so a second backend is a + second object rather than an edit to a singleton somebody else owns. + Every sink gets the whole stream either way, so the fleet exports to + all of them. + + A sink names the collector's exporter instance, and two destinations + naming a sink the same way would be one exporter with two meanings. + The destination that sorts first keeps the name and the other is + dropped with a warning, because taking the whole fleet's telemetry + down over a name collision is the worse failure. Extensions collide + the same way and resolve the same way. + """ + sinks: dict[str, tdv1alpha1.Sink] = {} + extensions: dict[str, Any] = {} + clashes: list[str] = [] + for dest in destinations: + name = _name(dest.metadata) + for sink in dest.spec.sinks: + if sink.name in sinks: + clashes.append(f"sink {sink.name} in TelemetryDestination {name}") + continue + sinks[sink.name] = sink + for key, value in (dest.spec.extensions or {}).items(): + if key in extensions: + clashes.append(f"extension {key} in TelemetryDestination {name}") + continue + extensions[key] = value + if clashes: + response.warning( + self.rsp, + f"Ignored {', '.join(clashes)}: already defined by a TelemetryDestination sorting earlier.", + ) + return list(sinks.values()), extensions + def compose_gateway_usages(self) -> None: """Compose Usages ordering the hand-rendered gateway teardown. diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index 85946d8a1..527f7f5f7 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -1588,3 +1588,75 @@ async def test_a_destination_asks_for_its_credential_in_one_namespace(self) -> N selector = got.requirements.resources["collector-secret-primary"] self.assertEqual(selector.match_name, "telemetry-credentials") self.assertEqual(selector.namespace, "modelplane-system") + + async def test_every_destination_contributes_its_sinks(self) -> None: + """A second backend is a second object, not an edit to a singleton. + + Picking one destination and warning about the rest means a team + adding an export has to edit an object another team owns, and gets + silence if they create their own instead. + """ + req = _request("GKE", "Standard", observed=_observed_pcs()) + for name, sink in ( + ("acme", {"name": "vendor", "type": "otlphttp", "endpoint": "https://otel.vendor.example"}), + ("zeta", {"name": "prom", "type": "prometheus_remote_write", "endpoint": "https://p.example/w"}), + ): + req.required_resources["destinations"].items.append( + fnv1.Resource( + resource=resource.dict_to_struct( + { + "apiVersion": "modelplane.ai/v1alpha1", + "kind": "TelemetryDestination", + "metadata": {"name": name}, + "spec": {"sinks": [sink]}, + } + ) + ) + ) + req.required_resources["mappings"].items.extend([]) + got = await self.runner.RunFunction(req, None) + config = yaml.safe_load( + resource.struct_to_dict(got.desired.resources["collector-config"].resource)["spec"]["forProvider"][ + "manifest" + ]["data"]["collector.yaml"] + ) + self.assertEqual( + sorted(config["exporters"]), + ["otlphttp/vendor", "prometheus_remote_write/prom"], + ) + self.assertEqual( + sorted(config["service"]["pipelines"]["metrics"]["exporters"]), + ["otlphttp/vendor", "prometheus_remote_write/prom"], + ) + + async def test_two_destinations_cannot_name_one_exporter(self) -> None: + """A sink names the collector's exporter instance. + + Two of them under one name is one exporter with two meanings. The + destination sorting first keeps it and the other is dropped with a + warning, rather than failing the whole fleet's telemetry over a name. + """ + req = _request("GKE", "Standard", observed=_observed_pcs()) + for name, endpoint in (("acme", "https://a.example"), ("zeta", "https://z.example")): + req.required_resources["destinations"].items.append( + fnv1.Resource( + resource=resource.dict_to_struct( + { + "apiVersion": "modelplane.ai/v1alpha1", + "kind": "TelemetryDestination", + "metadata": {"name": name}, + "spec": {"sinks": [{"name": "primary", "type": "otlphttp", "endpoint": endpoint}]}, + } + ) + ) + ) + req.required_resources["mappings"].items.extend([]) + got = await self.runner.RunFunction(req, None) + config = yaml.safe_load( + resource.struct_to_dict(got.desired.resources["collector-config"].resource)["spec"]["forProvider"][ + "manifest" + ]["data"]["collector.yaml"] + ) + self.assertEqual(list(config["exporters"]), ["otlphttp/primary"]) + self.assertEqual(config["exporters"]["otlphttp/primary"]["endpoint"], "https://a.example") + self.assertTrue([r for r in got.results if "zeta" in r.message]) From 9e6c9c78fb6e488500524b0365833b5a6e928f5b Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:53:30 -0700 Subject: [PATCH 35/42] Stop the composed Endpoints being mirrored into a second slice A cluster's gateway address is published as both an EndpointSlice and the legacy Endpoints kube-dns still reads. The endpointslice mirroring controller copies a hand-written Endpoints into an EndpointSlice of its own, and the slice composed beside it already is that slice, so two writers kept correcting each other's object for the same Service. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-inference-gateway/function/fn.py | 12 +++++++++++- .../compose-inference-gateway/tests/test_fn.py | 16 ++++++++++++++-- 2 files changed, 25 insertions(+), 3 deletions(-) diff --git a/functions/compose-inference-gateway/function/fn.py b/functions/compose-inference-gateway/function/fn.py index 33257c1e7..135d872db 100644 --- a/functions/compose-inference-gateway/function/fn.py +++ b/functions/compose-inference-gateway/function/fn.py @@ -878,6 +878,12 @@ def compose_cluster_name(self, hostname: str, address: str) -> None: # # Deprecated since 1.33 and still the only thing kube-dns reads. It # costs one object; getting it wrong costs every request to the cluster. + # + # Mirroring is turned off on it. The endpointslice mirroring controller + # copies a hand-written Endpoints into an EndpointSlice of its own, and + # the slice above already is that slice: left on, two controllers write + # the same address for the same Service and each keeps correcting the + # other's object. resource.update( self.rsp.desired.resources[f"cluster-name-endpoints-{label}"], _k8s_object( @@ -885,7 +891,11 @@ def compose_cluster_name(self, hostname: str, address: str) -> None: { "apiVersion": "v1", "kind": "Endpoints", - "metadata": {"name": label, "namespace": REMOTE_NAMESPACE}, + "metadata": { + "name": label, + "namespace": REMOTE_NAMESPACE, + "labels": {"endpointslice.kubernetes.io/skip-mirror": "true"}, + }, "subsets": [ { "addresses": [{"ip": address}], diff --git a/functions/compose-inference-gateway/tests/test_fn.py b/functions/compose-inference-gateway/tests/test_fn.py index f5b9b50cb..74703a0a9 100644 --- a/functions/compose-inference-gateway/tests/test_fn.py +++ b/functions/compose-inference-gateway/tests/test_fn.py @@ -819,7 +819,13 @@ async def test_resolves_each_cluster_gateway_name(self) -> None: "cluster-name-endpoints-prod-ipv4-gateway-aaaaa": { "apiVersion": "v1", "kind": "Endpoints", - "metadata": {"name": "prod-ipv4-gateway-aaaaa", "namespace": fn.REMOTE_NAMESPACE}, + "metadata": { + "name": "prod-ipv4-gateway-aaaaa", + "namespace": fn.REMOTE_NAMESPACE, + # Off, or the mirroring controller writes a second + # EndpointSlice over the one composed beside this. + "labels": {"endpointslice.kubernetes.io/skip-mirror": "true"}, + }, "subsets": [{"addresses": [{"ip": "203.0.113.7"}], "ports": [{"name": "https", "port": 443}]}], }, "cluster-name-slice-prod-ipv4-gateway-aaaaa": { @@ -843,7 +849,13 @@ async def test_resolves_each_cluster_gateway_name(self) -> None: "cluster-name-endpoints-prod-ipv6-gateway-bbbbb": { "apiVersion": "v1", "kind": "Endpoints", - "metadata": {"name": "prod-ipv6-gateway-bbbbb", "namespace": fn.REMOTE_NAMESPACE}, + "metadata": { + "name": "prod-ipv6-gateway-bbbbb", + "namespace": fn.REMOTE_NAMESPACE, + # Off, or the mirroring controller writes a second + # EndpointSlice over the one composed beside this. + "labels": {"endpointslice.kubernetes.io/skip-mirror": "true"}, + }, "subsets": [{"addresses": [{"ip": "2001:db8::1"}], "ports": [{"name": "https", "port": 443}]}], }, "cluster-name-slice-prod-ipv6-gateway-bbbbb": { From 8d00dc86c875d30d95881890c05d0c8d285b7a46 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 17:53:41 -0700 Subject: [PATCH 36/42] Correct what the telemetry APIs and guide claim Several descriptions described behaviour the code does not have. A later rename was said to win over a built-in. Each rename matches the name the component emitted, so once a built-in has renamed a metric the operator's rename matches nothing and the built-in stands. Taking a part of a histogram was said to leave the histogram exported under its own name, but only modelplane_* leaves a cluster, so it is dropped unless a second entry renames it too. The fleet-wide quantile example selected on a model label the collector does not add, and the MetricMapping example omitted acrossReplicas, which is required - as written the API would have rejected it. values is documented as meaningful only alongside from, and now a CEL rule says so rather than it being silently ignored beside a fixed value. A label's from naming the label itself is documented, since that is how a mapping rewrites values in place. The remote-write exporter is named prometheus_remote_write as of 0.161.0, with prometheusremotewrite a deprecated alias, so the examples use the canonical spelling. The oidc extension authenticates callers of a receiver rather than an exporter, so it is no longer offered as a way to authenticate a sink. endpoint and auth are documented as overriding config, which is what the renderer does, and the one default Modelplane applies to a sink is documented rather than left to be discovered. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 90 +++++++-------- apis/telemetrydestinations/definition.yaml | 104 +++++++++--------- docs/content/platform/telemetry.md | 29 +++-- e2e/README.md | 2 - schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 26 ++--- .../telemetrydestination/v1alpha1.py | 26 ++--- 7 files changed, 144 insertions(+), 135 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index 563476827..89f1b881a 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -37,20 +37,21 @@ spec: A mapping naming a component Modelplane already provides renames for is additive: its renames run after the built-in - ones, and a later rename of the same metric wins. + ones, on whatever those left behind. A metric a built-in + already renamed no longer answers to the name it was emitted + under, so a second mapping selecting on that name matches + nothing and the built-in stands. Select on the `modelplane_` + name instead to rename one of Modelplane's own. properties: metrics: type: array minItems: 1 maxItems: 128 description: >- - The metrics this component emits, and what Modelplane calls - them. - - Rename only where the measurements agree. Two engines' - histograms sharing a name are worth less than nothing if - their buckets disagree, because a quantile across them is - wrong rather than approximate. + The metrics to rename. Give two components' metrics the same + name only if they measure the same thing, and histograms only + if their buckets match too. A quantile across mismatched + buckets is wrong. items: type: object required: [from, to, acrossReplicas] @@ -60,24 +61,25 @@ spec: maxLength: 255 pattern: '^[a-zA-Z_:][a-zA-Z0-9_:]*$' description: >- - The metric's name as the component emits it, matched - exactly. Nothing here declares which engine a - deployment runs: a name that no component emits simply - matches nothing. + The metric's name as the component exposes it, matched + exactly wherever it appears in the fleet. - Held to the characters a metric name can contain. The - name is matched inside the collector's own query - language, so a quote here would end the comparison - early and rename whatever the rest of the line - matched. + A histogram is named by its base name, without the + _count, _sum or _bucket a Prometheus query would use: + the collector holds it as one metric, and `part` is + what reaches into it. to: type: string maxLength: 255 pattern: '^modelplane_[a-z0-9_]*[a-z0-9]$' description: >- - What Modelplane calls it. Only modelplane_* leaves a - cluster, so a metric with no name here is one nobody - downstream can read. + The name to export the metric under. Only modelplane_* + metrics leave a cluster, so a metric no mapping renames + never leaves its cluster. + + Name it in base units - seconds, bytes, joules, a ratio + from nought to one - because that is what `fromUnit` + converts to. acrossReplicas: type: string enum: [Sum, Mean, Max] @@ -114,26 +116,28 @@ spec: the histogram measures request duration. Sum is their total. - The histogram carries on unchanged under its own name. - This adds a series beside it. + The extraction leaves the histogram alone, but the + collector exports only what a mapping renames, so the + histogram itself is dropped unless another mapping + gives it a `modelplane_` name of its own. Write that + second mapping to keep both. labels: type: array maxItems: 16 description: >- - Labels to set on the series, for folding several - metrics into one that a label tells apart - tokens in - and out under one name with a direction, responses - under one name with the reason they ended. - - Two mappings writing the same `to` with a different - fixed value is how the fold is expressed: each renames - its own source and stamps its own value. + Labels to set on this metric's series. To fold several + metrics into one name, give each its own entry with + the same `to` and a different fixed value, such as + direction: input and direction: output on + modelplane_tokens_total. items: type: object required: [name] x-kubernetes-validations: - rule: "has(self.value) != has(self.from)" message: set either value, for a fixed label, or from, to carry one the component already emits. + - rule: "!has(self.values) || has(self.from)" + message: values remaps what from carries, so it needs from. A fixed value has nothing to remap. properties: name: type: string @@ -152,9 +156,13 @@ spec: maxLength: 63 pattern: '^[a-zA-Z_][a-zA-Z0-9_]*$' description: >- - A label the component already emits, carried onto - the new name and dropped from the series under - its old one. + A label the component already emits. Modelplane + copies its value into this label and removes the + original. + + Naming the label itself keeps it: that is how + `values` rewrites what a component writes without + renaming the label. values: type: object maxProperties: 32 @@ -166,8 +174,6 @@ spec: putting an engine's own vocabulary into Modelplane's. A value with no entry here is left as the component wrote it. - - Only meaningful alongside `from`. fromUnit: type: string enum: [Millijoules, Mebibytes, Milliseconds, Nanoseconds, Percent] @@ -176,27 +182,23 @@ spec: the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, percent by a - hundred, and mebibytes multiplied out to bytes. + hundred, and mebibytes multiplied out to bytes. A + histogram is converted whole - its sum, its bounds and + its bucket boundaries - so its quantiles come out in + the target unit too. Percent is for a component that counts a saturation from nought to a hundred where the name says a ratio. Check rather than assume: vLLM publishes kv_cache_usage_perc and the value is a fraction, so a name is no guide. - - Say it whenever the source disagrees with the target, - even where the factor looks obvious. A name ending in - _bytes that holds mebibytes is the kind of thing - nobody notices until a capacity review, and stating - the source unit is what makes the conversion happen at - all. status: type: object properties: clusters: type: integer description: >- - How many inference clusters have taken these statements. + How many inference clusters apply this mapping. conditions: type: array items: diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index 858034035..c2a57c7c8 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -27,14 +27,15 @@ spec: type: object required: [sinks] description: >- - Where the fleet's telemetry goes. Modelplane composes no - collectors until a TelemetryDestination exists: neither tier - stores anything, so collecting with nowhere to export would - spend GPU-cluster memory on samples nobody reads. Creating one - turns collection on everywhere at once. + Where the fleet's metrics go. Modelplane runs no collectors + until a TelemetryDestination exists, and creating one turns on + collection on every inference cluster. - There is no per-deployment opt-out. A ModelDeployment's author - owns neither the destination nor its bill. + Several can exist. Their sinks are concatenated into one + collector configuration, so adding a backend is a new object + rather than an edit to one somebody else owns. A sink's name is + the collector's name for its exporter, so it has to be unique + across destinations. properties: sinks: type: array @@ -68,7 +69,7 @@ spec: description: >- The collector exporter to send with, by the name OpenTelemetry gives it: otlphttp, otlp, - prometheusremotewrite, kafka, and every other one the + prometheus_remote_write, kafka, and every other one the collector provides. Not an enum, because enumerating them here would mean @@ -79,30 +80,24 @@ spec: type: string maxLength: 2048 description: >- - Where this sink writes. Typed rather than left to the - configuration below because it is the setting every - destination has to get right, and the one worth - catching here rather than in a collector that won't - start. + Where this sink writes. Leave it unset for an exporter + that doesn't take an endpoint, such as kafka or debug, + and configure it in `config` instead. - Optional, because not every exporter addresses its - destination this way: Kafka takes brokers, the file - exporter a path, and the debug exporter nothing at - all. Those go in the configuration below, under the - names that exporter gives them. + Set here, it wins: Modelplane applies it over + `config`, so a sink cannot be quietly redirected by + the configuration passed through beside it. auth: type: object description: >- - How to authenticate, for the schemes Modelplane - composes. The collector takes no credential inline: it - authenticates through an extension an exporter names, - so setting this composes that extension and wires the - reference. + Authentication Modelplane sets up for this sink, using + a credential from `secretRef`. For another scheme, + define an authenticator under spec.extensions and + reference it from `config`. - A scheme that isn't here is still reachable. Define - the extension yourself under spec.extensions and name - it from this sink's config, which is what Modelplane - does on your behalf. + Set here, it wins: Modelplane applies it over an + `auth` block in `config`, so a sink's credential + cannot be quietly unpicked. properties: bearerTokenKey: type: string @@ -116,46 +111,47 @@ spec: type: object x-kubernetes-preserve-unknown-fields: true description: >- - Anything else that exporter takes, passed through - unread: TLS, retry, queueing, compression, headers. + The rest of the exporter's configuration, passed + through as written: TLS, retries, queueing, + compression, headers. - Modelplane does not model an exporter's configuration, - because the schema is OpenTelemetry's and versioned - separately. Typing it would mean a Modelplane release - for each setting the collector gains, and would drop - the ones this has never heard of. What is typed above - is what belongs to Modelplane: which sinks exist, what - each is called, where it writes, and which Secret it - reads. + Modelplane sets one default, for the exporters that + flatten a series into labels: prometheus and + prometheus_remote_write get + resource_to_telemetry_conversion, or they would + receive every series stripped of the cluster, + deployment, engine and role it belongs to. Setting it + here overrides that. secretRef: type: object required: [name] description: >- - A Secret holding this sink's credential. Its keys - reach the collector as files under - /etc/modelplane/telemetry//, and as - environment variables, for configuration above - referring to ${env:TOKEN}. - - Per sink rather than per destination, so two sinks - with different credentials don't have to share one - Secret and tell their keys apart by prefix. The files - are per sink; the environment variables are not, so - two Secrets sharing a key name still collide there and - the file is the one to read. + A Secret holding this sink's credentials. Modelplane + mounts each key as a file under + /etc/modelplane/telemetry// and sets it as + an environment variable for ${env:KEY} references in + `config`. All sinks share one environment, so where + two Secrets have the same key, refer to the file. properties: name: type: string maxLength: 253 - description: Name of the Secret, in Modelplane's namespace. + description: >- + Name of the Secret, in modelplane-system on the + control plane. Modelplane copies it to every + cluster running a collector, so it doesn't have + to exist on each of them already. extensions: type: object x-kubernetes-preserve-unknown-fields: true description: >- - The collector's extensions block, passed through unread, - for the authenticator a sink references. Bearer token, - basic auth, OIDC and SigV4 all work, because none of them - is modelled here. + Collector extensions, passed through as written. Use it to + define an authenticator that a sink's `config` references. + + An exporter needs a client authenticator - basicauth, + oauth2client, sigv4auth, headers_setter. The oidc extension + authenticates callers of a receiver, so it is not one of + these. status: type: object properties: diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 1a5895475..53ae410a8 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -102,8 +102,10 @@ spec: bearerTokenKey: token ``` -It reads the token from a file rather than the environment, so rotating it doesn't need the -collector restarted. +Create that Secret once, in `modelplane-system` on the control plane. Modelplane copies it +to every cluster running a collector, so you don't put the credential on each GPU cluster +yourself. It reads the token from a file rather than the environment, so rotating it +doesn't need the collector restarted. If you run Prometheus, export to that instead and query the fleet there: @@ -111,7 +113,7 @@ If you run Prometheus, export to that instead and query the fleet there: spec: sinks: - name: prometheus - type: prometheusremotewrite + type: prometheus_remote_write endpoint: https://prom.example.internal/api/v1/write ``` @@ -129,11 +131,18 @@ spec: auth: bearerTokenKey: token - name: prometheus - type: prometheusremotewrite + type: prometheus_remote_write endpoint: https://prom.example.internal/api/v1/write ``` -That is two copies of the fleet's metrics, billed twice. +That is two copies of the fleet's metrics, so a vendor charging per sample charges for +both. + +Sinks can also come from more than one `TelemetryDestination`. Modelplane concatenates +them, so a team adding an export creates its own object rather than editing one somebody +else owns. Sink names are what the collector calls its exporters, so they have to be +unique across destinations; where two collide, the destination whose name sorts first +keeps it and Modelplane says so on the `ServingStack`. Anything else the exporter takes goes under `config`, passed through as you wrote it: @@ -171,7 +180,7 @@ produces no rates and no quantiles. Your backend does that. A fleet-wide p99: ```promql histogram_quantile(0.99, sum by (le) ( - rate(modelplane_frontend_ttft_seconds_bucket{model="Qwen/Qwen3-8B"}[5m]))) + rate(modelplane_frontend_ttft_seconds_bucket{deployment="qwen3-8b"}[5m]))) ``` Modelplane has no dashboards of its own. What it exports is counters and histogram buckets, and @@ -198,14 +207,20 @@ spec: metrics: - from: my_engine_queued_requests to: modelplane_requests_waiting + acrossReplicas: Sum - from: my_engine_kv_transfer_ms fromUnit: Milliseconds to: modelplane_request_kv_transfer_seconds + acrossReplicas: Mean ``` Modelplane renders every mapping into every cluster's collector, so you write one once. `from` is the name your engine emits and `to` is what Modelplane calls it. +`acrossReplicas` says how a query should combine the metric over a deployment's replicas, +since every pod publishes its own series. Use `Sum` for anything counted and `Mean` for a +ratio, where adding two replicas at half capacity would read as one at full. + Say `fromUnit` whenever the engine measures in something other than the unit the name claims, and Modelplane converts to the base one. Skipping it is the expensive mistake here: a series named `_seconds` that holds milliseconds reads a thousand times fast, and nothing @@ -242,7 +257,7 @@ store stops scraping and starts receiving. Same Prometheus, same retention, same spec: sinks: - name: prometheus - type: prometheusremotewrite + type: prometheus_remote_write config: endpoint: http://prometheus.monitoring.svc:9090/api/v1/write ``` diff --git a/e2e/README.md b/e2e/README.md index 43561583b..ff7aa6751 100644 --- a/e2e/README.md +++ b/e2e/README.md @@ -66,8 +66,6 @@ and `DCGM_FI_DEV_FB_USED` — so `--verify` asserts they arrive as MiB became 1073741824 bytes, that each carries its deployment, engine, role and cluster, and that the engine's own `vllm:` names did *not* leave the cluster. -It costs no extra wait: the destination is applied with every other manifest, so -the collector composes while the model is still rolling out. ### Why cloud provisioning cannot be tested here diff --git a/schemas/.lock.json b/schemas/.lock.json index 1e1c56c9c..7ae604486 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "7d6ebf1a8ad6d6797cff6bf790ff01378ba3ead27240e0ddb20f32536c9a66ea", + "fs://apis": "1d2cf24a31d015db1785d6f5921e560194a2985d8fe812e4c38f10972fe3c06e", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.8.1": "sha256:ca2e9e3b2e3a8b6ca44a9700d5abf7abd733cfa388d1afe9bb2bf7c847cd39ef", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.8.1": "sha256:ebb1bcd8dc9a7e60e97a324652fbd9d1b0609ddbb1107db518ce999cbc1113bd", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index 23d5a30d7..0d9318c62 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -47,7 +47,8 @@ class Label(BaseModel): None, alias='from' ) """ - A label the component already emits, carried onto the new name and dropped from the series under its old one. + A label the component already emits. Modelplane copies its value into this label and removes the original. + Naming the label itself keeps it: that is how `values` rewrites what a component writes without renaming the label. """ name: constr(pattern=r'^[a-zA-Z_][a-zA-Z0-9_]*$', max_length=63) """ @@ -60,7 +61,6 @@ class Label(BaseModel): values: dict[str, constr(max_length=253)] | None = Field(None, max_length=32) """ What each of that label's values becomes, for putting an engine's own vocabulary into Modelplane's. A value with no entry here is left as the component wrote it. - Only meaningful alongside `from`. """ @@ -76,31 +76,30 @@ class Metric(BaseModel): ..., alias='from' ) """ - The metric's name as the component emits it, matched exactly. Nothing here declares which engine a deployment runs: a name that no component emits simply matches nothing. - Held to the characters a metric name can contain. The name is matched inside the collector's own query language, so a quote here would end the comparison early and rename whatever the rest of the line matched. + The metric's name as the component exposes it, matched exactly wherever it appears in the fleet. + A histogram is named by its base name, without the _count, _sum or _bucket a Prometheus query would use: the collector holds it as one metric, and `part` is what reaches into it. """ fromUnit: ( Literal['Millijoules', 'Mebibytes', 'Milliseconds', 'Nanoseconds', 'Percent'] | None ) = None """ - What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, percent by a hundred, and mebibytes multiplied out to bytes. + What the component measures this in, when that isn't the unit the name claims. Modelplane converts to the base unit: millijoules and milliseconds are divided by a thousand, nanoseconds by a billion, percent by a hundred, and mebibytes multiplied out to bytes. A histogram is converted whole - its sum, its bounds and its bucket boundaries - so its quantiles come out in the target unit too. Percent is for a component that counts a saturation from nought to a hundred where the name says a ratio. Check rather than assume: vLLM publishes kv_cache_usage_perc and the value is a fraction, so a name is no guide. - Say it whenever the source disagrees with the target, even where the factor looks obvious. A name ending in _bytes that holds mebibytes is the kind of thing nobody notices until a capacity review, and stating the source unit is what makes the conversion happen at all. """ labels: list[Label] | None = Field(None, max_length=16) """ - Labels to set on the series, for folding several metrics into one that a label tells apart - tokens in and out under one name with a direction, responses under one name with the reason they ended. - Two mappings writing the same `to` with a different fixed value is how the fold is expressed: each renames its own source and stamps its own value. + Labels to set on this metric's series. To fold several metrics into one name, give each its own entry with the same `to` and a different fixed value, such as direction: input and direction: output on modelplane_tokens_total. """ part: Literal['Count', 'Sum'] | None = None """ Take a part of a histogram as a counter of its own, rather than the histogram itself. Count is how many observations it holds, which is a request count where the histogram measures request duration. Sum is their total. - The histogram carries on unchanged under its own name. This adds a series beside it. + The extraction leaves the histogram alone, but the collector exports only what a mapping renames, so the histogram itself is dropped unless another mapping gives it a `modelplane_` name of its own. Write that second mapping to keep both. """ to: constr(pattern=r'^modelplane_[a-z0-9_]*[a-z0-9]$', max_length=255) """ - What Modelplane calls it. Only modelplane_* leaves a cluster, so a metric with no name here is one nobody downstream can read. + The name to export the metric under. Only modelplane_* metrics leave a cluster, so a metric no mapping renames never leaves its cluster. + Name it in base units - seconds, bytes, joules, a ratio from nought to one - because that is what `fromUnit` converts to. """ @@ -111,8 +110,7 @@ class Spec(BaseModel): """ metrics: list[Metric] = Field(..., max_length=128, min_length=1) """ - The metrics this component emits, and what Modelplane calls them. - Rename only where the measurements agree. Two engines' histograms sharing a name are worth less than nothing if their buckets disagree, because a quantile across them is wrong rather than approximate. + The metrics to rename. Give two components' metrics the same name only if they measure the same thing, and histograms only if their buckets match too. A quantile across mismatched buckets is wrong. """ @@ -128,7 +126,7 @@ class Condition(BaseModel): class Status(BaseModel): clusters: int | None = None """ - How many inference clusters have taken these statements. + How many inference clusters apply this mapping. """ conditions: list[Condition] | None = None """ @@ -152,7 +150,7 @@ class MetricMapping(BaseModel): spec: Spec """ How one component's metrics become part of the modelplane_* surface. Modelplane renders every MetricMapping into each inference cluster's collector, so a mapping is written once on the control plane and reaches the whole fleet. - A mapping naming a component Modelplane already provides renames for is additive: its renames run after the built-in ones, and a later rename of the same metric wins. + A mapping naming a component Modelplane already provides renames for is additive: its renames run after the built-in ones, on whatever those left behind. A metric a built-in already renamed no longer answers to the name it was emitted under, so a second mapping selecting on that name matches nothing and the built-in stands. Select on the `modelplane_` name instead to rename one of Modelplane's own. """ status: Status | None = None diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index 11b278c92..2365e38fe 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -52,25 +52,25 @@ class Auth(BaseModel): class SecretRef(BaseModel): name: constr(max_length=253) """ - Name of the Secret, in Modelplane's namespace. + Name of the Secret, in modelplane-system on the control plane. Modelplane copies it to every cluster running a collector, so it doesn't have to exist on each of them already. """ class Sink(BaseModel): auth: Auth | None = None """ - How to authenticate, for the schemes Modelplane composes. The collector takes no credential inline: it authenticates through an extension an exporter names, so setting this composes that extension and wires the reference. - A scheme that isn't here is still reachable. Define the extension yourself under spec.extensions and name it from this sink's config, which is what Modelplane does on your behalf. + Authentication Modelplane sets up for this sink, using a credential from `secretRef`. For another scheme, define an authenticator under spec.extensions and reference it from `config`. + Set here, it wins: Modelplane applies it over an `auth` block in `config`, so a sink's credential cannot be quietly unpicked. """ config: dict[str, Any] | None = None """ - Anything else that exporter takes, passed through unread: TLS, retry, queueing, compression, headers. - Modelplane does not model an exporter's configuration, because the schema is OpenTelemetry's and versioned separately. Typing it would mean a Modelplane release for each setting the collector gains, and would drop the ones this has never heard of. What is typed above is what belongs to Modelplane: which sinks exist, what each is called, where it writes, and which Secret it reads. + The rest of the exporter's configuration, passed through as written: TLS, retries, queueing, compression, headers. + Modelplane sets one default, for the exporters that flatten a series into labels: prometheus and prometheus_remote_write get resource_to_telemetry_conversion, or they would receive every series stripped of the cluster, deployment, engine and role it belongs to. Setting it here overrides that. """ endpoint: constr(max_length=2048) | None = None """ - Where this sink writes. Typed rather than left to the configuration below because it is the setting every destination has to get right, and the one worth catching here rather than in a collector that won't start. - Optional, because not every exporter addresses its destination this way: Kafka takes brokers, the file exporter a path, and the debug exporter nothing at all. Those go in the configuration below, under the names that exporter gives them. + Where this sink writes. Leave it unset for an exporter that doesn't take an endpoint, such as kafka or debug, and configure it in `config` instead. + Set here, it wins: Modelplane applies it over `config`, so a sink cannot be quietly redirected by the configuration passed through beside it. """ name: constr(pattern=r'^[a-z0-9]([-a-z0-9]*[a-z0-9])?$', max_length=63) """ @@ -78,12 +78,11 @@ class Sink(BaseModel): """ secretRef: SecretRef | None = None """ - A Secret holding this sink's credential. Its keys reach the collector as files under /etc/modelplane/telemetry//, and as environment variables, for configuration above referring to ${env:TOKEN}. - Per sink rather than per destination, so two sinks with different credentials don't have to share one Secret and tell their keys apart by prefix. The files are per sink; the environment variables are not, so two Secrets sharing a key name still collide there and the file is the one to read. + A Secret holding this sink's credentials. Modelplane mounts each key as a file under /etc/modelplane/telemetry// and sets it as an environment variable for ${env:KEY} references in `config`. All sinks share one environment, so where two Secrets have the same key, refer to the file. """ type: constr(max_length=63) """ - The collector exporter to send with, by the name OpenTelemetry gives it: otlphttp, otlp, prometheusremotewrite, kafka, and every other one the collector provides. + The collector exporter to send with, by the name OpenTelemetry gives it: otlphttp, otlp, prometheus_remote_write, kafka, and every other one the collector provides. Not an enum, because enumerating them here would mean a Modelplane release for each exporter the collector gains, and the collector already refuses to start on a name it doesn't have. """ @@ -95,7 +94,8 @@ class Spec(BaseModel): """ extensions: dict[str, Any] | None = None """ - The collector's extensions block, passed through unread, for the authenticator a sink references. Bearer token, basic auth, OIDC and SigV4 all work, because none of them is modelled here. + Collector extensions, passed through as written. Use it to define an authenticator that a sink's `config` references. + An exporter needs a client authenticator - basicauth, oauth2client, sigv4auth, headers_setter. The oidc extension authenticates callers of a receiver, so it is not one of these. """ sinks: list[Sink] = Field(..., max_length=16, min_length=1) """ @@ -134,8 +134,8 @@ class TelemetryDestination(BaseModel): """ spec: Spec """ - Where the fleet's telemetry goes. Modelplane composes no collectors until a TelemetryDestination exists: neither tier stores anything, so collecting with nowhere to export would spend GPU-cluster memory on samples nobody reads. Creating one turns collection on everywhere at once. - There is no per-deployment opt-out. A ModelDeployment's author owns neither the destination nor its bill. + Where the fleet's metrics go. Modelplane runs no collectors until a TelemetryDestination exists, and creating one turns on collection on every inference cluster. + Several can exist. Their sinks are concatenated into one collector configuration, so adding a backend is a new object rather than an edit to one somebody else owns. A sink's name is the collector's name for its exporter, so it has to be unique across destinations. """ status: Status | None = None From 96b22dc7ed883fc1c9c66c9efcee78847eb4e245 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 19:38:56 -0700 Subject: [PATCH 37/42] Map only the SGLang metrics SGLang actually publishes Checked the built-ins against a running SGLang v0.4.9.post2 on a GPU cluster. Three of the eight named metrics it does not emit, at that version or at v0.4.6 or v0.5.0, so those mappings could never match and the fleet reported nothing for them with nothing to say why. sglang:queue_time_seconds does not exist. The nearest thing, sglang:avg_request_queue_latency, is a gauge of the mean over the last batch rather than a per-request histogram, so folding it into modelplane_request_queue_seconds beside vLLM's would make a quantile over the fleet meaningless. SGLang publishes no retraction counters at all, so sglang:num_retracted_requests_total and sglang:num_retracted_input_tokens_total named nothing - preemption is a vLLM concept here. modelplane_tokens_recomputed_total had no other source, so it goes with them. The two token histograms are real but need --collect-tokens-histogram alongside --enable-metrics: without it the engine publishes only the _total counters, and modelplane_request_input_tokens and modelplane_request_output_tokens stay empty for SGLang. The guide said to pass --enable-metrics and nothing else. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- docs/content/platform/telemetry.md | 9 +++++++- .../function/stacks/metrics.py | 11 +++++++--- .../tests/test_collector.py | 21 +++++++++++++++++++ 3 files changed, 37 insertions(+), 4 deletions(-) diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 53ae410a8..26d654350 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -231,7 +231,14 @@ less than nothing if their buckets disagree, because a quantile over them is wro than approximate. SGLang publishes `/metrics` only when it runs with `--enable-metrics`, so add that to its -engine args. vLLM needs nothing. +engine args. Add `--collect-tokens-histogram` too, or it publishes prompt and generation +tokens as plain counters and `modelplane_request_input_tokens` and +`modelplane_request_output_tokens` stay empty for that engine. vLLM needs nothing. + +SGLang publishes no queue time per request and no preemption counters, so +`modelplane_request_queue_seconds` and `modelplane_requests_preempted_total` carry vLLM +only. Its `sglang:avg_request_queue_latency` is a gauge of the mean over the last batch, +which is a different measurement, so Modelplane doesn't fold it in. ## Why engine latency and gateway latency differ diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index f856cced1..39dcd6922 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -57,13 +57,18 @@ "vllm:prefix_cache_queries_total": "modelplane_prefix_cache_lookups_total", } +# SGLang publishes none of the queue-latency or preemption measurements vLLM +# does, so there is nothing here to rename onto those names. +# sglang:avg_request_queue_latency is the nearest thing to a queue time and is +# not the same measurement - a gauge holding the mean over the last batch, +# where modelplane_request_queue_seconds is a per-request histogram - and one +# name holding both would make a quantile over the fleet meaningless. +# The two token histograms need --collect-tokens-histogram as well as +# --enable-metrics; the engine publishes only the _total counters without it. _SGLANG = { - "sglang:queue_time_seconds": "modelplane_request_queue_seconds", "sglang:num_running_reqs": "modelplane_requests_running", "sglang:num_queue_reqs": "modelplane_requests_waiting", "sglang:token_usage": "modelplane_kv_cache_utilization_ratio", - "sglang:num_retracted_requests_total": "modelplane_requests_preempted_total", - "sglang:num_retracted_input_tokens_total": "modelplane_tokens_recomputed_total", "sglang:prompt_tokens_histogram": "modelplane_request_input_tokens", "sglang:generation_tokens_histogram": "modelplane_request_output_tokens", } diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 98398878f..4f3f1dec5 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -382,6 +382,27 @@ def test_a_value_rewrite_never_lands_in_the_metric_context(self) -> None: self.assertNotIn("set(name,", st) self.assertNotIn("set(value_double,", st) + def test_sglang_carries_no_queue_time_or_preemption(self) -> None: + """SGLang publishes neither, so there is nothing to rename onto them. + + Checked against a running SGLang v0.4.9.post2: it has no per-request + queue-time metric and no retraction counters at all. The nearest + thing, sglang:avg_request_queue_latency, is a gauge of the mean over + the last batch - a different measurement from vLLM's per-request + histogram, and one name holding both makes a fleet quantile + meaningless. + """ + sglang = { + m.from_: m.to + for mapping in stacks.BUILTIN_MAPPINGS + for m in mapping.spec.metrics + if m.from_.startswith("sglang:") + } + self.assertTrue(sglang, "the SGLang built-in went missing") + self.assertNotIn("modelplane_request_queue_seconds", sglang.values()) + self.assertNotIn("modelplane_requests_preempted_total", sglang.values()) + self.assertFalse([k for k in sglang if "retracted" in k or "queue_time" in k]) + def test_sglang_latency_histograms_are_not_renamed(self) -> None: """Their buckets resolve to 100ms where vLLM's resolve to 1ms.""" joined = " ".join(_metric_statements()) From 0fcb57d70aa47317e7d314c7bd9e9b1b87d2a549 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Thu, 1 Oct 2026 22:47:53 -0700 Subject: [PATCH 38/42] Skip a telemetry object the current schema rejects A CRD validates on write, not on what it already stored, so an object written under an older schema still comes back whole on read. The function parsed every MetricMapping and TelemetryDestination straight into its Pydantic model, so one stored before a field was required raised a ValidationError - and that fails the pipeline step, which fails the whole ServingStack. The fleet stops placing replicas because a telemetry object is out of date. Hit on a live cluster: a MetricMapping written before acrossReplicas became required took the serving stack to ReconcileError, and nothing about the error named telemetry as the cause. An object that will not parse is now dropped with a warning naming it, and the rest of the fleet's telemetry carries on - the same reasoning that keeps the collector out of the stack's readiness. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- .../compose-serving-stack/function/fn.py | 42 +++++++++++++++---- .../compose-serving-stack/tests/test_fn.py | 34 +++++++++++++++ 2 files changed, 68 insertions(+), 8 deletions(-) diff --git a/functions/compose-serving-stack/function/fn.py b/functions/compose-serving-stack/function/fn.py index 9303c0dd5..e58cf0b70 100644 --- a/functions/compose-serving-stack/function/fn.py +++ b/functions/compose-serving-stack/function/fn.py @@ -34,9 +34,10 @@ ahead of the Envoy Gateway release. """ -from typing import Any +from typing import Any, TypeVar import grpc +import pydantic from crossplane.function import logging, request, resource, response from crossplane.function.proto.v1 import run_function_pb2 as fnv1 from crossplane.function.proto.v1 import run_function_pb2_grpc as grpcv1 @@ -54,6 +55,9 @@ from function import collector, gateway, stacks +# The Pydantic model one required resource is parsed into. +_T = TypeVar("_T", bound=pydantic.BaseModel) + # Label key every rendered Release and Object carries, valued with its # composed-resource key, so Usage resourceSelectors can name any # component (or one doc of a bundle) mechanically. @@ -801,10 +805,7 @@ def compose_collector(self) -> None: # restarts on a change to its rendered config, so an unstable order # would redeploy it on alternate reconciles. destinations = sorted( - ( - tdv1alpha1.TelemetryDestination.model_validate(d) - for d in request.get_required_resources(self.req, "destinations") - ), + self._parse("TelemetryDestination", tdv1alpha1.TelemetryDestination, "destinations"), key=lambda d: _name(d.metadata), ) if not destinations: @@ -843,9 +844,7 @@ def compose_collector(self) -> None: # Modelplane's own mappings first, then the operator's, which add to # them rather than replacing them. mappings = list(stacks.BUILTIN_MAPPINGS) - mappings += [ - mmv1alpha1.MetricMapping.model_validate(m) for m in request.get_required_resources(self.req, "mappings") - ] + mappings += self._parse("MetricMapping", mmv1alpha1.MetricMapping, "mappings") pc_observed = self.provider_configs_observed() pc = _pc_name(self.xr) @@ -890,6 +889,33 @@ def compose_collector(self) -> None: ) self.rsp.desired.resources[key].ready = fnv1.READY_TRUE + def _parse(self, kind: str, model: type[_T], key: str) -> list[_T]: + """Parse the required resources under `key`, skipping what won't. + + An object the API server stored under an older schema still comes back + on read - a CRD's validation runs on write, not on what is already + there - so one MetricMapping written before a field was required is + enough to raise here. Raising fails the whole pipeline step, which + takes down the serving stack: the fleet stops placing replicas because + a telemetry object is out of date. + + So a parse failure drops that object and says so, the same reasoning + that keeps the collector out of the stack's readiness. The rest of the + fleet's telemetry carries on without it. + """ + out: list[_T] = [] + for obj in request.get_required_resources(self.req, key): + try: + out.append(model.model_validate(obj)) + except pydantic.ValidationError as err: + name = (obj.get("metadata") or {}).get("name", "") + response.warning( + self.rsp, + f"Ignoring {kind} {name}: it does not match the current schema " + f"({err.error_count()} problems), so nothing it asks for is collected.", + ) + return out + def merge_destinations( self, destinations: list[tdv1alpha1.TelemetryDestination] ) -> tuple[list[tdv1alpha1.Sink], dict[str, Any]]: diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index 527f7f5f7..72814a431 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -1660,3 +1660,37 @@ async def test_two_destinations_cannot_name_one_exporter(self) -> None: self.assertEqual(list(config["exporters"]), ["otlphttp/primary"]) self.assertEqual(config["exporters"]["otlphttp/primary"]["endpoint"], "https://a.example") self.assertTrue([r for r in got.results if "zeta" in r.message]) + + async def test_a_stale_mapping_does_not_break_the_stack(self) -> None: + """A CRD validates on write, not on what it already stored. + + A MetricMapping written before acrossReplicas was required still + comes back on read without it. Parsing it raises, and raising fails + the whole pipeline step - so the serving stack composes nothing and + the fleet stops placing replicas, because one telemetry object is out + of date. Seen on a real cluster. + """ + req = self._with_destination(_request("GKE", "Standard", observed=_observed_pcs())) + req.required_resources["mappings"].items.append( + fnv1.Resource( + resource=resource.dict_to_struct( + { + "apiVersion": "modelplane.ai/v1alpha1", + "kind": "MetricMapping", + "metadata": {"name": "stale"}, + # No acrossReplicas: the schema requires it now. + "spec": {"metrics": [{"from": "old_engine_waiting", "to": "modelplane_requests_waiting"}]}, + } + ) + ) + ) + got = await self.runner.RunFunction(req, None) + # The stack still composes, and says what it dropped. + self.assertIn("collector", got.desired.resources) + self.assertTrue([r for r in got.results if "stale" in r.message]) + config = yaml.safe_load( + resource.struct_to_dict(got.desired.resources["collector-config"].resource)["spec"]["forProvider"][ + "manifest" + ]["data"]["collector.yaml"] + ) + self.assertNotIn("old_engine_waiting", yaml.safe_dump(config)) From 2b7ce70c136c35eb4216e7af9a94ccdce9a0b754 Mon Sep 17 00:00:00 2001 From: Rae Sharp Date: Fri, 2 Oct 2026 11:18:50 -0400 Subject: [PATCH 39/42] Leans down guide, clarifying some sections, adds examples Signed-off-by: Rae Sharp --- docs/content/platform/telemetry.md | 320 ++++++++++-------- .../config/vocabularies/Modelplane/accept.txt | 2 + 2 files changed, 188 insertions(+), 134 deletions(-) diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 1a5895475..b96aea36b 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -2,42 +2,60 @@ title: Monitor the Fleet weight: 37 aliases: -- /guides/collecting-engine-metrics/ + - /guides/telemetry/ description: Collect normalized metrics across the fleet and send them anywhere that speaks OTLP. --- -Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from -every component Modelplane installs, which is more than your engines. It renames each -component's series to a single `modelplane_*` vocabulary and pushes to a collector on your -control plane. That collector is your fleet's -single egress point, and it sends to any backend that speaks OTLP. +Modelplane runs an OpenTelemetry collector on every inference cluster. It +collects from every component Modelplane installs. This includes the inference +server engine, inference gateway and Envoy proxy, router, and the GPU exporter +your cloud provides. It renames each component's series to a single +`modelplane_*` vocabulary and pushes to a collector on your control plane. That +collector is your fleet's single egress point and sends data to any +collector exporter backend. + +Modelplane allows you to write one destination for your metrics. You don't need +to manage per-deployment configurations or update your configuration when a +deployment changes. The OpenTelemetry collector can find pods itself and leader/worker +splits or a prefill/decode pairs get collected the same as a single pod. + +## Telemetry workflow -Modelplane has no API for this: nothing to write, and nothing to keep in sync as your -deployments change. +Every series carries `cluster`, `job`, and `instance` labels of the target +resource. A series about a deployment also carries `deployment`, `replica`, +`namespace`, `engine`, and `role` labels. -## What you get +Each replica publishes its own series, so combine them in your query. Use the +aggregation that the metric's `acrossReplicas` field in its `MetricMapping` names: -Every series carries `cluster`, and `job` and `instance` naming the target it was scraped -from. A series about a deployment also carries `deployment`, `replica`, `namespace`, -`engine`, and `role`. + - `sum by (deployment)`, for anything counted, such as requests, tokens, or queue depth. + - `avg by (deployment)`, for a ratio. + - `max by (deployment)`, for a saturation figure an alert fires on. -Each replica publishes its own series. Combine them in the query, the way the metric's -`acrossReplicas` says: `sum by (deployment)` for anything counted, `avg by (deployment)` -for a ratio, `max by (deployment)` for a saturation figure an alert fires on. The -collector doesn't add them up for you, because a scrape of one replica is one batch, and -adding readings taken at different moments is not the traffic that happened. +To combine: -The replica is an index rather than a pod, so it is bounded by the replica count and -survives a restart and a rolling update. Group by it, not by `instance`. +```promql +sum by (deployment) (rate(modelplane_frontend_request_duration_seconds_count[5m])) +``` + +The replica is an index rather than a pod, so it's bounded by the replica count +and survives a restart and a rolling update. Group by `replica`, not by `instance`. + +For example: +```promql +# One line per replica, stable across rolling updates +max by (deployment, replica) (modelplane_kv_cache_utilization_ratio) +``` +The `instance` label is the pod's address. Without the `instance` label two pods writing to the same +series would collide and a deployment with several pods per +replica or two gateway pods couldn't distinguish between the pods. -`instance` is the pod's address, and it is there because two pods writing one series is one -series with one of them lost - a deployment running several pods per replica, or two -gateway pods, have nothing else to tell them apart. It does turn over on a rolling update, -so a query that groups by it grows a series every time you deploy. Aggregate it away. +The `instance` label changes on a rolling update so queries that group by this +label gain a new series every time you deploy. Group by `replica` instead. -Some of what you can read: +Some examples of the available metrics: | Metric | Means | | --- | --- | @@ -53,24 +71,25 @@ Some of what you can read: | `modelplane_energy_joules_total` | Energy drawn since the driver last reloaded | Latency appears twice on purpose. The `frontend_` series are what your caller experienced, -measured at the gateway. The engine's own series are what the engine spent. When the -frontend number is slow and the engine number isn't, the problem is routing, queueing, or -the network rather than the model. +measured at the gateway. The engine's own series are what the engine spent. For +example, if the frontend metric is slow and the engine isn't, you can +troubleshoot routing, queueing, or networking issues instead of the model. + Saturation gauges come as a pair. The average is what you plan capacity against; the `_max` is what you alert on, because three replicas at 0.3 and one at 0.99 average to something comfortable while the fourth evicts and recomputes. A high `_max` beside `modelplane_requests_preempted_total` climbing is one replica thrashing. - -No series names a pod. Replicas are interchangeable, so they're summed before the metrics -leave the cluster; a rolling update would otherwise leave a dead series behind for every pod -it replaced. - +For example: -## Sending it somewhere +```promql +max by (deployment) (modelplane_kv_cache_utilization_ratio) > 0.95 +``` -Create a `TelemetryDestination` naming whatever you already run: +## Send telemetry to a destination + +Create a `TelemetryDestination` for your OpenTelemetry-compatible endpoint: ```yaml apiVersion: modelplane.ai/v1alpha1 @@ -86,9 +105,17 @@ spec: `type` names a collector exporter, by the name OpenTelemetry gives it. -Put the credential in a Secret, name it with the sink's `secretRef`, and say which key holds -the token. Modelplane composes the authenticator and wires it up, and the token never -appears in `kubectl get -o yaml`: +To authenticate with a bearer token, store the token in a Secret in Modelplane's +namespace: + +```shell +kubectl create secret generic telemetry-credentials \ + --namespace \ + --from-literal=token= +``` + +Reference the Secret from the sink with `secretRef`, and set `auth.bearerTokenKey` to +the key that holds the token: ```yaml spec: @@ -102,10 +129,11 @@ spec: bearerTokenKey: token ``` -It reads the token from a file rather than the environment, so rotating it doesn't need the -collector restarted. +Modelplane configures the collector to send the token with every export. -If you run Prometheus, export to that instead and query the fleet there: +The collector reads the token from a file rather than the environment. + +If you run Prometheus, export to your Prometheus endpoint instead and query the fleet there: ```yaml spec: @@ -115,8 +143,8 @@ spec: endpoint: https://prom.example.internal/api/v1/write ``` -Name more than one sink and every one gets the whole stream. Each carries its own -credential, so a vendor and your own Prometheus don't have to share a Secret: +If you create more than one sink, all get the entire stream. Each sink +carries it's own credential so you don't have to share a Secret. ```yaml spec: @@ -133,7 +161,7 @@ spec: endpoint: https://prom.example.internal/api/v1/write ``` -That is two copies of the fleet's metrics, billed twice. +That's two copies of the fleet's metrics, billed twice. Anything else the exporter takes goes under `config`, passed through as you wrote it: @@ -148,15 +176,37 @@ Anything else the exporter takes goes under `config`, passed through as you wrot tls: ca_file: /etc/ssl/certs/internal.pem ``` +Modelplane doesn't define a schema for an exporter's settings so anything under +the `config` is passed to the collector exactly as written. TLS, retries, +querying, compression and headers all work and any new settings in the collector +are respected and the sink keeps working. + +To use an authentication scheme Modelplane doesn't compose, define the extension +yourself under `spec.extensions`. Then reference it by its key from the sink's +`config.auth.authenticator`. The `auth` block does the same wiring for you when +you use a bearer token. + +```yaml +spec: + sinks: + - name: vendor + type: otlphttp + endpoint: https://otel.vendor.example + secretRef: + name: vendor-oauth + config: + auth: + authenticator: oauth2client/vendor + extensions: + oauth2client/vendor: + client_id: modelplane + client_secret: ${env:CLIENT_SECRET} + token_url: https://issuer.example/oauth2/token +``` -Modelplane doesn't model what an exporter is, so its TLS, retry and queue settings all work, -and a sink keeps working when the collector gains a setting Modelplane has never heard of. -An authentication scheme Modelplane doesn't compose works the same way: define the extension -under `spec.extensions` and name it from the sink's `config`, which is what the `auth` block -above does for you. +Modelplane doesn't run any collectors until you create a +`TelemetryDestination`. -Until you create one, Modelplane composes no collectors: nothing here stores anything, so -collecting with nowhere to send it would spend GPU-cluster memory on samples nobody reads. Creating a destination turns collection on everywhere at once, and there's no per-deployment opt-out. @@ -181,7 +231,39 @@ export to Prometheus and write recording rules there. ## Engines Modelplane renames vLLM's and SGLang's own metrics for you, so neither needs a mapping. -SGLang needs one flag to publish them at all, below. +SGLang requires the `--enable-metrics` to publish them at all. + + +For example: + +```yaml +apiVersion: modelplane.ai/v1alpha1 +kind: ModelDeployment +metadata: + name: my-sglang-model + namespace: ml-team +spec: + template: + spec: + engines: + - name: engine + members: + - role: Standalone + template: + spec: + containers: + - name: engine + image: lmsysorg/sglang:v0.5.10.post1-runtime + command: + - /bin/sh + - -c + - >- + exec python3 -m sglang.launch_server + --model-path + --host 0.0.0.0 + --port 8000 + --enable-metrics +``` Any other OpenAI-compatible engine reports its top-line numbers with no configuration. The gateway measures those, not the engine, so `modelplane_frontend_*` works for an engine @@ -207,7 +289,7 @@ Modelplane renders every mapping into every cluster's collector, so you write on `from` is the name your engine emits and `to` is what Modelplane calls it. Say `fromUnit` whenever the engine measures in something other than the unit the name -claims, and Modelplane converts to the base one. Skipping it is the expensive mistake here: +claims, and Modelplane converts to the base one. Skipping this is the expensive mistake here: a series named `_seconds` that holds milliseconds reads a thousand times fast, and nothing downstream can tell. @@ -215,8 +297,58 @@ Rename only where the measurements agree. Two engines' histograms under one name less than nothing if their buckets disagree, because a quantile over them is wrong rather than approximate. -SGLang publishes `/metrics` only when it runs with `--enable-metrics`, so add that to its -engine args. vLLM needs nothing. +### Examples + +```yaml +apiVersion: modelplane.ai/v1alpha1 +kind: MetricMapping +metadata: + name: my-engine +spec: + metrics: + # A plain rename. + - from: my_engine_queued_requests + to: modelplane_requests_waiting + acrossReplicas: Sum + + # A unit conversion. The engine reports milliseconds; the name says seconds. + - from: my_engine_kv_transfer_ms + to: modelplane_request_kv_transfer_seconds + fromUnit: Milliseconds + acrossReplicas: Sum + + # A request count taken out of a duration histogram. The histogram + # keeps its own name; this adds a counter beside it. + - from: my_engine_request_duration_seconds + part: Count + to: modelplane_requests_total + acrossReplicas: Sum + + # Two counters folded into one name, told apart by a fixed label. + - from: my_engine_prompt_tokens_total + to: modelplane_tokens_total + acrossReplicas: Sum + labels: + - name: direction + value: input + - from: my_engine_generated_tokens_total + to: modelplane_tokens_total + acrossReplicas: Sum + labels: + - name: direction + value: output + + # A label the engine already emits, renamed and its values translated. + - from: my_engine_finished_requests_total + to: modelplane_responses_total + acrossReplicas: Sum + labels: + - name: reason + from: finish_reason + values: + eos: stop + max_tokens: length +``` ## Why engine latency and gateway latency differ @@ -228,83 +360,3 @@ one engine against itself, and the `frontend_` series for anything fleet-wide. Some measurements don't translate at all. SGLang's inter-token latency isn't vLLM's time per output token, so neither is renamed onto a shared name. The gateway measures time per output token for both. - -## Migrating from a hand-written `PodMonitor` - -Modelplane used to have you write a `PodMonitor` and reach an in-cluster Prometheus over a -`port-forward`. Both are gone. Three steps to move across, and two of them fail quietly if -you skip them. - -**Keep your Prometheus, and point a destination at it.** Collection becomes a push, so your -store stops scraping and starts receiving. Same Prometheus, same retention, same Grafana: - -```yaml -spec: - sinks: - - name: prometheus - type: prometheusremotewrite - config: - endpoint: http://prometheus.monitoring.svc:9090/api/v1/write -``` - -**Delete the monitors you wrote.** A `PodMonitor` or `ScrapeConfig` pointed at your engines -keeps working against your own Prometheus, so nothing appears to break and you collect -everything twice, under `vllm:*` and under `modelplane_*`, paying for both. One written -against Modelplane's Prometheus stops being read by anything, because the operator goes with -the stack. - -**Rewrite your dashboard queries.** Names change, and so do the labels: group by -`deployment` rather than `model_name`, and every series carries `cluster`, `replica`, and -the `instance` it was scraped from. A panel that showed one engine now shows one pod, so -wrap it in `sum by (deployment)` or the aggregation that metric's `acrossReplicas` names. - -| Was | Is | -| --- | --- | -| `vllm:time_to_first_token_seconds` | `modelplane_request_ttft_seconds` | -| `vllm:e2e_request_latency_seconds` | `modelplane_request_duration_seconds` | -| `vllm:request_queue_time_seconds` | `modelplane_request_queue_seconds` | -| `vllm:request_prefill_time_seconds` | `modelplane_request_prefill_seconds` | -| `vllm:request_decode_time_seconds` | `modelplane_request_decode_seconds` | -| `vllm:num_requests_running` | `modelplane_requests_running` | -| `vllm:num_requests_waiting` | `modelplane_requests_waiting` | -| `vllm:kv_cache_usage_perc` | `modelplane_kv_cache_utilization_ratio` | -| `vllm:num_preemptions_total` | `modelplane_requests_preempted_total` | -| `vllm:prefix_cache_hits_total` | `modelplane_prefix_cache_hits_total` | -| `DCGM_FI_DEV_FB_USED` | `modelplane_gpu_memory_used_bytes` | -| `DCGM_FI_DEV_GPU_TEMP` | `modelplane_gpu_temperature_celsius` | -| `DCGM_FI_DEV_POWER_USAGE` | `modelplane_gpu_power_watts` | -| `DCGM_FI_PROF_PIPE_TENSOR_ACTIVE` | `modelplane_gpu_tensor_active_ratio` | -| `envoy_cluster_upstream_rq_time` | `modelplane_frontend_request_duration_seconds` | - -Some have no replacement. A rename carries one metric to one name, so the counters that -would fold several series under one label - tokens by direction, responses by reason, -requests by status - aren't part of this surface yet. Keep reading those from your engine -and your gateway directly. For prompt and output size, `modelplane_request_input_tokens` -and `modelplane_request_output_tokens` carry the same measurement as histograms. - -`vllm:inter_token_latency_seconds` isn't renamed, because -SGLang publishes a metric of the same name measuring something else; use -`modelplane_frontend_tpot_seconds`, which the gateway measures the same way for every -engine. `DCGM_FI_DEV_GPU_UTIL` isn't renamed either, because it only tells you the card -wasn't idle; use `modelplane_gpu_compute_active_ratio` and -`modelplane_gpu_tensor_active_ratio`. - -You can also defer the rewrite. A recording rule rebuilds an old name from a new one, so a -dashboard keeps working untouched while you migrate it: - -```yaml -- record: vllm:time_to_first_token_seconds_bucket - expr: label_replace(modelplane_request_ttft_seconds_bucket, - "model_name", "$1", "deployment", "(.*)") -``` - -Load that into the Prometheus you already run and nothing on the dashboard changes. It's one -rule evaluation per metric over series your backend already holds, so it costs far less than -collecting everything twice. Write one per name in the table above, and delete them once the -panels use the new names. - -A series no statement renames doesn't leave the cluster. If a panel needs an engine's -own name, write a `MetricMapping` that renames it onto the `modelplane_*` surface: a -mapping for an engine Modelplane already knows adds to the built-in renames rather -than replacing them. - diff --git a/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt b/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt index 9486fc527..a5fc7bb87 100644 --- a/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt +++ b/docs/utils/vale/styles/config/vocabularies/Modelplane/accept.txt @@ -149,6 +149,8 @@ Baseten WekaIO NetApp minikube +Monitor +Aggregate # AI assistants and agent tooling MCP From 56c804054017661816d759cfd017f52c0c0bd41a Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Fri, 2 Oct 2026 08:28:22 -0700 Subject: [PATCH 40/42] Correct what the telemetry guide claims after the edit pass Four claims in the guide aren't true of the code, two of them long-standing and two lost in the edit. There is no collector on the control plane. Only compose-serving-stack composes one, onto each inference cluster, and it exports to the destination's sinks directly - so there is no "single egress point" tier, and the sinks are whatever the collector has an exporter for rather than OTLP alone. Nothing emits a `_max` series. A saturation gauge is one series per replica and the choice between an average and a maximum belongs to the query, which is the whole point of `acrossReplicas`. The example beside the paragraph already said so. SGLang needs `--collect-tokens-histogram` as well as `--enable-metrics`, or the two token histograms stay empty, and it publishes neither a per-request queue time nor preemption counters. Both were checked against a running engine. The credential Secret goes in modelplane-system on the control plane and Modelplane copies it to every cluster running a collector, which is the question a reader has at that point. Also restores the collecting-engine-metrics alias, so the old URL keeps resolving. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- docs/content/platform/telemetry.md | 59 ++++++++++++++++-------------- 1 file changed, 31 insertions(+), 28 deletions(-) diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 185e9837a..e4392e805 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -2,9 +2,9 @@ title: Monitor the Fleet weight: 37 aliases: - +- /guides/collecting-engine-metrics/ - /guides/telemetry/ -description: Collect normalized metrics across the fleet and send them anywhere that speaks OTLP. +description: Collect normalized metrics across the fleet and send them to any backend the collector can export to. --- @@ -12,14 +12,13 @@ Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from every component Modelplane installs. This includes the inference server engine, inference gateway and Envoy proxy, router, and the GPU exporter your cloud provides. It renames each component's series to a single -`modelplane_*` vocabulary and pushes to a collector on your control plane. That -collector is your fleet's single egress point and sends data to any -collector exporter backend. +`modelplane_*` vocabulary and exports them to wherever you say - any backend the +collector has an exporter for, not only OTLP. Modelplane allows you to write one destination for your metrics. You don't need to manage per-deployment configurations or update your configuration when a -deployment changes. The OpenTelemetry collector can find pods itself and leader/worker -splits or a prefill/decode pairs get collected the same as a single pod. +deployment changes. The collector finds pods itself, so a leader/worker split or a +prefill/decode pair is collected the same as a single pod. ## Telemetry workflow @@ -41,9 +40,8 @@ sum by (deployment) (rate(modelplane_frontend_request_duration_seconds_count[5m] ``` The replica is an index rather than a pod, so it's bounded by the replica count -and survives a restart and a rolling update. Group by `replica`, not by `instance`. +and survives a restart and a rolling update. Group by `replica`, not by `instance`: -For example: ```promql # One line per replica, stable across rolling updates max by (deployment, replica) (modelplane_kv_cache_utilization_ratio) @@ -76,12 +74,10 @@ example, if the frontend metric is slow and the engine isn't, you can troubleshoot routing, queueing, or networking issues instead of the model. -Saturation gauges come as a pair. The average is what you plan capacity against; the `_max` -is what you alert on, because three replicas at 0.3 and one at 0.99 average to something -comfortable while the fourth evicts and recomputes. A high `_max` beside -`modelplane_requests_preempted_total` climbing is one replica thrashing. - -For example: +Saturation gauges are per replica, so how you combine them decides what you see. Average +across a deployment to plan capacity, and take the maximum to alert: three replicas at 0.3 +and one at 0.99 average to something comfortable while the fourth evicts and recomputes. A +high maximum beside `modelplane_requests_preempted_total` climbing is one replica thrashing. To alert on it: ```promql max by (deployment) (modelplane_kv_cache_utilization_ratio) > 0.95 @@ -105,12 +101,14 @@ spec: `type` names a collector exporter, by the name OpenTelemetry gives it. -To authenticate with a bearer token, store the token in a Secret in Modelplane's -namespace: +To authenticate with a bearer token, store the token in a Secret in +`modelplane-system` on your control plane. Create it once: Modelplane copies it to +every cluster running a collector, so you don't put the credential on each GPU +cluster yourself. ```shell kubectl create secret generic telemetry-credentials \ - --namespace \ + --namespace modelplane-system \ --from-literal=token= ``` @@ -129,9 +127,8 @@ spec: bearerTokenKey: token ``` -Modelplane configures the collector to send the token with every export. - -The collector reads the token from a file rather than the environment. +Modelplane configures the collector to send the token with every export. It reads the +token from a file rather than the environment, so rotating it needs no restart. If you run Prometheus, export to your Prometheus endpoint instead and query the fleet there: @@ -143,8 +140,8 @@ spec: endpoint: https://prom.example.internal/api/v1/write ``` -If you create more than one sink, all get the entire stream. Each sink -carries it's own credential so you don't have to share a Secret. +If you create more than one sink, all get the entire stream. Each sink carries its +own credential, so a vendor and your own Prometheus don't have to share a Secret. ```yaml spec: @@ -185,7 +182,7 @@ Anything else the exporter takes goes under `config`, passed through as you wrot ``` Modelplane doesn't define a schema for an exporter's settings so anything under the `config` is passed to the collector exactly as written. TLS, retries, -querying, compression and headers all work and any new settings in the collector +queueing, compression and headers all work and any new settings in the collector are respected and the sink keeps working. To use an authentication scheme Modelplane doesn't compose, define the extension @@ -238,10 +235,11 @@ export to Prometheus and write recording rules there. ## Engines Modelplane renames vLLM's and SGLang's own metrics for you, so neither needs a mapping. -SGLang requires the `--enable-metrics` to publish them at all. - - -For example: +SGLang needs two flags: `--enable-metrics` to publish `/metrics` at all, and +`--collect-tokens-histogram` for the prompt and generation histograms behind +`modelplane_request_input_tokens` and `modelplane_request_output_tokens`. Without the +second it publishes those as plain counters and both series stay empty. vLLM needs +nothing. ```yaml apiVersion: modelplane.ai/v1alpha1 @@ -270,8 +268,13 @@ spec: --host 0.0.0.0 --port 8000 --enable-metrics + --collect-tokens-histogram ``` +SGLang publishes no queue time per request and no preemption counters, so +`modelplane_request_queue_seconds` and `modelplane_requests_preempted_total` carry vLLM +only. + Any other OpenAI-compatible engine reports its top-line numbers with no configuration. The gateway measures those, not the engine, so `modelplane_frontend_*` works for an engine Modelplane has never seen. From cc09442e96941d869ddefdd73e49591dea1ee411 Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Fri, 2 Oct 2026 08:44:19 -0700 Subject: [PATCH 41/42] Answer the rest of the guide's review comments Sweeps what was left of the second review pass over the telemetry guide. `acrossReplicas` arrived with no introduction, in a sentence that assumed the reader already knew the field. It now says what it is where it first appears. The `instance` paragraphs spent six lines justifying a label's existence; one sentence says what it is and what to group by instead. How the collector reads a credential off disk is our business rather than the reader's, so that goes, keeping only the part an operator acts on: rotating a token needs no restart. The Engines section told you vLLM and SGLang need no mapping a page before anything said what a mapping was. And "top-line numbers" wasn't a phrase anyone outside this repo uses - they're the frontend numbers, which is what the metrics are called. The TelemetryDestination schema justified `type` not being an enum by the Modelplane release a new exporter would otherwise need. That isn't true: a new exporter needs a new collector image, which needs a release anyway. The honest reason is narrower - the collector already rejects a name it doesn't have, so a list here would only be a second thing to go stale. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/telemetrydestinations/definition.yaml | 7 +++-- docs/content/platform/telemetry.md | 27 +++++++++---------- schemas/.lock.json | 2 +- .../telemetrydestination/v1alpha1.py | 2 +- 4 files changed, 17 insertions(+), 21 deletions(-) diff --git a/apis/telemetrydestinations/definition.yaml b/apis/telemetrydestinations/definition.yaml index c2a57c7c8..0d24e9a83 100644 --- a/apis/telemetrydestinations/definition.yaml +++ b/apis/telemetrydestinations/definition.yaml @@ -72,10 +72,9 @@ spec: prometheus_remote_write, kafka, and every other one the collector provides. - Not an enum, because enumerating them here would mean - a Modelplane release for each exporter the collector - gains, and the collector already refuses to start on a - name it doesn't have. + Not an enum: the collector already refuses to start + on a name it doesn't have, so repeating the list here + would only add a second place for it to go stale. endpoint: type: string maxLength: 2048 diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index e4392e805..9d32a8557 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -26,8 +26,9 @@ Every series carries `cluster`, `job`, and `instance` labels of the target resource. A series about a deployment also carries `deployment`, `replica`, `namespace`, `engine`, and `role` labels. -Each replica publishes its own series, so combine them in your query. Use the -aggregation that the metric's `acrossReplicas` field in its `MetricMapping` names: +Each replica publishes its own series, so combine them in your query. Every metric +records which combination is the right one for it, as `acrossReplicas` in its +`MetricMapping`, so you don't have to work it out per metric: - `sum by (deployment)`, for anything counted, such as requests, tokens, or queue depth. - `avg by (deployment)`, for a ratio. @@ -46,12 +47,8 @@ and survives a restart and a rolling update. Group by `replica`, not by `instanc # One line per replica, stable across rolling updates max by (deployment, replica) (modelplane_kv_cache_utilization_ratio) ``` -The `instance` label is the pod's address. Without the `instance` label two pods writing to the same -series would collide and a deployment with several pods per -replica or two gateway pods couldn't distinguish between the pods. - -The `instance` label changes on a rolling update so queries that group by this -label gain a new series every time you deploy. Group by `replica` instead. +`instance` is the pod's address, which keeps two pods of the same replica apart. It +turns over on every rolling update, so group by `replica` rather than by `instance`. Some examples of the available metrics: @@ -127,8 +124,8 @@ spec: bearerTokenKey: token ``` -Modelplane configures the collector to send the token with every export. It reads the -token from a file rather than the environment, so rotating it needs no restart. +Modelplane configures the collector to send the token with every export. Rotating the +token needs no restart. If you run Prometheus, export to your Prometheus endpoint instead and query the fleet there: @@ -234,8 +231,8 @@ export to Prometheus and write recording rules there. ## Engines -Modelplane renames vLLM's and SGLang's own metrics for you, so neither needs a mapping. -SGLang needs two flags: `--enable-metrics` to publish `/metrics` at all, and +Modelplane already knows vLLM's and SGLang's metric names and renames them for you, so +neither needs anything from you here. SGLang needs two flags: `--enable-metrics` to publish `/metrics` at all, and `--collect-tokens-histogram` for the prompt and generation histograms behind `modelplane_request_input_tokens` and `modelplane_request_output_tokens`. Without the second it publishes those as plain counters and both series stay empty. vLLM needs @@ -275,9 +272,9 @@ SGLang publishes no queue time per request and no preemption counters, so `modelplane_request_queue_seconds` and `modelplane_requests_preempted_total` carry vLLM only. -Any other OpenAI-compatible engine reports its top-line numbers with no configuration. The -gateway measures those, not the engine, so `modelplane_frontend_*` works for an engine -Modelplane has never seen. +Any other OpenAI-compatible engine reports its frontend numbers with no configuration. +The gateway measures those, not the engine, so `modelplane_frontend_*` works for an +engine Modelplane has never seen. To normalize that engine's own metrics as well, create a `MetricMapping`: diff --git a/schemas/.lock.json b/schemas/.lock.json index 7ae604486..3edfa7b5e 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "1d2cf24a31d015db1785d6f5921e560194a2985d8fe812e4c38f10972fe3c06e", + "fs://apis": "828906d23cc6da5be87e2b7692a96a5249f77c44ab14e012760807e09add5d96", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.8.1": "sha256:ca2e9e3b2e3a8b6ca44a9700d5abf7abd733cfa388d1afe9bb2bf7c847cd39ef", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.8.1": "sha256:ebb1bcd8dc9a7e60e97a324652fbd9d1b0609ddbb1107db518ce999cbc1113bd", diff --git a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py index 2365e38fe..3414851d8 100644 --- a/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/telemetrydestination/v1alpha1.py @@ -83,7 +83,7 @@ class Sink(BaseModel): type: constr(max_length=63) """ The collector exporter to send with, by the name OpenTelemetry gives it: otlphttp, otlp, prometheus_remote_write, kafka, and every other one the collector provides. - Not an enum, because enumerating them here would mean a Modelplane release for each exporter the collector gains, and the collector already refuses to start on a name it doesn't have. + Not an enum: the collector already refuses to start on a name it doesn't have, so repeating the list here would only add a second place for it to go stale. """ From 191b0bfaa4e79463d7dfec0fb6748169610a19fa Mon Sep 17 00:00:00 2001 From: Dennis Ramdass Date: Fri, 2 Oct 2026 11:08:33 -0700 Subject: [PATCH 42/42] Drop acrossReplicas from MetricMapping Nothing read it. It was required on every entry of every mapping, and no code anywhere consumed the value - it existed to tell whoever wrote a dashboard query whether to sum or average a metric. It fails at that. Finding out what a metric says means fetching the MetricMapping and reading the field, which is more work than deciding from the metric itself, and for everything Modelplane ships the name already answers it: a _total sums, a _ratio averages. So the field taxed every author of a mapping to serve a reader who is better off without it. Removing it now is cheap. Taking a required field out of a published API later is not, and adding it back if something ever generates dashboards or recording rules from mappings would be purely additive. The reasoning it carried stays in the guide, because it is about the collector rather than the field: Modelplane doesn't combine a deployment's replicas, since a scrape of one replica is one batch and summing readings taken at different moments is not the traffic that happened. Co-Authored-By: Claude Opus 5 Signed-off-by: Dennis Ramdass --- apis/metricmappings/definition.yaml | 28 +------------------ docs/content/platform/telemetry.md | 20 ++++--------- .../compose-metric-mapping/tests/test_fn.py | 2 +- .../function/collector.py | 3 +- .../function/stacks/metrics.py | 22 +-------------- .../tests/test_collector.py | 12 ++------ .../compose-serving-stack/tests/test_fn.py | 24 +++++++++++----- schemas/.lock.json | 2 +- .../ai/modelplane/metricmapping/v1alpha1.py | 7 ----- 9 files changed, 30 insertions(+), 90 deletions(-) diff --git a/apis/metricmappings/definition.yaml b/apis/metricmappings/definition.yaml index 89f1b881a..63b512801 100644 --- a/apis/metricmappings/definition.yaml +++ b/apis/metricmappings/definition.yaml @@ -54,7 +54,7 @@ spec: buckets is wrong. items: type: object - required: [from, to, acrossReplicas] + required: [from, to] properties: from: type: string @@ -80,32 +80,6 @@ spec: Name it in base units - seconds, bytes, joules, a ratio from nought to one - because that is what `fromUnit` converts to. - acrossReplicas: - type: string - enum: [Sum, Mean, Max] - description: >- - How this metric combines over a deployment's replicas. - - Every pod publishes its own series, told apart by the - replica it belongs to and the instance it was scraped - from, and a query over a deployment combines them. This says which combination is the right - one: Sum for anything counted - requests, tokens, - joules, a queue's depth. Mean for a ratio, where - summing reads two replicas at half capacity as one at - full. Max for a saturation figure an alert fires on, - where a mean hides the replica in trouble. - - Modelplane does not combine them in the collector. A - scrape of one replica is one batch, so a collector that - added them up would be adding readings taken at - different moments, and two readings of one cumulative - counter sum to twice the traffic that happened. The - backend holds every replica's series and combines them - at query time, where the arithmetic is right. - - Required, with no default, because the wrong - combination is silent: a deployment reports a number - that looks entirely plausible. part: type: string enum: [Count, Sum] diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 9d32a8557..cc5d0d560 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -26,9 +26,8 @@ Every series carries `cluster`, `job`, and `instance` labels of the target resource. A series about a deployment also carries `deployment`, `replica`, `namespace`, `engine`, and `role` labels. -Each replica publishes its own series, so combine them in your query. Every metric -records which combination is the right one for it, as `acrossReplicas` in its -`MetricMapping`, so you don't have to work it out per metric: +Each replica publishes its own series, so combine them in your query. Which +combination is right follows from what the metric measures: - `sum by (deployment)`, for anything counted, such as requests, tokens, or queue depth. - `avg by (deployment)`, for a ratio. @@ -287,19 +286,18 @@ spec: metrics: - from: my_engine_queued_requests to: modelplane_requests_waiting - acrossReplicas: Sum - from: my_engine_kv_transfer_ms fromUnit: Milliseconds to: modelplane_request_kv_transfer_seconds - acrossReplicas: Mean ``` Modelplane renders every mapping into every cluster's collector, so you write one once. `from` is the name your engine emits and `to` is what Modelplane calls it. -`acrossReplicas` says how a query should combine the metric over a deployment's replicas, -since every pod publishes its own series. Use `Sum` for anything counted and `Mean` for a -ratio, where adding two replicas at half capacity would read as one at full. +Modelplane leaves the combining to your backend. A scrape of one replica is one batch, so +a collector that added them up would be summing readings taken at different moments, and +two readings of one cumulative counter come to twice the traffic that happened. Your +backend holds every replica's series and combines them at query time. Say `fromUnit` whenever the engine measures in something other than the unit the name claims, and Modelplane converts to the base one. Skipping this is the expensive mistake here: @@ -322,31 +320,26 @@ spec: # A plain rename. - from: my_engine_queued_requests to: modelplane_requests_waiting - acrossReplicas: Sum # A unit conversion. The engine reports milliseconds; the name says seconds. - from: my_engine_kv_transfer_ms to: modelplane_request_kv_transfer_seconds fromUnit: Milliseconds - acrossReplicas: Sum # A request count taken out of a duration histogram. The histogram # keeps its own name; this adds a counter beside it. - from: my_engine_request_duration_seconds part: Count to: modelplane_requests_total - acrossReplicas: Sum # Two counters folded into one name, told apart by a fixed label. - from: my_engine_prompt_tokens_total to: modelplane_tokens_total - acrossReplicas: Sum labels: - name: direction value: input - from: my_engine_generated_tokens_total to: modelplane_tokens_total - acrossReplicas: Sum labels: - name: direction value: output @@ -354,7 +347,6 @@ spec: # A label the engine already emits, renamed and its values translated. - from: my_engine_finished_requests_total to: modelplane_responses_total - acrossReplicas: Sum labels: - name: reason from: finish_reason diff --git a/functions/compose-metric-mapping/tests/test_fn.py b/functions/compose-metric-mapping/tests/test_fn.py index 51e87bcea..5399e5dd1 100644 --- a/functions/compose-metric-mapping/tests/test_fn.py +++ b/functions/compose-metric-mapping/tests/test_fn.py @@ -52,7 +52,7 @@ async def test_compose(self) -> None: "kind": "MetricMapping", "metadata": {"name": "my-engine"}, "spec": { - "metrics": [{"from": "my_engine_queued", "to": "modelplane_requests_waiting", "acrossReplicas": "Sum"}], + "metrics": [{"from": "my_engine_queued", "to": "modelplane_requests_waiting"}], }, } cluster = resource.dict_to_struct( diff --git a/functions/compose-serving-stack/function/collector.py b/functions/compose-serving-stack/function/collector.py index de81414cc..706f2fa47 100644 --- a/functions/compose-serving-stack/function/collector.py +++ b/functions/compose-serving-stack/function/collector.py @@ -83,8 +83,7 @@ # number that looks right. # # It is the pod's address, so it does churn on a rolling update, which is - # the cost. Aggregate it away in the query: the identity above is what to - # group by, and `acrossReplicas` names how. + # the cost. Aggregate it away in the query, grouping by the identity above. "service.name", "service.instance.id", ) diff --git a/functions/compose-serving-stack/function/stacks/metrics.py b/functions/compose-serving-stack/function/stacks/metrics.py index 39dcd6922..66dac9f22 100644 --- a/functions/compose-serving-stack/function/stacks/metrics.py +++ b/functions/compose-serving-stack/function/stacks/metrics.py @@ -98,25 +98,6 @@ } -# What describes a piece of hardware or a replica rather than a fleet. A ratio -# summed reads two replicas at half their KV cache as one at capacity, and a -# temperature summed is not a temperature at all. -# -# Everything else here is counted - requests, tokens, joules, watts drawn, -# queue depth - and a histogram can only be summed, which merges its buckets. -_MEAN = { - "modelplane_kv_cache_utilization_ratio", - "modelplane_gpu_compute_active_ratio", - "modelplane_gpu_tensor_active_ratio", - "modelplane_gpu_memory_bandwidth_ratio", - "modelplane_gpu_temperature_celsius", -} - - -def _across_replicas(target: str) -> str: - return "Mean" if target in _MEAN else "Sum" - - def _mapping( name: str, pairs: dict[str, str], @@ -128,8 +109,7 @@ def _mapping( spec=v1alpha1.Spec( metrics=[ v1alpha1.Metric.model_validate( - {"from": src, "to": dst, "acrossReplicas": _across_replicas(dst)} - | ({"fromUnit": units[src]} if src in units else {}) + {"from": src, "to": dst} | ({"fromUnit": units[src]} if src in units else {}) ) for src, dst in pairs.items() ] diff --git a/functions/compose-serving-stack/tests/test_collector.py b/functions/compose-serving-stack/tests/test_collector.py index 4f3f1dec5..f1b8217c6 100644 --- a/functions/compose-serving-stack/tests/test_collector.py +++ b/functions/compose-serving-stack/tests/test_collector.py @@ -139,7 +139,6 @@ def test_a_part_is_extracted_before_anything_selects_on_it(self) -> None: { "from": "my_engine_duration_ms", "to": "modelplane_requests_total", - "acrossReplicas": "Sum", "part": "Count", "fromUnit": "Milliseconds", "labels": [{"name": "status", "value": "ok"}], @@ -298,7 +297,6 @@ def test_a_percentage_is_divided_into_a_ratio(self) -> None: { "from": "my_engine_cache_percent", "to": "modelplane_kv_cache_utilization_ratio", - "acrossReplicas": "Mean", "fromUnit": "Percent", } ] @@ -311,14 +309,10 @@ def test_a_percentage_is_divided_into_a_ratio(self) -> None: def test_a_metric_name_cannot_end_the_comparison_early(self) -> None: """A quote in `from` would rename whatever the rest of the line matched.""" with self.assertRaises(ValidationError): - mmv1alpha1.Metric.model_validate( - {"from": 'x" or true or name == "y', "to": "modelplane_x", "acrossReplicas": "Sum"} - ) + mmv1alpha1.Metric.model_validate({"from": 'x" or true or name == "y', "to": "modelplane_x"}) for mapping in stacks.BUILTIN_MAPPINGS: for m in mapping.spec.metrics: - round_tripped = mmv1alpha1.Metric.model_validate( - {"from": m.from_, "to": m.to, "acrossReplicas": m.acrossReplicas} - ) + round_tripped = mmv1alpha1.Metric.model_validate({"from": m.from_, "to": m.to}) self.assertEqual(round_tripped.from_, m.from_) def test_a_label_value_cannot_end_the_string_it_sits_in(self) -> None: @@ -336,7 +330,6 @@ def test_a_label_value_cannot_end_the_string_it_sits_in(self) -> None: { "from": "my_engine_finish", "to": "modelplane_requests_total", - "acrossReplicas": "Sum", "labels": [{"name": "reason", "from": "finish", "values": {'ab"c': 'x"y'}}], } ] @@ -362,7 +355,6 @@ def test_carrying_a_label_onto_itself_keeps_it(self) -> None: { "from": "my_engine_finish", "to": "modelplane_requests_total", - "acrossReplicas": "Sum", "labels": [{"name": "reason", "from": "reason", "values": {"eos": "stop"}}], } ] diff --git a/functions/compose-serving-stack/tests/test_fn.py b/functions/compose-serving-stack/tests/test_fn.py index 72814a431..70fcbef17 100644 --- a/functions/compose-serving-stack/tests/test_fn.py +++ b/functions/compose-serving-stack/tests/test_fn.py @@ -1664,11 +1664,13 @@ async def test_two_destinations_cannot_name_one_exporter(self) -> None: async def test_a_stale_mapping_does_not_break_the_stack(self) -> None: """A CRD validates on write, not on what it already stored. - A MetricMapping written before acrossReplicas was required still - comes back on read without it. Parsing it raises, and raising fails - the whole pipeline step - so the serving stack composes nothing and - the fleet stops placing replicas, because one telemetry object is out - of date. Seen on a real cluster. + A MetricMapping written against an older schema comes back on read + exactly as it was stored, so a value the enum no longer carries + reaches the parser. Parsing it raises, and raising fails the whole + pipeline step - so the serving stack composes nothing and the fleet + stops placing replicas, because one telemetry object is out of date. + Seen on a real cluster, where a mapping predating a required field + did it. """ req = self._with_destination(_request("GKE", "Standard", observed=_observed_pcs())) req.required_resources["mappings"].items.append( @@ -1678,8 +1680,16 @@ async def test_a_stale_mapping_does_not_break_the_stack(self) -> None: "apiVersion": "modelplane.ai/v1alpha1", "kind": "MetricMapping", "metadata": {"name": "stale"}, - # No acrossReplicas: the schema requires it now. - "spec": {"metrics": [{"from": "old_engine_waiting", "to": "modelplane_requests_waiting"}]}, + # A unit the enum no longer carries. + "spec": { + "metrics": [ + { + "from": "old_engine_transfer", + "to": "modelplane_request_kv_transfer_seconds", + "fromUnit": "Centiseconds", + } + ] + }, } ) ) diff --git a/schemas/.lock.json b/schemas/.lock.json index 3edfa7b5e..b8c2f548b 100644 --- a/schemas/.lock.json +++ b/schemas/.lock.json @@ -1,6 +1,6 @@ { "packages": { - "fs://apis": "828906d23cc6da5be87e2b7692a96a5249f77c44ab14e012760807e09add5d96", + "fs://apis": "8718ada2c08c9f8bdad4a455f3f05e1c1cc5a83ef633259fd01848d8b5a8be84", "git://https://github.com/crossplane/crossplane/cluster/crds": "90d8b72ad8b829f0bcd7d7d5a98eaa0d579f244a", "xpkg://xpkg.upbound.io/upbound/provider-aws-ec2:v2.8.1": "sha256:ca2e9e3b2e3a8b6ca44a9700d5abf7abd733cfa388d1afe9bb2bf7c847cd39ef", "xpkg://xpkg.upbound.io/upbound/provider-aws-efs:v2.8.1": "sha256:ebb1bcd8dc9a7e60e97a324652fbd9d1b0609ddbb1107db518ce999cbc1113bd", diff --git a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py index 0d9318c62..f924b9700 100644 --- a/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py +++ b/schemas/python/models/ai/modelplane/metricmapping/v1alpha1.py @@ -65,13 +65,6 @@ class Label(BaseModel): class Metric(BaseModel): - acrossReplicas: Literal['Sum', 'Mean', 'Max'] - """ - How this metric combines over a deployment's replicas. - Every pod publishes its own series, told apart by the replica it belongs to and the instance it was scraped from, and a query over a deployment combines them. This says which combination is the right one: Sum for anything counted - requests, tokens, joules, a queue's depth. Mean for a ratio, where summing reads two replicas at half capacity as one at full. Max for a saturation figure an alert fires on, where a mean hides the replica in trouble. - Modelplane does not combine them in the collector. A scrape of one replica is one batch, so a collector that added them up would be adding readings taken at different moments, and two readings of one cumulative counter sum to twice the traffic that happened. The backend holds every replica's series and combines them at query time, where the arithmetic is right. - Required, with no default, because the wrong combination is silent: a deployment reports a number that looks entirely plausible. - """ from_: constr(pattern=r'^[a-zA-Z_:][a-zA-Z0-9_:]*$', max_length=255) = Field( ..., alias='from' )