Skip to content

Avoid parent lock for existing metric labelsets - #1216

Open
skulitom wants to merge 1 commit into
prometheus:masterfrom
skulitom:codex/reduce-label-lookup-contention
Open

skulitom wants to merge 1 commit into
prometheus:masterfrom
skulitom:codex/reduce-label-lookup-contention

Conversation

@skulitom

@skulitom skulitom commented Oct 8, 2026

Copy link
Copy Markdown

Repeated labels(...).inc() and labels(...).observe() calls currently acquire the parent metric lock even when their labeled child already exists. An existing-child lookup also waits for unrelated child construction to finish.

Return an existing child with a single dictionary lookup after the existing validation/normalization. On a miss, retain the locked recheck and child creation. Removal and collection keep their existing synchronization. The single read avoids a check-then-read race; Python documents this dictionary operation as atomic, including in free-threaded builds: https://docs.python.org/3/builtins/threadsafety.html#thread-safety-for-dict-objects.

Fixes #1215. @csmarchbanks

Validation:

  • The new blocking regression fails on the base revision and passes with the fix. Additional tests cover concurrent creation/update of the same child and recreation after each removal API.
  • Full suite, with benchmark timing disabled: Windows CPython 3.12.13: 420 passed, 12 skipped; Linux/WSL CPython 3.9.25 with optional dependencies: 432 passed; CPython 3.14.8: 420 passed, 12 skipped; free-threaded CPython 3.14.8: 420 passed, 12 skipped; PyPy 3.9.19: 429 passed, 3 skipped.
  • tox -e flake8,isort,mypy: all passed.
  • Additional free-threaded stress check: 120,000 updates concurrent with 20,000 removals/clears and 5,000 collections, followed by checking that clearing and recreating a child resets its value.

Local benchmark (Linux/WSL, CPython 3.14.8, 32 logical CPUs): median of five samples, 50,000 operations per thread, one pre-created labelset per thread. The before/after runs used the same interpreter and separate source trees, without simultaneous test runs. The free-threaded process reported the GIL disabled. Throughput below is million operations/second, with eight threads:

Operation Default Python before → after Free-threaded Python before → after
labels() 2.348 → 3.162 1.168 → 2.069
labels().inc() 1.527 → 1.823 0.913 → 5.235
labels().observe(0.2) 1.066 → 1.223 0.969 → 4.206
Cached child .inc() (control) 5.336 → 5.266 12.726 → 13.143

These are positional-label microbenchmarks, not an application-wide speedup claim. Caching a child remains faster. The script also measures one and four threads; at one thread, labels().inc() improved from 1.545 to 1.837 Mops/s on default Python and 1.147 to 1.395 Mops/s on free-threaded Python. Unchanged control measurements varied by up to about 9% across the thread counts.

Reproduction script

Save as benchmark_labels.py outside the checkout and run with each source revision on PYTHONPATH, using the same interpreter:

PYTHONPATH=/path/to/checkout python benchmark_labels.py before --iterations 50000 --repeats 5
"""Bounded repeated measurements of existing-label instrumentation paths."""
import argparse
from concurrent.futures import ThreadPoolExecutor
import json
import os
import platform
import statistics
import sys
from threading import Barrier
from time import perf_counter

from prometheus_client import CollectorRegistry, Counter, Histogram
import prometheus_client


def measure(mode, threads, iterations):
    registry = CollectorRegistry()
    cls = Histogram if mode == "histogram" else Counter
    metric = cls("bench", "Benchmark", ["worker"], registry=registry)
    children = [metric.labels(str(index)) for index in range(threads)]
    barrier = Barrier(threads + 1)

    def run(index):
        label = str(index)
        child = children[index]
        barrier.wait()
        if mode == "lookup":
            for _ in range(iterations):
                assert metric.labels(label) is child
        elif mode == "counter":
            for _ in range(iterations):
                metric.labels(label).inc()
        elif mode == "histogram":
            for _ in range(iterations):
                metric.labels(label).observe(0.2)
        else:
            for _ in range(iterations):
                child.inc()

    with ThreadPoolExecutor(max_workers=threads) as pool:
        futures = [pool.submit(run, index) for index in range(threads)]
        started = perf_counter()
        barrier.wait()
        for future in futures:
            future.result(timeout=60)
        elapsed = perf_counter() - started
    if mode != "lookup":
        sample = "bench_count" if mode == "histogram" else "bench_total"
        for index in range(threads):
            assert registry.get_sample_value(sample, {"worker": str(index)}) == iterations
    return elapsed


parser = argparse.ArgumentParser()
parser.add_argument("label")
parser.add_argument("--iterations", type=int, default=30000)
parser.add_argument("--repeats", type=int, default=3)
args = parser.parse_args()
print(json.dumps({"label": args.label, "python": sys.version, "platform": platform.platform(),
                  "gil_enabled": getattr(sys, "_is_gil_enabled", lambda: True)(),
                  "cpu_count": os.cpu_count(), "module": prometheus_client.__file__,
                  "iterations_per_thread": args.iterations, "repeats": args.repeats}), flush=True)
for mode in ("lookup", "counter", "histogram", "cached_counter"):
    for threads in (1, 4, 8):
        measure(mode, threads, 1000)
        samples = [measure(mode, threads, args.iterations) for _ in range(args.repeats)]
        elapsed = statistics.median(samples)
        print(json.dumps({"mode": mode, "threads": threads, "seconds": samples,
                          "median_ops_per_second": args.iterations * threads / elapsed}), flush=True)

AI assistance: implementation, tests, and benchmark execution were prepared with Codex.

Signed-off-by: skulitom <artemskulimovskiy@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optimize labels() existing-child lookups to avoid parent lock contention

1 participant