Skip to content

fix: add drift detection to ProtocolMapper, Role and ClientScope - #2

Merged
NerdySoftPaw merged 1 commit into
mainfrom
fix/drift-detection-mapper-role-scope
Sep 24, 2026
Merged

NerdySoftPaw merged 1 commit into
mainfrom
fix/drift-detection-mapper-role-scope

Conversation

@NerdySoftPaw

Copy link
Copy Markdown
Member

Problem

Since the rollout of v0.12.1 on prod (2026-09-20), the ProtocolMapper, Role and ClientScope controllers PUT their definition to Keycloak on every sync (5m), even when nothing changed.

Operator writes per day on prod (Loki):

up to 18.09. (0.8.1) from 22.09. (0.12.1)
ProtocolMapper ~1,270 ~11,650
Role ~1,290 ~580

Every protocol mapper PUT invalidates the owning client in Keycloak's realm cache on all nodes, so the following token/userinfo requests reload from the DB. Result: 2–13s token latencies (~95 requests >2s per day, previously 0–3). This pushed the pace-id SLA v1 from ~99.98% down to 99.82% / 99.77% / 99.74% (21.–23.09.). Of the slow requests, 61% start within 5s of an operator write, versus 11% for random timestamps (small sample, n=28).

Fix

Compare the desired definition against the current Keycloak representation before updating, using the existing definitionsMatch (subset compare, string arrays unordered). This is the same pattern the client controller already uses:

  • ProtocolMapper: list mappers via Get{Client,ClientScope}ProtocolMappersRaw and keep the representation matched by name (findMapperByName)
  • Role: Get{Realm,Client}RoleRaw. composites is stripped beforehand as before and synced separately through the composites endpoint
  • ClientScope: GetClientScopeRaw

If the fetch fails, the controller falls through to the update, so error behaviour is unchanged. Create paths are untouched.

Verification

  • go test ./internal/... and go vet pass. New tests in keycloakprotocolmapper_drift_test.go, using the mapper shape returned by a live Keycloak 26:
    • an unchanged mapper does not count as drift
    • a config change is detected
    • a missing current representation forces an update
    • composites are stripped before the role compare
  • Compared against live prod with read-only GETs: all 40 KeycloakProtocolMapper CRs and both client scopes compare as in sync, so zero PUTs after rollout.

Rollout / check

After deploying, protocol mapper updated successfully should drop from ~480/h to ~0. Slow token requests (>2s) should disappear from SLA v1 within a day.

Related (separate PR in keycloak-realm): man-idp uses config.hideOnLoginPage, which Keycloak 26 drops in favour of the top-level hideOnLogin. That causes a perpetual IdP diff.

🤖 Generated with Claude Code

These controllers PUT their definition to Keycloak on every sync (5m)
even when nothing changed. Each protocol mapper PUT invalidates the
owning client in Keycloak's realm cache on all nodes, so the next
token/userinfo requests reload from the DB. On prod (40 mapper CRs)
this meant ~11.6k no-op writes per day and 2-13s token latencies,
pushing the pace-id SLA (v1) from ~99.98% to 99.74%.

Compare the desired definition against the current representation with
the existing definitionsMatch (subset, unordered string arrays) and skip
the PUT when in sync, as the client controller already does:

- ProtocolMapper: list mappers raw and keep the matched representation
- Role: compare via Get{Realm,Client}RoleRaw; composites are stripped
  beforehand and still synced through the composites endpoint
- ClientScope: compare via GetClientScopeRaw

A failed fetch falls through to the update, so behaviour on errors is
unchanged. Verified against live prod state: all 40 mappers compare as
in sync.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@NerdySoftPaw
NerdySoftPaw merged commit 5f7186f into main Sep 24, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant