Conversation
Propose condition-based access control — grants can optionally specify conditions that must be satisfied for the grant to take effect. Two condition types: - Resource conditions (tags, aliases): control which resources a grant applies to - Request conditions: restrict what values a user can set Enables dynamic permission boundaries within a workspace (dev/prod lifecycle, promotion protection) without moving resources between workspaces. Reuses MLflow's search filter syntax, preserves the allow-only model, and adds no extra DB queries for resource attributes. Related: mlflow/mlflow#24742
Rework the condition model so conditions attach to a grant keyed by capability (can_read/can_use/can_update/can_delete/can_manage) rather than the grant as a whole. Capabilities are atomic, so each condition is independent by construction with no cross-capability inheritance; an absent capability key is unconditional. Add capability composition shorthand (higher-references-lower, expanded 5-condition limit) and a level ceiling. Recast condition-kind compatibility as (kind x capability x resource_type) validity derived from the request-context extractor registry (fail-closed at grant creation), resolving the missing-evaluator open question. Add a "Relationship to RFC 0008" section: capabilities map to 0008's action verbs; request_attributes/resource_attributes are proposed requirement extensions; a predicate channel extends AuthorizedResources for search. Add a Use cases section mirroring the sub-resource RFC, and make grant loading lazy (conditions loaded on the fine pass for surviving grants only).
The "Condition shape" and "Condition syntax" subsections both documented the <entity>.<key> <operator> <value> format. Consolidate the entity and operator tables into "Condition shape" alongside the examples and remove the redundant "Condition syntax" section.
… map Replace the parallel REQUEST_CONTEXT_EXTRACTORS registry with the per-operation validator that already exists in BEFORE_REQUEST_HANDLERS and already reads the request. The validator surfaces the attributes it read on request_attributes, dispatched from the before-request hook — so there is no separate map to keep in sync, and the RFC now states where extraction runs. Update the two-phase resolver, the 0008 relationship section, the condition-kind validity rule, and open question 2 to reference validator- surfaced attributes instead of the registry.
Relocate "Relationship to RFC 0008" to the top of Detailed design (right after Condition model, before API change) so the contract framing precedes the API/schema/performance mechanics that depend on it. Replace the resolve/authorize() code block with a concise four-step prose description; the two-phase flow is already shown by the mermaid diagram above, so the code was redundant.
…e example Replace the abstract can_read/can_update example with a concrete one: a can_read resource-tag condition (tags.stage='dev') gates reading existing resources but does not apply to can_use in its create-in-workspace sense (no resource exists yet). The point is that per-capability keying means the read gate never has to be reasoned out of applying to create — not leak prevention.
…eferred Add the considered alternative — a flat clause set the permission model filters per action at evaluation time — and explain why per-capability keying is preferred: the flat approach is opaque to readers/admins (which clause governs which action lives in engine logic, not the grant), so it does not visibly honor the authored clauses. Explicit keying makes the action-condition mapping authored and visible: what you read is what is enforced.
…ated content Split Detailed design into "Condition model (authoring & customer experience)" and "Implementation" (starting at the validity rule). Group creation-time rules (capability scoping, shape, limits, validity + grant-creation processing) in the CX/validity sections and the evaluation/wiring in Implementation. Genericize request-context extraction to one paragraph + sequence diagram (reuse the operation's existing request parsing; no parallel registry). Reuse SearchUtils (parse_search_filter, get_comparison_func, _does_*_match_clause) for evaluation instead of a bespoke engine. Rewrite search filtering around RFC 0008 list_authorized push-down (predicate channel + all=None fallback), dropping the post-response fast/slow-path code. Show the concrete AuthorizationRequirement / AuthorizedResources shape additions (request_attributes, resource_attributes, predicate). Distribute the RFC 0008 relationship into the relevant sections, keeping a cross-cutting notes subsection. Fold the flat-condition-set alternative into Alternatives §E. Remove the redundant API-change section (covered by CX + creation processing), the outdated allow-only discussion (Drawbacks bullets retained), and resolved open question 2.
- Fold Implementation details into Detailed design to match the template. - Store conditions as the filter string; add an open question on persisting the decomposed representation for optimization. - Rework Drawbacks: merge the cross-role non-monotonic points, reframe as type-level authoring vs operation-level enforcement, add no-read-restriction, drop the metadata-layer bullet. - Remove 'Why a new resource type'; fold its rationale into Alternative A. - Merge Deferred into Out of scope; mark per-operation target granularity an intentional non-goal rather than future work. - Replace the order-of-evaluation step list with a lead-in plus diagram.
Describe displaying a role's mutation conditions alongside its grants and editing them via the add/update/remove/list API, using the stored filter strings directly. No new UI framework or data shape is required.
| is deliberate, so conditions are never mistaken for read gates. Read conditions and search or | ||
| list prefiltering need a queryable scope plus predicate pushdown into the store query, and are | ||
| a follow-up. | ||
| - **Parent-resource conditions.** A condition reads `tags.*`/`aliases.*` only from the resource |
There was a problem hiding this comment.
I think we should allow parent-resource conditions for registered models and prompts. For example, it seems reasonable to say, “On registered model A, data scientists cannot modify versions with this tag or alias.” Without that, the permission feels either too broad (all versions) or not dynamic (a separate permission for each version).
There was a problem hiding this comment.
Will poc this and updated the design accordingly
| both and can mutate neither. This is the safe direction for a restriction, but admins should | ||
| be aware of it. | ||
| - **At most 5 clauses per condition** (value and target each). | ||
| - **At most one value condition and one target condition per `(role, resource_type)`.** |
There was a problem hiding this comment.
I think we either need to allow more than 5 clauses per condition or allow multiple conditions per role and resource type. Otherwise, it becomes limiting if you have a complex workflow.
There was a problem hiding this comment.
Multiple conditions per role makes sense, maybe at a limit of 10 to put an upper bound on auth eval time. Load test could confirm final numbers. Having a lot of clauses in a condition statement will make it hard to display in the UI and possibly the SDK response.
| supported type determines the target vocabulary. Each operation contributes only its value | ||
| (settable) fields. | ||
|
|
||
| | Resource type | Target fields | Value-setting operations | |
There was a problem hiding this comment.
Let's include prompts too.
| supported type determines the target vocabulary. Each operation contributes only its value | ||
| (settable) fields. | ||
|
|
||
| | Resource type | Target fields | Value-setting operations | |
There was a problem hiding this comment.
How does "create" work here? Only experiments with a specific tag can be created?
There was a problem hiding this comment.
Yes, or can also be used to restrict setting certain tags on create.
|
|
||
| | Resource type | Target fields | Value-setting operations | | ||
| |---|---|---| | ||
| | `experiment` | `tags.*` | set-tag, create. Other mutations (update, delete, restore, delete-tag) carry no value fields. | |
There was a problem hiding this comment.
If the target condition is tags.lifecycle = 'dev' and a user has that role, can they delete the lifecycle tag?
There was a problem hiding this comment.
If they have Manage perms then yes, and then they would loose access similar to if tags.lifecycle changes from dev to prod.
| authorize(subject, requirement) -> Decision # unchanged signature; one call, one decision | ||
| ``` | ||
|
|
||
| Because both fields default to empty, existing callers and backends are unaffected: a backend |
There was a problem hiding this comment.
Could we have the API reject adding conditions if the authorization backend does not support them?
| Core-->>Client: 200 OK or 403 Forbidden | ||
| ``` | ||
|
|
||
| ## MLflow UI |
There was a problem hiding this comment.
It'd be helpful if there was a UI change that allowed you to set common conditions on say the production alias when viewing a registered model or prompt.
There was a problem hiding this comment.
Sure, might be a fast follow depending on release timing/ my bandwidth
There was a problem hiding this comment.
Tracking as fast follow (out of scope)
| - A batch set-tags operation sets multiple tags in one call, so the value condition is | ||
| evaluated against each tag, and if any tag fails the operation is denied. | ||
|
|
||
| ## Adaptability to Pluggable Auth (RFC 8) |
There was a problem hiding this comment.
Would it make sense for MLflow to always own the condition based grants and essentially do the filtering after the baseline resource permission is validated by the authorization backend? Otherwise, I think many authorization backends would not support this or have to implement their own storage for these grants.
There was a problem hiding this comment.
I think that the auth backend should own the storing and condition evaluation given purpose of the backend being customizable and opt in for different auth features.
We can have core support FeatureNotSupported response from auth backend for condition CRUD apis. The backend is free to respond with an AuthorizationDecision without integrating any conditions loading/processing logic (e.g. make no changes there).
| CONSTRAINT unique_role_resource_type UNIQUE (role_id, resource_type) | ||
| ); | ||
|
|
||
| CREATE INDEX idx_type_conditions_role_id ON type_conditions(role_id); |
There was a problem hiding this comment.
Is there a unique constraint on role_id and resource_type based on the described limitations above? Also, an index with both role_id and resource_type would be preferable to just this standalone one.
There was a problem hiding this comment.
Updated the index requirements
| - A batch set-tags operation sets multiple tags in one call, so the value condition is | ||
| evaluated against each tag, and if any tag fails the operation is denied. | ||
|
|
||
| ## Adaptability to Pluggable Auth (RFC 8) |
There was a problem hiding this comment.
Could the backend and therefore the UI show a clear error message of why the permission is denied (e.g. this has the production alias)?
There was a problem hiding this comment.
Given that a Read grant is minimum necessary before conditions are evaluated we could include details in the error message
| The condition carries no workspace of its own, exactly as a grant (`role_permission`) carries | ||
| none and inherits its workspace from the role. | ||
|
|
||
| ### Admin Apis |
There was a problem hiding this comment.
Could we add a preview API so an admin can test a proposed condition against a particular resource before applying it to a role? I’d want to see whether a specific operation would be allowed or denied, and which condition caused the denial. Since conditions combine across roles, it would also be useful to distinguish “does this condition match?” from “would this user actually have access?” This would make it much easier to check a condition before rolling it out.
There was a problem hiding this comment.
Leaning towards having this as a fast follow but it would definitely improve the UX
There was a problem hiding this comment.
Tracking as fast follow
|
|
||
| # Open questions | ||
|
|
||
| 1. **Storing the decomposed representation for optimization.** Conditions are stored as the |
There was a problem hiding this comment.
I think it's worth caching at least. I say just keep this in the back of your mind while implementing. Something obvious may come to you.
mprahl
left a comment
There was a problem hiding this comment.
Seems like a great feature!
Support up to 100 role-owned conditions per resource type with optional exact direct-parent scope. Document workspace-scoped lookup, parent-aware indexing, parser caching, prompt support, and updated backend and UI behavior.
Seven additions, each a rule that currently lives only in the code. The most important is a correction rather than an addition. Tag deletion exposes the tag key. The resource-type table listed only set-tag as value-setting and swept delete-tag into "other mutations carry no value fields", and the closing note then described deleting a governing tag as authorized with "subsequent mutations no longer match the condition". Read literally that documents a two-step bypass: block setting a tag, permit deleting it, then operate unrestricted. The implementation has never behaved that way -- a deletion is matched against tag_key, and on the target side an absent tag fails every comparator, so removing the governing tag loses access rather than escaping the condition. The table and the note now say so, including that `tags.lifecycle != 'prod'` denies an untagged resource. Reserved `mlflow.*` keys are refused in both namespaces, with the reason: MLflow writes them itself, so a value condition naming one blocks its own bookkeeping -- and because no value condition can gate them, they stay freely settable, so a target condition reading one restricts nothing. A rule that cannot be enforced should not be storable. Ordering comparators are rejected, because tag values are strings and `>` would compare lexicographically while reading as a numeric range. The two MCP types are added, with the note that aliases live on the server so an alias condition is never keyed on the version. An unresolved parent must fail closed. The previous wording made a scoped condition merely inapplicable when a validator could not resolve the parent, which is the fail-open direction and indistinguishable from no restriction. The conjunctive-objects hazard is stated at the authoring surface, not only in the limits list: a list of up to 100 objects reads like alternatives and most permission systems treat a list that way, but "env = dev" plus "owner = me" permits only resources that are both. Slot allocation is server-side and the loser of a race retries another slot -- the UNIQUE constraint makes the bound race-safe but does not pick the slot.
| # - a filter string provided sets or replaces that condition | ||
| # - null clears that condition | ||
| # - omitted leaves that condition unchanged | ||
| add(role_id, resource_type, parent_resource_type?, parent_resource_id?, |
There was a problem hiding this comment.
Does it make sense to allow for conditions to apply globally to all workspaces?
There was a problem hiding this comment.
It makes sense, but would need to a cross-workspace role or some other permission binding to work which we don't have today. Will treat it as a feature enhancement
There was a problem hiding this comment.
I was thinking it could be restricted to only admins setting that. It's not a blocker though.
| value_condition TEXT, | ||
| target_condition TEXT, | ||
|
|
||
| CHECK (condition_slot BETWEEN 1 AND 100), |
There was a problem hiding this comment.
Optional: I'd consider dropping this since if there is a race that allows 101 conditions, it's not that big of a deal IMO.
There was a problem hiding this comment.
Makes sense, will remove
mprahl
left a comment
There was a problem hiding this comment.
I left a few non-blocking comments.
Clarify the role-owned workspace lookup, exclude cross-workspace conditions, and retain the role/type cap without a database range check. Document backend slot allocation and parent-aware indexed lookup behavior.
Summary
Proposes mutation conditions for MLflow RBAC: role-owned restrictions that layer on top of existing grants and gate create/mutation operations by resource state and by values in the request. Conditions only ever further-restrict what a grant already allows; they never confer access, and reads are never gated.
Problem: MLflow RBAC grants access by resource type and identity, not by resource state. Once a user has EDIT on registered models in a workspace, they can modify any model regardless of lifecycle stage and set operationally significant tags or aliases. There is no way to express "edit dev models but not production ones," or "may not set the
championalias," without moving resources between workspaces and breaking lineage.Condition object: A role can have up to 100 condition objects per target resource type. Each object contains optional filters plus an optional exact direct-parent scope:
tags.lifecycle = 'dev') gates which existing resources may be mutated, evaluated against the target's current tags/aliases. It is mutation-only and vacuous on create.tag_key != 'lifecycle' AND alias != 'champion') gates what values may be set, evaluated against the request. It applies to create as well as mutation.parent_resource_type+parent_resource_idrestrict the object to children of one exact direct parent, e.g. versions under registered modelfraud-model.Each filter is an AND-ed expression in MLflow's existing search filter grammar.
INandNOT INexpress same-field alternatives; general OR is intentionally excluded to preserve reuse of the existing parser/evaluator and avoid special absent-field semantics for value conditions.Composition — grants add, conditions subtract: a user's capability remains the union of their roles' grants. Every condition applicable to the target's type and scope across every held role must pass. Parent scope selects conditions before they combine, so a condition scoped to model A does not affect a version of model B.
Key use cases:
Design highlights:
mutation_conditionstable with a race-safe 100-object bound and a composite lookup index on role, target type, and parent scope. Existing deployments have no rows and remain default-allow.add/update/remove/list, including partial updates, parent scope, automatic removal when both filters are cleared, and a feature-not-supported response for auth backends that do not implement conditions.experiment,registered_model,prompt,run,trace,logged_model,registered_model_version, andprompt_version. Trace conditions can also gate assessment mutations because those operations authorize against their trace.Explicitly out of scope: read-scoping / hiding resources on read; predicates over mutable parent attributes; per-operation target granularity; fields beyond tags and aliases; assessment-owned metadata/feedback conditions; and owner-based conditions.
Related: mlflow/mlflow#24742