Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions docs/resources/cluster.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,20 +74,20 @@ const euWest = new devzero.Cluster("eu-west", { name: "production-eu-west-1" });
| Name | Type | Description |
|---------|--------|-------------------------------------------------------------------------------------------|
| `id` | string | Unique identifier of the cluster. Managed by the provider. |
| `token` | string | **(Secret)** Authentication token for the cluster agent. Stored encrypted in Pulumi state. Automatically rotated when the resource is imported and then updated. |
| `token` | string | **(Secret)** Authentication token for the cluster agent. Stored encrypted in Pulumi state. Not retrievable via `Read`/import — see the Import note below. |

## Import

An existing cluster can be imported using its cluster ID:

```shell
pulumi import devzero:index/cluster:Cluster my-cluster <cluster-id>
pulumi import devzero:resources:Cluster my-cluster <cluster-id>

# Example
pulumi import devzero:index/cluster:Cluster production "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
pulumi import devzero:resources:Cluster production "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
```

> **Note:** After importing, the `token` field will be empty in state. The next `pulumi up` will automatically call `ResetClusterToken` to obtain a fresh token.
> **Note:** After importing, the `token` field will be empty in state and stays that way — the provider does not rotate it automatically on update (doing so used to silently invalidate the credential the running in-cluster agent was using whenever an unrelated field changed). If you need the token in state, rotate it deliberately from the DevZero UI, or recreate the resource.

## Notes

Expand Down
2 changes: 1 addition & 1 deletion docs/resources/node_policy.md
Original file line number Diff line number Diff line change
Expand Up @@ -313,7 +313,7 @@ Used by all instance/zone/architecture selectors:
## Import

```shell
pulumi import devzero:index/nodePolicy:NodePolicy my-policy <policy-id>
pulumi import devzero:resources:NodePolicy my-policy <policy-id>
```

> **Note:** Because there is no delete API, `pulumi destroy` only removes the resource from Pulumi state. The policy continues to exist on the DevZero platform.
4 changes: 2 additions & 2 deletions docs/resources/node_policy_target.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,10 +127,10 @@ func main() {
An existing node policy target can be imported using its target ID:

```shell
pulumi import devzero:index/nodePolicyTarget:NodePolicyTarget my-target <target-id>
pulumi import devzero:resources:NodePolicyTarget my-target <target-id>

# Example
pulumi import devzero:index/nodePolicyTarget:NodePolicyTarget production "c84ccd96-d3f6-439d-9976-360577123fe0"
pulumi import devzero:resources:NodePolicyTarget production "c84ccd96-d3f6-439d-9976-360577123fe0"
```

## Notes
Expand Down
2 changes: 1 addition & 1 deletion docs/resources/workload_policy.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,5 +162,5 @@ Each of `cpuVerticalScaling`, `memoryVerticalScaling`, `gpuVerticalScaling`, and
## Import

```shell
pulumi import devzero:index/workloadPolicy:WorkloadPolicy my-policy <policy-id>
pulumi import devzero:resources:WorkloadPolicy my-policy <policy-id>
```
2 changes: 1 addition & 1 deletion docs/resources/workload_policy_target.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,5 +124,5 @@ target = devzero.WorkloadPolicyTarget("production-target",
## Import

```shell
pulumi import devzero:index/workloadPolicyTarget:WorkloadPolicyTarget my-target <target-id>
pulumi import devzero:resources:WorkloadPolicyTarget my-target <target-id>
```
194 changes: 194 additions & 0 deletions docs/resources/workload_rule.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,194 @@
# devzero:resources:WorkloadRule

Pins explicit resource rules directly to a single Kubernetes workload (a specific `kind`/`namespace`/`name` on a cluster). Unlike `WorkloadPolicy`, which applies a shared policy to many workloads via a `WorkloadPolicyTarget`, a `WorkloadRule` targets one workload and lets you override CPU, memory, GPU, and HPA settings with precise values — or set `autoGenerate: true` to let the engine compute them from observed usage.

## Example Usage

### Minimal example (auto-generated)

```typescript
import * as devzero from "@devzero/pulumi-provider-devzero";

const rule = new devzero.WorkloadRule("my-app-rule", {
clusterId: "cluster-abc123",
namespace: "production",
kind: "Deployment",
name: "my-api",
autoGenerate: true,
});
```

### Manual CPU, memory, and emergency-response rules

```typescript
import * as devzero from "@devzero/pulumi-provider-devzero";

const rule = new devzero.WorkloadRule("my-app-rule", {
clusterId: "cluster-abc123",
namespace: "production",
kind: "Deployment",
name: "my-api",

actionTriggers: ["on_schedule", "on_detection"],
cronSchedule: "0 2 * * *",
detectionTriggers: ["pod_creation", "pod_update"],

cpuRule: {
enabled: true,
minRequest: 10, // millicores
maxRequest: 32000, // millicores (32 cores)
targetPercentile: 0.95, // P95 of observed CPU usage
limitsAdjustmentEnabled: true,
limitMultiplier: 1.0,
},
memoryRule: {
enabled: true,
minRequest: 67108864, // 64 MiB
maxRequest: 68719476736, // 64 GiB
targetPercentile: 0.95,
limitsAdjustmentEnabled: true,
},
emergencyResponse: {
oomEnabled: true,
oomMemoryMultiplier: 1.5,
cpuThrottlingEnabled: true,
cpuThrottlingThreshold: 0.20,
cpuThrottlingMultiplier: 1.25,
},
});

export const ruleId = rule.id;
```

```python
import pulumi_devzero as devzero

rule = devzero.WorkloadRule(
"my-app-rule",
cluster_id="cluster-abc123",
namespace="production",
kind="Deployment",
name="my-api",
cpu_rule=devzero.ResourceRuleConfigArgs(
enabled=True,
min_request=10,
max_request=32000,
target_percentile=0.95,
limits_adjustment_enabled=True,
),
)
```

## Schema

### Required

| Name | Type | Description |
|-------------|--------|---------------------------------------------------------------------------------------|
| `clusterId` | string | ID of the cluster the workload lives in. |
| `namespace` | string | Kubernetes namespace of the workload. |
| `kind` | string | Workload kind. Values: `"Deployment"`, `"StatefulSet"`, `"DaemonSet"`, `"CronJob"`, `"Job"`. |
| `name` | string | Name of the Kubernetes workload. |

### Optional — Rules

| Name | Type | Description |
|---------------------|----------------------------|----------------------------------------------------------------------------------|
| `autoGenerate` | `boolean` | When `true`, the engine fills all rule fields from observed usage; manual field overrides below are ignored. |
| `cpuRule` | `ResourceRuleConfigArgs` | CPU vertical scaling rule. |
| `memoryRule` | `ResourceRuleConfigArgs` | Memory vertical scaling rule. |
| `gpuRule` | `ResourceRuleConfigArgs` | GPU vertical scaling rule (units: GPU millicores). |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Quality: Contradictory GPU unit description in workload_rule.md

In docs/resources/workload_rule.md the gpuRule row states units are "GPU millicores" (line 100), while the ResourceRuleConfigArgs.minRequest row says "bytes for memory/GPU" (line 133). These two statements contradict each other for GPU. Pick the correct unit and make both rows agree so users don't misconfigure GPU requests.

Was this helpful? React with 👍 / 👎

| `hpaRule` | `HPARuleConfigArgs` | Horizontal (replica) scaling rule. |
| `emergencyResponse` | `EmergencyResponseConfigArgs` | OOM and CPU-throttle emergency reactions. |
| `containers` | `ContainerResourceRuleConfigArgs[]` | Per-container resource overrides. When empty, workload-level rules apply to all containers. |
| `disabled` | `boolean` | Whether the rule is currently disabled. |

### Optional — Triggers & Timing

| Name | Type | Description |
|---------------------------|------------|-----------------------------------------------------------------------------------|
| `actionTriggers` | `string[]` | When to apply recommendations. Values: `"on_detection"`, `"on_schedule"`. |
| `cronSchedule` | `string` | Cron expression for scheduled application (5-field UTC). Required when `actionTriggers` includes `"on_schedule"`. |
| `detectionTriggers` | `string[]` | Events that trigger a recommendation. Values: `"pod_creation"`, `"pod_update"`, `"pod_reschedule"`. |
| `startupPeriodSeconds` | `number` | Seconds after workload start to exclude from usage data. |
| `cooldownMinutes` | `number` | Minimum minutes between consecutive recommendation applications. |
| `lookbackPeriodSeconds` | `number` | Seconds to look back for resource usage data. |

### Optional — Scaling Behaviour

| Name | Type | Description |
|-----------------------------|-----------|--------------------------------------------------------------------------|
| `schedulerPlugins` | `string[]`| Kubernetes scheduler plugins to activate. Example: `["binpacking"]`. |
| `defragmentationSchedule` | `string` | Cron expression for background node defragmentation. |
| `liveMigrationEnabled` | `boolean` | Allow live pod migration when applying recommendations without a restart.|
| `useInPlaceVerticalScaling` | `boolean` | Use in-place pod vertical scaling instead of pod restarts. |

### `ResourceRuleConfigArgs`

Used for `cpuRule`, `memoryRule`, and `gpuRule` at both the workload and per-container level. `maxScaleUpPercent`/`maxScaleDownPercent` are **not** supported on per-container rules.

| Field | Type | Description |
|---------------------------|-----------|---------------------------------------------------------------------|
| `enabled` | `boolean` | Enable this resource axis rule. |
| `minRequest` | `number` | Minimum resource request (millicores for CPU, bytes for memory/GPU).|
| `maxRequest` | `number` | Maximum resource request. |
| `targetPercentile` | `number` | Percentile of observed usage to target (0–1). Example: `0.95`. |
| `maxScaleUpPercent` | `number` | Maximum % to scale up in one step *(workload-level only)*. |
| `maxScaleDownPercent` | `number` | Maximum % to scale down in one step *(workload-level only)*. |
| `limitsAdjustmentEnabled` | `boolean` | Whether to also adjust resource limits. |
| `limitMultiplier` | `number` | Limits = request × `limitMultiplier`. |
| `limitsRemovalEnabled` | `boolean` | Actively remove limits from workloads (CPU only). |

### `HPARuleConfigArgs`

| Field | Type | Description |
|----------------------------|-----------|------------------------------------------------------------------------------------------------|
| `enabled` | `boolean` | Enable horizontal (replica) scaling. |
| `minReplicas` | `number` | Minimum number of replicas. |
| `maxReplicas` | `number` | Maximum number of replicas. |
| `maxReplicaChangePercent` | `number` | Maximum percentage change in replica count per cycle. |
| `scaleDownCooldownSeconds` | `number` | Seconds to wait between scale-down events. |
| `metrics` | `HPAMetricTriggerArgs[]` | External metric triggers only (Prometheus, queue depth, etc). CPU/Memory/Network are engine-generated. |
| `compositeFormula` | `string` | Expression combining multiple metric ratios into one scaling signal. Example: `"0.6*cpu + 0.4*memory"`. |
| `behavior` | `HPABehaviorArgs` | Fine-grained scale-up and scale-down behavior policies. |
| `fallback` | `HPAFallbackArgs` | Replica fallback when metrics become unavailable. |

### `EmergencyResponseConfigArgs`

| Field | Type | Description |
|---------------------------|-----------|---------------------------------------------------------------------|
| `oomEnabled` | `boolean` | React to OOM kills by increasing the memory request. |
| `oomMemoryMultiplier` | `number` | Multiplier applied to memory on each OOM event. |
| `oomMaxReactions` | `number` | Maximum OOM reactions before giving up. |
| `oomCooldownSeconds` | `number` | Seconds to wait between OOM reactions. |
| `cpuThrottlingEnabled` | `boolean` | React to CPU throttling by increasing the CPU request. |
| `cpuThrottlingThreshold` | `number` | Throttle ratio threshold that triggers a reaction (0–1). |
| `cpuThrottlingMultiplier` | `number` | Multiplier applied to the CPU request on a throttle reaction. |

### `ContainerResourceRuleConfigArgs`

| Field | Type | Description |
|-----------------|--------------------------|-----------------------------------------------------------|
| `containerName` | `string` | Name of the container this config applies to. |
| `cpuRule` | `ResourceRuleConfigArgs` | CPU resource rule for this container. |
| `memoryRule` | `ResourceRuleConfigArgs` | Memory resource rule for this container. |
| `gpuRule` | `ResourceRuleConfigArgs` | GPU resource rule for this container. |

### Read-Only

| Name | Type | Description |
|------|--------|------------------------------------------------------------|
| `id` | string | Unique identifier of the rule. Managed by the provider. |

## Import

An existing workload rule can be imported using its rule ID:

```shell
pulumi import devzero:resources:WorkloadRule my-app-rule <rule-id>

# Example
pulumi import devzero:resources:WorkloadRule production-api "b7f3c1a2-9d4e-4a1b-8c3f-2e5d6f7a8b90"
```

> **Note:** `autoGenerate` is not stored as its own field on the API side — it is inferred on every `Read` (including at import time) from the rule's `current_source`: `true` when the rule is engine-managed (`auto_optimization`), unset otherwise. This keeps an imported rule's mode consistent with how it actually behaves on the platform.
14 changes: 14 additions & 0 deletions provider/pkg/resources/fake_server_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -172,6 +172,20 @@ func (f *fakeBackend) GetWorkloadRecommendationPolicy(_ context.Context, req *co
return connect.NewResponse(&apiv1.GetWorkloadRecommendationPolicyResponse{Policy: p}), nil
}

func (f *fakeBackend) UpdateWorkloadRecommendationPolicy(_ context.Context, req *connect.Request[apiv1.UpdateWorkloadRecommendationPolicyRequest]) (*connect.Response[apiv1.UpdateWorkloadRecommendationPolicyResponse], error) {
f.mu.Lock()
defer f.mu.Unlock()
p := req.Msg.Policy
if p == nil {
return nil, connect.NewError(connect.CodeInvalidArgument, fmt.Errorf("policy required"))
}
if _, ok := f.wp[p.PolicyId]; !ok {
return nil, connect.NewError(connect.CodeNotFound, fmt.Errorf("workload policy not found"))
}
f.wp[p.PolicyId] = p
return connect.NewResponse(&apiv1.UpdateWorkloadRecommendationPolicyResponse{Policy: p}), nil
}

func (f *fakeBackend) DeleteWorkloadRecommendationPolicy(_ context.Context, req *connect.Request[apiv1.DeleteWorkloadRecommendationPolicyRequest]) (*connect.Response[apiv1.DeleteWorkloadRecommendationPolicyResponse], error) {
f.mu.Lock()
defer f.mu.Unlock()
Expand Down
Loading
Loading