What problem are you facing?
A prefill/decode pair is one serving instance. A request crosses both phases, the KV cache moves between them, and a replica doesn't serve until both are up. On a Dynamo cluster we never tell Grove or KAI that.
We compose one PodCliqueSet per engine, so the two phases are two independent gangs. Each is all-or-nothing on its own, but nothing co-schedules the pair, so KAI can place prefill and leave decode pending on GPUs that can't serve.
Also, an engine with a Standalone member composes a plain Deployment whatever stack its cluster runs:
|
def select_backend(engine: v1alpha1.Engine, stack: str) -> str: |
|
"""Pick the serving path for an engine from its member roles. |
|
|
|
A single Standalone member is a self-contained pod, served natively as a |
|
Deployment. A Leader plus Worker gang coordinates across nodes, served by |
|
the cluster's chosen stack: Standard (a LeaderWorkerSet, the LLMD backend) |
|
or Dynamo (a Grove PodCliqueSet). |
|
""" |
|
if engine_member(engine, ROLE_STANDALONE) is not None: |
|
return NATIVE |
|
return GROVE if stack == "Dynamo" else LLMD |
Both phases of the Kimi-K2 recipe are Standalone, so that pair reaches neither Grove nor KAI.
The fleet scheduler takes some of the sting out of this, because it only places a replica on a cluster that should have room for all of its replicas.
How could Modelplane help solve your problem?
Grove's levels look like the right tool. A PodCliqueSet holds more than two roles, and a scaling group scales its cliques as a unit while holding their ratio, which is also how you'd say one prefill per three decodes. I don't have a strong view on the layout.
We should keep in mind that currently a worker finds its leader through MODELPLANE_LEADER_ADDRESS, which we build from the scaling group's name and index, because the set-scoped variables are identical across an engine's copies:
|
def grove_leader_address_env() -> dict: |
|
"""The MODELPLANE_LEADER_ADDRESS env entry for the Grove backend. |
|
|
|
Concatenating the PCSG name and index reproduces Grove's own PodClique name, |
|
and the leader clique holds one pod, so this resolves to the leader of *this |
|
gang*. The PCS-scoped vars are identical across gangs and would point every |
|
engine.copies at gang 0's leader. |
|
|
|
Needs Grove to inject its vars before template env for the expansion to see |
|
them (grove#753, first released in v0.1.0-alpha.12-rc2, which the serving |
|
stack pins). |
|
""" |
|
leader_pod = f"$({_GROVE_PCSG_NAME_ENV})-$({_GROVE_PCSG_INDEX_ENV})-{GROVE_LEADER_CLIQUE}-0" |
|
address = leader_pod + f".$({_GROVE_HEADLESS_SERVICE_ENV})" |
|
return {"name": LEADER_ADDRESS_ENV, "value": address} |
So a set holding both phases needs a distinct leader clique per phase, and an address that still resolves to the right one.
Ranks are the other constraint. A gang command derives its own rank from a within-clique pod index, pending ai-dynamo/grove#755. Any re-layout should leave that no worse.
What problem are you facing?
A prefill/decode pair is one serving instance. A request crosses both phases, the KV cache moves between them, and a replica doesn't serve until both are up. On a Dynamo cluster we never tell Grove or KAI that.
We compose one
PodCliqueSetper engine, so the two phases are two independent gangs. Each is all-or-nothing on its own, but nothing co-schedules the pair, so KAI can place prefill and leave decode pending on GPUs that can't serve.Also, an engine with a
Standalonemember composes a plain Deployment whatever stack its cluster runs:modelplane/functions/compose-model-replica/function/backends/base.py
Lines 473 to 483 in 960c645
Both phases of the Kimi-K2 recipe are
Standalone, so that pair reaches neither Grove nor KAI.The fleet scheduler takes some of the sting out of this, because it only places a replica on a cluster that should have room for all of its replicas.
How could Modelplane help solve your problem?
Grove's levels look like the right tool. A
PodCliqueSetholds more than two roles, and a scaling group scales its cliques as a unit while holding their ratio, which is also how you'd say one prefill per three decodes. I don't have a strong view on the layout.We should keep in mind that currently a worker finds its leader through
MODELPLANE_LEADER_ADDRESS, which we build from the scaling group's name and index, because the set-scoped variables are identical across an engine's copies:modelplane/functions/compose-model-replica/function/backends/base.py
Lines 291 to 305 in 960c645
So a set holding both phases needs a distinct leader clique per phase, and an address that still resolves to the right one.
Ranks are the other constraint. A gang command derives its own rank from a within-clique pod index, pending ai-dynamo/grove#755. Any re-layout should leave that no worse.