Repository navigation
[Bug]: Cluster node sandbox routing fails - OrchestratorIP overwritten with empty string #1763
Description
Activity
@alansyang can you please describe your deployment more?
Store sandbox should be called only once in API when the sandbox was successfully created.
From your diagram, it looks like you are using an edge proxy? In that case, the router catalog should be accessed only via the API for the local cluster or the edge proxy when your deployment is in cluster mode.
Theoretically, it could happen if you do an edge-type deployment and share the same Redis instance between the API and the edge API. Theoretically, to omit this case, an additional condition check would be needed in the API's
addSandboxToRoutingTablemethod to prevent the sandbox routing record from being inserted if the sandbox is in a remote cluster.Is this your case? In the latest release, we removed the edge API and edge gRPC proxy from the open-source repository.
@alansyang can you please describe your deployment more?
Store sandbox should be called only once in API when the sandbox was successfully created.
From your diagram, it looks like you are using an edge proxy? In that case, the router catalog should be accessed only via the API for the local cluster or the edge proxy when your deployment is in cluster mode.
Theoretically, it could happen if you do an edge-type deployment and share the same Redis instance between the API and the edge API. Theoretically, to omit this case, an additional condition check would be needed in the API's
addSandboxToRoutingTablemethod to prevent the sandbox routing record from being inserted if the sandbox is in a remote cluster.Is this your case? In the latest release, we removed the edge API and edge gRPC proxy from the open-source repository.
It's exactly as you described. I haven't seen any related changes in the latest version.
Between 2026.02 and 2026.03 releases we did two major changes. We removed the need to deploy edge api for open source deployments (where compute clusters run along the control plane), and then we removed edge api and gRPC proxy from the open source repository.
This can somehow break your deployment. You mentioned you are running via edge cluster and your infra is not managed by Nomad, can you elaborate here more, please?
@alansyang, can you try to apply changes to your version to see if it helps?
Currently, I am attempting to deploy infra at the edge using containerization. The API discovers the Client-Proxy nodes based on the configuration in the postgre clusters table, and the Client-Proxy discovers the Orchestrator nodes through DNS. In the Orchestrator, I return the container IP where the Envd and Code-Server have been deployed.
The changes are correct. I've tested this before, but I'm unsure if there could be any unexpected impacts from these changes.
Currently, I am attempting to deploy infra at the edge using containerization
Is this a k8s deployment?
The changes are correct. I've tested this before, but I'm unsure if there could be any unexpected impacts from these changes.
Okay, there should be no problem with merging this.
Is this a k8s deployment?
yes
Okay, there should be no problem with merging this.
thanks
Sandbox ID or Build ID
No response
Environment
infra== 2026.03
Timestamp of the issue
2026-01-23 01:47 UTC+8
Frequency
Happens every time
Expected behavior
The sandbox should be accessible via Client Proxy. The catalog should contain the correct OrchestratorIP value obtained from the Edge Proxy's OrchestratorsPool.
Actual behavior
The OrchestratorIP in Redis catalog is an empty string, causing:
Issue reproduction
Steps to reproduce the behavior:
Additional context
packages/client-proxy/internal/proxy/proxy.go#L43-L44
The current architecture has Edge Proxy correctly handling this, but the API Service's redundant write breaks it.