Skip to content

[Bug]: Cluster node sandbox routing fails - OrchestratorIP overwritten with empty string #1763

Description

@alansyang

Sandbox ID or Build ID

No response

Environment

infra== 2026.03

Timestamp of the issue

2026-01-23 01:47 UTC+8

Frequency

Happens every time

Expected behavior

The sandbox should be accessible via Client Proxy. The catalog should contain the correct OrchestratorIP value obtained from the Edge Proxy's OrchestratorsPool.

Actual behavior

The OrchestratorIP in Redis catalog is an empty string, causing:

  • catalogResolution() returns ""
  • Client Proxy constructs invalid URL: http://:5007
  • All HTTP/WebSocket requests to cluster sandboxes fail

Issue reproduction

Steps to reproduce the behavior:

  • Deploy E2B infrastructure with a cluster/edge setup (non-Nomad managed nodes)
  • Create a sandbox on a cluster node
  • Attempt to access the sandbox via Client Proxy
  • Request fails because the routing URL becomes http://:5007 (invalid)

Additional context

Image

packages/client-proxy/internal/proxy/proxy.go#L43-L44

The current architecture has Edge Proxy correctly handling this, but the API Service's redundant write breaks it.

Activity

  1. linear commented on Jan 22, 2026

    @linear
  2. self-assigned this
    on Jan 23, 2026
  3. sitole commented on Jan 23, 2026

    @sitole
    Member

    @alansyang can you please describe your deployment more?

    Store sandbox should be called only once in API when the sandbox was successfully created.

    From your diagram, it looks like you are using an edge proxy? In that case, the router catalog should be accessed only via the API for the local cluster or the edge proxy when your deployment is in cluster mode.

    Theoretically, it could happen if you do an edge-type deployment and share the same Redis instance between the API and the edge API. Theoretically, to omit this case, an additional condition check would be needed in the API's addSandboxToRoutingTable method to prevent the sandbox routing record from being inserted if the sandbox is in a remote cluster.

    Is this your case? In the latest release, we removed the edge API and edge gRPC proxy from the open-source repository.

  4. alansyang commented on Jan 23, 2026

    @alansyang
    Author

    @alansyang can you please describe your deployment more?

    Store sandbox should be called only once in API when the sandbox was successfully created.

    From your diagram, it looks like you are using an edge proxy? In that case, the router catalog should be accessed only via the API for the local cluster or the edge proxy when your deployment is in cluster mode.

    Theoretically, it could happen if you do an edge-type deployment and share the same Redis instance between the API and the edge API. Theoretically, to omit this case, an additional condition check would be needed in the API's addSandboxToRoutingTable method to prevent the sandbox routing record from being inserted if the sandbox is in a remote cluster.

    Is this your case? In the latest release, we removed the edge API and edge gRPC proxy from the open-source repository.

    It's exactly as you described. I haven't seen any related changes in the latest version.

  5. sitole commented on Jan 23, 2026

    @sitole
    Member

    Between 2026.02 and 2026.03 releases we did two major changes. We removed the need to deploy edge api for open source deployments (where compute clusters run along the control plane), and then we removed edge api and gRPC proxy from the open source repository.

    This can somehow break your deployment. You mentioned you are running via edge cluster and your infra is not managed by Nomad, can you elaborate here more, please?

  6. sitole commented on Jan 23, 2026

    @sitole
    Member

    @alansyang, can you try to apply changes to your version to see if it helps?

  7. alansyang commented on Jan 23, 2026

    @alansyang
    Author

    Currently, I am attempting to deploy infra at the edge using containerization. The API discovers the Client-Proxy nodes based on the configuration in the postgre clusters table, and the Client-Proxy discovers the Orchestrator nodes through DNS. In the Orchestrator, I return the container IP where the Envd and Code-Server have been deployed.

    The changes are correct. I've tested this before, but I'm unsure if there could be any unexpected impacts from these changes.

  8. sitole commented on Jan 23, 2026

    @sitole
    Member

    Currently, I am attempting to deploy infra at the edge using containerization

    Is this a k8s deployment?

    The changes are correct. I've tested this before, but I'm unsure if there could be any unexpected impacts from these changes.

    Okay, there should be no problem with merging this.

  9. alansyang commented on Jan 23, 2026

    @alansyang
    Author

    Is this a k8s deployment?

    yes

    Okay, there should be no problem with merging this.

    thanks

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions