Skip to content

fix(topology): ask for bus topology when adding a gateway, prune whole duplicate devices (#524) - #525

Merged
GreenGrassBlueOcean merged 2 commits into
OpenWebNet-HA:v2-phase1-architecturefrom
GreenGrassBlueOcean:fix/shared-bus-onboarding
Sep 28, 2026
Merged

GreenGrassBlueOcean merged 2 commits into
OpenWebNet-HA:v2-phase1-architecturefrom
GreenGrassBlueOcean:fix/shared-bus-onboarding

Conversation

@GreenGrassBlueOcean

@GreenGrassBlueOcean GreenGrassBlueOcean commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #524 (split out of #523).

Problem

Adding a second gateway on a bus that a configured gateway already serves:

  1. The new entry starts as a standalone primary, because topology was only in the options flow. Its startup sweep and passive discovery create every device the other gateway already has. Live check on a single MH200: one *#2*0## returns 11 cover addresses (11#4#02, 85, 12–16, 18, 19, 21, 22 behind #4#02). That's 11 cover devices plus 33 buttons before the user can do anything.
  2. Once the entry is made a follower, PlatformDiscovery.restore() prunes the duplicate actuators (Two gateways in one installation? Help us support shared buses (MH200N + MH201, MHS1 + LN4890, …) #453). It left their lock / unlock / calibrate buttons registered, and button.py no longer rebuilds them, so they stay behind as unavailable orphans on a device without its actuator. This is what @anotherjulien saw.

Changes

A. Topology in the initial config flow. When another IP gateway is configured (not a follower, not USB/serial), a new bus_topology step runs before the entry is created:

  • Standalone (default): creates the entry with bus_topology: standalone stored.
  • Shared: creates the entry as secondary / warm standby of the chosen gateway, so its first setup already skips the sweep and discovers only delegated WHOs.
    • Role and delegated WHOs are suggested by recommend_follower(), which keeps the configured gateway as primary because it already owns the devices. For example, an MH202 next to a MyHomeServer1 takes WHO 16+22; next to an F454 it becomes a standby.
    • A standalone primary is promoted to shared primary.
    • The answer is checked with the existing validate_shared_bus_topology(), and any matching shared-bus repair issue is dismissed.
  • A single-gateway install never sees the step.

B. Prune whole devices, not just the entity.

  • discovery.prune_entity() removes the entity together with its -disable / -enable / -calibrate buttons, then the device once nothing is left on it. restore() now uses it.
  • prune_orphaned_companions() runs at follower button setup. It removes buttons left behind by earlier versions, but only when this gateway no longer has the actuator and the primary does. Buttons of an actuator a user deleted are left alone.
  • The device-level "Calibrate all covers" button (one per gateway, not per actuator) sat outside this pruning entirely and could stay present-but-pointless on a follower with no covers of its own. available now also requires at least one cover entity on this config entry (found live by @anotherjulien on fix(topology): ask for bus topology when adding a gateway, prune whole duplicate devices (#524) #525, fixed in d2e2fa32).

topology.infer_shared_bus_topology() now shares its follower-delegation logic (_follower_delegation) with recommend_follower(). Its behaviour is unchanged.

Strings: new step and errors in strings.json / en.json, with nl/fr/it translations.

How it works

A. Adding a gateway (config flow)

  user pick / SSDP / manual IP
               │
               ▼
        test_connection ──── fails ───▶ password step / abort
               │ ok
               ▼
  another gateway configured that could be the primary?
  (IP gateway, not ignored, not itself a secondary/standby)
               │
       no ─────┴───── yes
       │               │
       ▼               ▼
  create entry     bus_topology step
  (unchanged)      suggested role + WHOs = recommend_follower(existing, new)
                       │
       standalone ─────┴───── shared
            │                    │
            ▼                    ▼
  create entry with     validate_shared_bus_topology(new entry,
  bus_topology =          primary as it will be after promotion)
  standalone                     │
                     errors ─────┴───── ok
                        │                │
                        ▼                ▼
                re-show the form   primary standalone? ─▶ make it shared primary
                with the answers   dismiss the shared-bus repair issue, if any
                                   create entry as secondary / standby
                                         │
                                         ▼
                            first setup is already a follower:
                            no startup sweep, discovers only its delegated WHOs

B. Follower setup: removing duplicates without leaving orphans

  follower gateway set up
     │
     ├─▶ light / switch / cover platform: PlatformDiscovery.restore()
     │      for each registry entity of this gateway
     │         │
     │         ▼
     │      same unique id exists on the primary's MAC? ── no ──▶ restore it
     │         │ yes
     │         ▼
     │      prune_entity()
     │         ├─ remove the entity
     │         ├─ remove <uid>-disable / -enable / -calibrate buttons
     │         └─ device left empty? ── yes ──▶ remove it from this entry
     │
     └─▶ button platform: prune_orphaned_companions()
            (cleans up what earlier versions left behind)
            for each <actuator uid>-disable / -enable / -calibrate button
               │
               ▼
            actuator still on this gateway? ── yes ──▶ keep (restore() handles it)
               │ no
               ▼
            actuator on the primary? ── no ──▶ keep (the user deleted it)
               │ yes
               ▼
            prune_entity(button) ─▶ device removed once empty

Either platform may set up first. If the buttons are rebuilt before restore() prunes their actuator, removing the registry entries takes the live button entities down too.

C. Holding back discovery on an unconfirmed entry: feasibility

As proposed ("hold discovery while topology is unconfirmed and other gateways exist"): feasible, but I don't recommend it.

  • The code is small: gate PlatformDiscovery._discovers() and initial_discovery() on CONF_BUS_TOPOLOGY in entry.options.
  • After A, every entry created next to another gateway has bus_topology stored. The only unconfirmed entries left are ones created by earlier versions, where the duplicates already exist.
  • Gating those would silently stop discovery on upgrade for every existing multi-gateway user with separate plants. The benefit is close to zero and the risk is real.

Better variant: detect instead of hold, to prefill A's answer.

  • During the bus_topology step, the new gateway would send a few point status requests (*#1*<addr>##) for addresses the configured gateway already has entities for. It would then listen on that gateway's myhome_message_<mac> signal for the replies for about 2 s. Hearing them means same bus, so Shared is preselected.

  • It is read-only (status requests only) and runs while both gateways are connected.

  • The existing TX→RX echo detector (_correlate_shared_bus_traffic) can't do this job. It needs 3 command echoes within 5 min, which a quiet plant may never produce, and commands change state, so they can't be used as a probe.

  • Confirmed on a real bus, live: same-bus replies do cross, but only in one direction. @anotherjulien tested this on his F454 + MH202 (#524), sending point status requests (light *#1*33##, cover *#2*31##) from each gateway while tracing both, plus *#2*99## (a nonexistent address) as a control:

    Direction Result
    F454 sends → reply on MH202? 4/4, 81–128 ms later
    MH202 sends → reply on F454? 0/4
    *#2*99## (nonexistent) → either side? none, in either trace

    The MH202→F454 misses aren't a capture gap: in that exact window the F454's trace caught one unrelated bus frame (proving its monitor was live), just never a reply to anything the MH202 sent. So this is a real, repeatable asymmetry on that bus — not a trace-capture artifact — and the control frame rules out a false "same bus" from either direction.

  • What this means for the probe: my original sketch ("new gateway sends, existing gateway listens") is exactly the direction that failed on this plant. The reliable direction is the other way round: the already-configured gateway sends, the new gateway listens. That's slightly more work than the original estimate, since the new gateway only has a short test connection during onboarding, not a running event listener yet — it needs a temporary monitor session for those couple of seconds:

    bus_topology step shown ─▶ pick 2-3 point addresses the existing gateway
                               has entities for (never general/area/group, no WHO 13)
                                    │
                                    ▼
         existing gateway ── *#1*<addr>## ──▶ bus
                                    │
                                    ▼
         temporary monitor session on the new gateway, ~2 s:
         *1*<what>*<addr>## ?
                                    │
                 every probe answered ──┴── none / some answered
                          │                          │
                          ▼                          ▼
              preselect "shared" (user still   keep the "standalone" default
              confirms)
    
  • Next step: a follow-up PR once A/B have settled, roughly 150 lines plus tests (the extra size over my original estimate is the temporary monitor session). Whether this asymmetry is specific to Julien's plant (a bus coupler, a firmware difference in how the two models arbitrate the bus) or general to F454+MH202 pairs is still open; it doesn't block the design above either way, since only the working direction is needed.

Tests

  • tests/test_shared_bus_onboarding.py (15 tests):

    • the recommendation;
    • the flow: first gateway, candidates, standalone, shared plus promotion, second-standby refusal, primary removed while the form is open;
    • pruning with buttons and the device, keeping a device that still has entities, orphan cleanup, follower-only button setup, and the calibrate-all-covers button's availability.
  • Full suite (WSL, core 2026.9.2): 2618 passed, 100.0 % coverage. Ruff and verify_ha_standards.py pass.

  • mypy-ratchet passes in CI. Locally, the newer WSL core reports 8 strict errors in code this PR doesn't touch; the two this PR introduced are fixed.

  • V2 Architecture: New entities implement handle_event() and do not manually subscribe via async_dispatcher_connect. (No new entities.)

🤖 Generated with Claude Code

…e duplicate devices (OpenWebNet-HA#524)

A gateway added next to a configured one was set up as a standalone
primary: its startup sweep discovered every device the other gateway
already had. Once it was made a follower, only the duplicate actuators
were pruned, leaving their lock / unlock / calibrate buttons unavailable
on devices without an actuator.

- Config flow: when another IP gateway is configured, a bus_topology step
  asks standalone or shared before the entry exists. Shared creates the
  entry as secondary / standby of the chosen gateway (role and delegated
  WHOs suggested from both models) and promotes a standalone primary to
  shared primary.
- Pruning a follower's duplicate now also removes its companion buttons
  and the device once empty; button setup cleans up orphans left by
  earlier versions (only when the primary has the actuator).
- topology: split the follower delegation out of
  infer_shared_bus_topology; recommend_follower keeps the configured
  gateway as primary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@GreenGrassBlueOcean

Copy link
Copy Markdown
Contributor Author

Testing this PR: one line from the Home Assistant terminal

@anotherjulien, you can try this on your F454 + MH202 without a release or HACS. From the Terminal & SSH add-on:

curl -fsSL https://github.com/GreenGrassBlueOcean/MyHOME/releases/download/pr-525-test/install.sh | bash

It backs up your current integration to /config/myhome_backup/myhome-<timestamp>, installs the test build, runs ha core check and restarts Home Assistant. The build logs at DEBUG level, so ha core logs shows every frame.

What's in it (pre-release pr-525-test):

Part Version
MyHOME this PR, 64a08bb2
OWNd master @ 6002fd8 + open PR OpenWebNet-HA/OWNd#63 @ 1012021, as wheel 2.0.0b9.dev202609281413 (the manifest pins it; HA installs it on restart)

The full MyHOME suite passes against that OWNd (2617 tests).

Important

We urgently need an OWNd release. master is 47 commits ahead of v2.0.0-b8, and OpenWebNet-HA/OWNd#63 (MH200N WHO 16 audio support, generated profiles table) is still open. MyHOME pins OWNd==2.0.0b8, so none of these fixes reach users: the MH200 profile with WHO 16 (#53), no retry after a status-request NACK (#57), unknown lighting WHATs (#59), WHO 4 dimension 7 zone state (#60/#62), shutter levels out of range (#56) and more. Only test builds like this one get them. @anotherjulien, could you review/merge #63 and cut a 2.0.0b9? Then we can bump the MyHOME pin and drop the patched wheel.

What to check:

  1. Leftovers from your first try (B). After the install's restart, with your MH202 still set up as secondary, the cover devices that only had unavailable Lock / Unlock / Calibrate travel time buttons should be gone. The log line is Pruned orphaned button …: its actuator lives on primary gateway ….
  2. The new setup step (A). Delete the MH202 entry and add it again. After the connection test you should get "Is this MH202 on the same bus as another gateway?". Pick Shared SCS bus, primary = F454; the role and subsystems are pre-filled. Expected:
    • no *#2*0## sweep from the MH202 and no duplicate cover devices;
    • the F454 is switched to shared primary if it wasn't already.
  3. If you have 5 more minutes, the same build is fine for the two-gateway probe traces I asked for in #524.

Roll back:

rm -rf /config/custom_components/myhome && cp -a /config/myhome_backup/<the backup dir> /config/custom_components/myhome && ha core restart

The restored manifest pins OWNd==2.0.0b8 again, and HA reinstalls it on restart.

@anotherjulien

Copy link
Copy Markdown
Member

Thanks for the neat package 🙂

So:

  1. Yes, the cover devices are gone, with the mentioned log line "Pruned orphaned button button.cover_..."
    BUT, there is still an available "calibrate all covers" button on the MH202 Device under the MH202 config
  2. Config workflow works as described
Skipping WHO=2 discovery: follower gateway on shared bus.
Skipping WHO=4 discovery: follower gateway on shared bus.
Skipping WHO=16 discovery: follower gateway on shared bus.
TCP keepalive enabled on event socket (idle=30s, intvl=10s, cnt=3).

If the F454 is standalone, it is correctly switched to principal on a shared bus
BUT, the "calibrate all covers" button on the MH202 Device under the MH202 config is still there, but now unavailable.

On my way to perform the tests in #524 !

@anotherjulien

Copy link
Copy Markdown
Member

Quick note: if I add it through the config flow as a secondary with delegated WHO2, the cover devices still appear under the F454.
At this point I'm not clear if this is the intended way to be displayed or not 😅

…gateway

anotherjulien found it on OpenWebNet-HA#525: after a follower's duplicate covers are
pruned in favour of the primary, its device-level "Calibrate all covers"
button stayed present and looked available even though this gateway has
nothing left to calibrate. `available` now also requires at least one
cover entity of this config entry, matching what `async_press` already
checks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@GreenGrassBlueOcean

Copy link
Copy Markdown
Contributor Author

Thanks for both — pushed a fix for the first, and the second is expected behaviour.

1. "Calibrate all covers" button unavailable-but-present: fixed in d2e2fa32. It's a device-level button (one per gateway, not per actuator), so it sat outside the per-actuator pruning entirely — nothing removed or hid it. available now also requires at least one cover entity still on this config entry, same check async_press already used. On your MH202 (no covers of its own once the duplicates are pruned) it will now consistently show unavailable instead of a live-looking button that silently does nothing when pressed. The updated build is on the same pr-525-test release — re-run the one-liner from my earlier comment to pick it up.

2. Covers "still appear under the F454" with WHO 2 delegated to the MH202: this is existing #453 behaviour, not something #525 changes. Delegation only decides who discovers new devices going forward — it never moves devices a gateway already owns. Your F454 discovered those covers when it was still standalone, long before the MH202 was configured; delegating WHO 2 to the MH202 afterwards doesn't retroactively transfer them. Concretely: the MH202 does try to create them from the bus, but _create()'s follower check (_owned_by_primary) sees the F454 already has that unique id and skips creating a duplicate — so they simply stay where they are. If you want them to actually live on the MH202, the current way is to delete them from the F454's registry (or delete/re-add the F454 device) so they get rediscovered by whichever gateway currently owns that WHO. Retroactively migrating ownership on a delegation change would be a reasonable follow-up, but it's a separate piece of work from this PR — happy to open an issue for it if you'd like it tracked.

@GreenGrassBlueOcean

Copy link
Copy Markdown
Contributor Author

Updated test build — same one-liner

@anotherjulien, pr-525-test is rebuilt at d2e2fa32 (adds the "Calibrate all covers" fix from your report above, on top of everything you already tested). Same install:

curl -fsSL https://github.com/GreenGrassBlueOcean/MyHOME/releases/download/pr-525-test/install.sh | bash

Backs up your current integration first, then restarts Home Assistant. Release is MyHOME @ d2e2fa32 + OWNd master + OpenWebNet-HA/OWNd#63, same combination as before.

What's new to check:

  1. The MH202's "Calibrate all covers" button should now show unavailable (not present-but-live) once its own covers are gone. You shouldn't need to do anything to trigger this — it re-evaluates from the existing state.

Nothing else changed — A (the topology step) and the rest of B (device/button pruning) are exactly what you already confirmed working. No need to redo those unless you want to.

Rollback is the same as before:

rm -rf /config/custom_components/myhome && cp -a /config/myhome_backup/<the backup dir> /config/custom_components/myhome && ha core restart

@GreenGrassBlueOcean
GreenGrassBlueOcean merged commit 4463e23 into OpenWebNet-HA:v2-phase1-architecture Sep 28, 2026
16 checks passed
@GreenGrassBlueOcean
GreenGrassBlueOcean deleted the fix/shared-bus-onboarding branch September 28, 2026 15:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants