Skip to content

[Bug]: Thermal tiers 2 and 3 run in the wrong order, and a non-elevated backend floods the console with Access denied stack traces #205

Description

@G00dS0ul

What happened?

Two defects, both visible the moment the backend starts without Administrator rights.

1. The tier order in code does not match the documented fallback chain.

beta-technical-spec defines the chain as:

  • Tier 1 — LibreHardwareMonitorLib (in-process, broadest coverage): CPU load, per-core temps, clocks, package power, some fans, drive temps.
  • Tier 2 — WMI MSAcpi_ThermalZoneTemperature: ACPI thermal-zone / ambient temperature where LHM is blank.
  • Tier 3 — OEM WMI bridges: vendor namespaces for fan RPM and OEM temps — Dell root\dcim\sysman (DCIM_NumericSensor), HP root\hp\instrumentedbios, Lenovo ACPI VPC.

The code runs Tier 1 → Dell OEM → WMI, i.e. tiers 2 and 3 are transposed. Dell OEM is queried
before the ACPI thermal zone, not after it.

The codebase contradicts itself about this. The Dell provider's own summary is correct:

core/backend/Services/Oem/Dell/DellOemTelemetry.cs:9

/// Tier-3 fallback: reads fan RPM from Dell Command | Monitor's WMI provider

But the call site labels it Tier 2 and labels WMI Tier 3:

core/backend/Services/LibreThermalProvider.cs:196-197

// Tier 2: DELL VBS
var dell = _dellOemTelemetry.TryGetDellOemTelemetry();

core/backend/Services/LibreThermalProvider.cs:218-222

// Tier 3: Standard WMI Fallback
if (payload.CpuPackageCelsius == 0 || payload.CpuPackageCelsius == null)
{
    payload.CpuPackageCelsius = _wmiFallback.GetCpuTemperatureCelsius();
}

This is not only a comment problem — it is the execution order. prd.md:164 (FR-S4) also places
WMI immediately after LibreHardwareMonitor:

FR-S4: Provide a WMI ACPI fallback (GetAcpiTempFallback) when LibreHardwareMonitor sensors are unavailable.

2. The Dell tier is unconditional, and neither failing tier backs off.

LibreThermalProvider.cs:197 calls the Dell provider on every poll regardless of whether
LibreHardwareMonitor already produced a reading. Only the WMI tier at :219 is guarded by a
"do we still need this?" check. So on a machine that is not a Dell — or any machine where the
backend is not elevated — the Dell query runs forever, fails forever, and logs a full stack trace
every time.

Both failing tiers log at Error level with the exception attached:

core/backend/Services/Oem/Dell/DellOemTelemetry.cs:93-97

catch (Exception ex)
{
    _logger.LogError(ex, "Dell OEM telemetry query failed");
    return null;
}

core/backend/Services/WmiThermalFallback.cs:33-36

catch (Exception ex)
{
    _logger.LogError(ex, "WMI thermal fallback query failed");
}

The Dell provider has a cache, but it only caches success. The 3-second throttle at
DellOemTelemetry.cs:28 can never engage on a failing machine, because _lastPollTime and
_cachedReading are assigned only inside the success branch:

core/backend/Services/Oem/Dell/DellOemTelemetry.cs:85-91

if (found)
{
    _cachedReading = reading;
    _lastPollTime = DateTime.UtcNow;
    return _cachedReading;
}
return null;

WmiThermalFallback has no caching or back-off of any kind.

Rate. The thermal loop polls every ThermalPollIntervalMs, which defaults to 2000 ms
(core/backend/Models/SettingDtos/MonitoringSettingDto.cs:7), validated to a 500–60000 ms range
(core/backend/Models/SettingDtos/AppSettingDto.cs:21). The loop is
core/backend/Engine/ThermalMonitoringEngine.cs:43 inside a while that delays by _pollInterval
at :64. That is two stack traces every 2 seconds — roughly 3,600 per hour — which is what
makes the console unusable.

There is no vendor gate. The factory that was meant to select a vendor-specific reader is
entirely commented out — core/backend/Services/Oem/Factory/OemReaderFactory.cs:6-19 is a single
block comment containing the whole class. Nothing checks the machine's manufacturer before querying
the Dell namespace.

All three providers are registered as singletons, so the state needed for a one-shot availability
probe already has somewhere to live:

core/backend/Program.cs:89-91

builder.Services.AddSingleton<IThermalProvider, LibreThermalProvider>();
builder.Services.AddSingleton<IWmiThermalFallback, WmiThermalFallback>();
builder.Services.AddSingleton<IDellOemTelemetry, DellOemTelemetry>(); // User needs to have Dell OEM telemetry installed for this to work

Steps to reproduce

  1. On a Windows machine, open a non-elevated terminal (this is the key difference from [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181 —
    no special hardware or missing software is required).
  2. Run the backend: dotnet run --project backend
  3. Wait about 5 seconds for the thermal loop to start
    (ThermalMonitoringEngine.cs:33 delays 5000 ms before its first poll).
  4. Watch the console.

Every 2 seconds, two full System.Management.ManagementException: Access denied stack traces
appear — one from DellOemTelemetry.TryGetDellOemTelemetry(), one from
WmiThermalFallback.GetCpuTemperatureCelsius() — and they continue for the life of the process.


Expected behaviour

  1. Run the tiers in the documented order: LibreHardwareMonitor → WMI ACPI thermal zone →
    OEM/Dell, matching beta-technical.md:54-56. Correct the two call-site comments at
    LibreThermalProvider.cs:196 and :218 to match, so the file stops contradicting
    DellOemTelemetry.cs:9.
  2. Gate each tier on need, the way the WMI tier already is at LibreThermalProvider.cs:219 —
    do not query a lower tier when a higher one already supplied the value.
  3. Probe once, then disable for the session. On a deterministic, non-transient failure, set an
    availability flag and return early on every later call. The failure must be logged once, at
    Debug level, as a short message with no stack trace.
  4. Treat Access denied as non-transient, alongside Invalid class and Invalid namespace.
    Elevation cannot change mid-process, so retrying it is pointless. This is the specific gap that
    [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181's proposed fix leaves open — see below.
  5. Degrade visibly rather than silently. When every tier fails, the user should be told that
    thermal data needs an elevated backend, rather than being shown an empty panel and a console
    full of stack traces.

Area

Backend (/backend, C#/SignalR)

Operating system

Windows

.NET / Flutter versions

.NET 10.0.x (net10.0-windows), Flutter stable — exact patch versions to be confirmed by whoever
files this.

Logs / screenshots

fail: GSSystemAnalyzer.Services.Oem.Dell.DellOemTelemetry[0]
      Dell OEM telemetry query failed
      System.Management.ManagementException: Access denied
         at System.Management.ManagementException.ThrowWithExtendedInfo(ManagementStatus errorCode)
         at System.Management.ManagementScope.InitializeGuts(Object o)
         at System.Management.ManagementScope.Initialize()
         at System.Management.ManagementObjectSearcher.Initialize()
         at System.Management.ManagementObjectSearcher.Get()
         at GSSystemAnalyzer.Services.Oem.Dell.DellOemTelemetry.TryGetDellOemTelemetry() in ...\core\backend\Services\Oem\Dell\DellOemTelemetry.cs:line 35

fail: GSSystemAnalyzer.Services.WmiThermalFallback[0]
      WMI thermal fallback query failed
      System.Management.ManagementException: Access denied
         at System.Management.ManagementException.ThrowWithExtendedInfo(ManagementStatus errorCode)
         at System.Management.ManagementObjectCollection.ManagementObjectEnumerator.MoveNext()
         at GSSystemAnalyzer.Services.WmiThermalFallback.GetCpuTemperatureCelsius() in ...\core\backend\Services\WmiThermalFallback.cs:line 21

(both repeat every 2 seconds, indefinitely)

Note the reported line numbers differ slightly from the using statements cited above because the
throw surfaces on the enumeration/initialisation line rather than the declaration:
DellOemTelemetry.cs:35 is using var results = searcher.Get(); and WmiThermalFallback.cs:21 is
the foreach over searcher.Get().

Pre-flight


Constraints any fix must respect

  • Administrator elevation is a documented requirement, not a bug.
    docs/notion/techspec-v2.md:572 states: "All Windows sensor reads use LibreHardwareMonitor
    (LibreHardwareMonitorLib NuGet). Do NOT use raw WMI for sensor data — WMI does not expose fan
    RPM or per-core temps on most consumer hardware. … Requires the backend process to run with
    Administrator privileges."
    techspec-v2.md:316 repeats it. So the fix is not to make sensors
    work unelevated — it is to fail quietly and tell the user why.
  • prd.md:400 records this exact question as still open on the product side:
    "What is the fallback UX when both LibreHardwareMonitor and ACPI fail? | Product — degraded-mode
    design"
    . Point 5 of Expected behaviour above needs a product decision before it can be built; the
    first four points do not.
  • prd.md:161 (FR-S1) — "Read thermal and fan sensors via LibreHardwareMonitorLib."
    LibreHardwareMonitor stays Tier 1. Reordering must not demote it.
  • Do not replace LHM with WMI for sensor data — techspec-v2.md:572 forbids it explicitly.
  • Existing tests touch this area. core/backend/GSSystemAnalyzer.Tests/Engine/LibreThermalProviderTests.cs
    constructs LibreThermalProvider with mocked IWmiThermalFallback and IDellOemTelemetry
    throughout (:45-49, :75, :99, :120, :146, :185, :227-229, :249-252). Reordering
    the tiers will change what those mocks are expected to receive.
    core/backend/GSSystemAnalyzer.Tests/Services/DellOemTelemetryTests.cs also exists.
    I did not determine whether any existing test asserts tier order — that needs checking
    before the change, because if none does, nothing will catch a future re-transposition.

Proposed acceptance criteria

  • Tier execution order is LibreHardwareMonitor → WMI ACPI → OEM/Dell, matching
    beta-technical.md:54-56, with a test that pins the order.
  • Each tier is skipped when a higher tier already supplied the value.
  • A non-elevated backend logs the unavailability of each tier once, at Debug level, with no
    stack trace, and never retries it for the life of the process.
  • Access denied is treated as non-transient alongside Invalid class / Invalid namespace.
  • Console output from a 10-minute non-elevated run contains no repeated sensor stack traces.
  • On an elevated Dell machine, fan RPM and OEM temps still populate as before — no regression of
    the Tier-3 data itself.
  • Backend test suite green, in particular LibreThermalProviderTests and DellOemTelemetryTests.

Relationship to #181 — read before filing

#181 (open, labelled bug, backend, cleanup & refactor, created 2026-08-07) is titled
"OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with
stack traces"
. It already covers the retry-and-log-noise half of this report, and its proposed fix
(probe once, cache an _isAvailable flag, log once at Debug, ideally hoisted onto a shared
ISensorProvider) is the right shape.

What #181 does NOT cover:

  1. The tier-order defect. [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181 does not mention ordering at all. It states the opposite of a
    problem: "The tiered fallback itself works correctly: the chain degrades to the WMI provider and
    the Thermal Radar panel populates normally."
    The transposition against beta-technical.md:54-56
    is genuinely new and is the substantive half of this report.
  2. The Access denied trigger. [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181 is written entirely around
    ManagementException: **Invalid class** (0x80041010) on machines without Dell Command | Monitor.
    Its fix list is explicit: "On ManagementException with Invalid class or Invalid namespace,
    set an _isAvailable = false flag"
    . Access denied is not in that list, so [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181 as written
    would ship and the non-elevated flood would continue unchanged.
  3. The WMI tier flooding too. [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181 concerns only the Dell/OEM provider and assumes WMI succeeds.
    In the non-elevated case WmiThermalFallback fails identically and has no back-off at all
    (WmiThermalFallback.cs:33-36).
  4. [Bug]: OEM telemetry tier retries permanently-unavailable WMI class on every poll, flooding logs with stack traces #181's pre-flight says the backend was run as Administrator, so it cannot have observed this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    backendC#bugSomething isn't workingcleanup & refactorCleanup unused method or function and refactoringgood first issueGood for newcomers

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions