Skip to content

[Bug]: Aborting a directory scan raises OperationCanceledException, and the follow-up rescan is stuck on "INITIALIZING / Walking up the Engine" forever #206

Description

@G00dS0ul

What happened?

Aborting a long directory scan and then pressing refresh to rescan leaves the UI permanently stuck
on INITIALIZING... with the target Walking up the Engine.... and a stale progress bar
carried over from the previous scan (observed: SECTORS SCANNED: 12 / 13, 92.3%). It never
advances and never returns to idle. Separately, the abort itself surfaces an
OperationCanceledException in the debugger.

These are two distinct problems with independent causes. The backend recovers correctly; the UI
does not
, and it cannot, because of four separate defects in the progress contract between the
two sides.


Part 1 — the exception on abort

System.OperationCanceledException: 'The operation was canceled.' breaks at:

core/backend/Engine/DiskScannerEngine.cs:321-323

private long GetDirectorySize(DirectoryInfo dir, CancellationToken token, string scanRoot, int currentDepth = 1)
{
    token.ThrowIfCancellationRequested();

This is a first-chance exception, not a crash — say so plainly so nobody hunts a phantom. The
request is handled end to end:

  • per-item cancellation is caught inside the parallel body at DiskScannerEngine.cs:269-272
    (catch (OperationCanceledException) { return; });
  • the Parallel.ForEachAsync cancellation is caught at :306-311, which logs, emits a CANCELED
    progress message and rethrows;
  • StorageController catches it at core/backend/Controllers/StorageController.cs:101-108 and
    returns HTTP 499 with "Directory Scan Aborted by User.".

The defect is the design, not a missing handler: cancellation is signalled by throwing on a hot,
deeply recursive path. GetDirectorySize recurses per directory and re-checks the token at every
level, and it is invoked once per top-level directory inside a Task.Run
(DiskScannerEngine.cs:257) under a Parallel.ForEachAsync with
MaxDegreeOfParallelism = ProcessorCount / 2 (:245-249). Cancelling a deep tree therefore throws
a burst of exceptions, one per in-flight branch per level. Under the Visual Studio debugger each one
breaks execution, which is what makes a clean abort look like a fault.

Suggested direction: in the recursive walk, return early on token.IsCancellationRequested
instead of calling ThrowIfCancellationRequested(), and let the single outer catch at :306
remain the one place cancellation becomes an exception. Keep the HTTP 499 contract — it is correct
and StorageControllerCancellationTests already exercises this area.


Part 2 — why the rescan sticks on "INITIALIZING"

The backend emits three different ScanProgress shapes, and the frontend mishandles all three. Each
defect below is independently sufficient to freeze the UI.

2a. Progress messages carry no status, and copyWith treats absent as unchanged

The per-directory progress payload has no status field:

core/backend/Engine/DiskScannerEngine.cs:277-284

_ = _hub.Clients.All.SendAsync("ScanProgress", new
{
    scanId = scanId,
    completed = completed,
    total = totalNodes,
    percentageComplete = percentage,
    currentTarget = dir.Name
});

The handler reads data['status'] → null
(core/frontend/gs_analyzer_ui/lib/services/telemetry_service.dart:112), and the reducer preserves
the previous value on null:

core/frontend/gs_analyzer_ui/lib/providers/telemetry_provider.dart:51

status: status ?? this.status,

So once INITIALIZING is set by DiskScannerEngine.cs:243, no message in the progress stream can
ever change it.
The UI is pinned to "INITIALIZING..." for the entire duration of every scan, not
just after an abort. The only way out is a COMPLETED or CANCELED message — and 2c below breaks
CANCELED.

2b. The percentage field name does not match

Backend sends percentageComplete (DiskScannerEngine.cs:282). Frontend reads
percentComplete:

core/frontend/gs_analyzer_ui/lib/services/telemetry_service.dart:115

final percentageComplete = (data['percentComplete'] as num?)?.toDouble();

That is always null, so percentComplete: percentComplete ?? this.percentComplete
(telemetry_provider.dart:54) keeps whatever was last set. This is why the bar freezes at a stale
92.3% — the value is never updated from the progress stream at all.

2c. The backend says CANCELED; the frontend listens for ABORTED

core/backend/Engine/DiskScannerEngine.cs:309

await _hub.Clients.All.SendAsync("ScanProgress", new { scanId = scanId, status = "CANCELED", count = _scannedFilesCount, currentTarget = "Scan canceled" });

core/frontend/gs_analyzer_ui/lib/providers/telemetry_provider.dart:88-91

final isDone =
    status == 'COMPLETED' ||
    status == 'ABORTED' ||
    status == 'FAILED';

CANCELED is not in the list. The backend never sends ABORTED or FAILED on this path — the only
three statuses it emits are INITIALIZING (:243), COMPLETED (:304) and CANCELED (:309).
So an aborted scan is never recognised as finished.

2d. currentScanId can never be cleared, which then drops the next scan's updates

core/frontend/gs_analyzer_ui/lib/providers/telemetry_provider.dart:98

currentScanId: isDone ? null : scanId,

but

core/frontend/gs_analyzer_ui/lib/providers/telemetry_provider.dart:56

currentScanId: currentScanId ?? this.currentScanId,

Passing null to signal "clear this" preserves the old value. The clear-on-finish branch is a
no-op, for COMPLETED just as much as for an abort. That matters because of the acceptance gate:

core/frontend/gs_analyzer_ui/lib/providers/telemetry_provider.dart:85-87

if (status == 'INITIALIZING' ||
    state.currentScanId == null ||
    state.currentScanId == scanId) {

With a permanently non-null, stale currentScanId, messages from any later scan are accepted only
when their status is literally INITIALIZING — and per 2a the progress stream has no status at all.

2e. INITIALIZING carries no counters, so the old ones survive

core/backend/Engine/DiskScannerEngine.cs:243

await _hub.Clients.All.SendAsync("ScanProgress", new { scanId = scanId, status = "INITIALIZING", count = 0, currentTarget = "Walking up the Engine...." });

It sends count, which nothing reads — the frontend reads completed and total
(telemetry_service.dart:113-114). Both are null, so copyWith preserves the previous scan's
12 / 13. This is the exact stale reading in the screenshot, and it is why a fresh scan appears to
start at 92.3%.

Together: a new scan sets INITIALIZING and Walking up the Engine...., keeps the old
counters, then receives progress messages that update completed/total/target but can never
change status or percentComplete, and can never reach a terminal state after an abort. That is
precisely the observed screen.


Part 3 — the queued-scan pile-up

The backend log showed three scans cancelled back to back:

info: GSSystemAnalyzer.Engine.DiskScannerEngine[0]  Scan f602d806-c02f-4e5f-a716-ee541f32edb6 was canceled by user.
info: GSSystemAnalyzer.Engine.DiskScannerEngine[0]  Scan 5784742c-3318-4d06-ac33-1db82cea583f was canceled by user.
info: GSSystemAnalyzer.Engine.DiskScannerEngine[0]  Scan 621c2d21-9d1a-470d-ad69-4868b56355b5 was canceled by user.

Each refresh press starts another scan that blocks on a semaphore with no timeout and no
cancellation token
:

core/backend/Engine/DiskScannerEngine.cs:217

await _scanLock.WaitAsync();

So repeated presses queue up rather than being rejected or coalesced, and a scan already queued
cannot be aborted while it waits — it must first acquire the lock, then immediately throw. Then
one abort press cancels all of them at once, because the UI never sends a scan id:

core/frontend/gs_analyzer_ui/lib/services/api_service.dart:241-248

Future<void> abortScan({String? scanId}) async {
  try {
    appLogger.i('SENDING SCAN ABORT SIGNAL... (scanId: $scanId)');
    var uri = Uri.parse('$storageUrl/abort-scan');
    if (scanId != null) {
      uri = uri.replace(queryParameters: {'scanId': scanId});
    }
    await _client.post(uri);

and the HUD calls it with no argument at
core/frontend/gs_analyzer_ui/lib/widgets/telemetry_hud_widget.dart:87
(await ApiService().abortScan();), so scanId is always null. That reaches the global branch,
which cancels every active session:

core/backend/Engine/DiskScannerEngine.cs:663-670

else
{
    foreach (var session in _activeSessions.Values)
    {
        session.Cts.Cancel();
    }
    _logger.LogInformation("Global scan abort signal received, all active scans canceled");
}

Ruled out — do not chase this. Session cleanup is correct. TriggerScanAbort only cancels and
does not remove sessions (DiskScannerEngine.cs:653-671), but EndScanSession runs in a finally
that covers the cancelled path:

core/backend/Services/DiskOperationsService.cs:170-173

finally
{
    _scanner.EndScanSession(scanId);
}

and _scanLock is likewise released in a finally at DiskScannerEngine.cs:314-317. So neither a
leaked session nor a permanently-held lock explains the stall. The stall is entirely Part 2.

UNVERIFIED, worth checking during implementation: EndScanSession has only two call sites
(DiskOperationsService.cs:172 and core/backend/Services/DuplicateFileDetector.cs:100), yet
BeginScan is also called from core/backend/BackgroundWorkers/ScheduledScanWorker.cs:131 and
core/backend/Controllers/SchedulesController.cs:112. I did not trace whether those two paths
always route through a method that ends the session. If they do not, they leak a ScanSession per
scheduled scan.


Part 4 — a fourth stuck-state path

core/backend/Engine/DiskScannerEngine.cs:226

if (totalNodes > 0)

guards the entire progress-emitting block. directoriesToScan is filtered at :220-222 to
directories missing from the cache. If a rescan finds everything already cached, totalNodes is
0 and no ScanProgress message is sent at all — not INITIALIZING, not COMPLETED. The UI is
left on whatever it last displayed, with no terminal message ever arriving. A fast, fully-cached
rescan is therefore indistinguishable from a hung one.


Steps to reproduce

  1. Start the backend and the Flutter app.
  2. Navigate to the storage/directory screen and scan a large root such as C:/.
  3. While it is running, press the refresh icon two or three more times.
  4. Press ABORT SCAN.
  5. Observe the backend console: several Scan <guid> was canceled by user. lines, and
    POST /api/storage/abort-scan returning 200.
  6. Press refresh again to rescan.
  7. Observe: the panel shows INITIALIZING... / TARGET: Walking up the Engine.... with the
    progress figures from the previous scan, and stays there indefinitely.

Under a debugger attached to the backend, step 4 also breaks on
System.OperationCanceledException at DiskScannerEngine.cs:323.


Expected behaviour

  1. Aborting a scan returns the UI to a terminal state — idle or an explicit "Scan cancelled" —
    within a second, and clears the progress figures.
  2. A rescan starts from a visibly reset progress bar, leaves INITIALIZING as soon as the first
    directory completes, and shows a live percentage.
  3. Aborting does not produce exception breaks in a debugger for an expected user action.
  4. Repeatedly pressing refresh does not queue additional scans; it either coalesces them or the
    in-flight scan is cancelled and replaced.
  5. A fully-cached rescan still emits a terminal COMPLETED.

Area

Backend (/backend, C#/SignalR) — and Frontend (/frontend/gs_analyzer_ui, Flutter). Most of the
user-visible damage is in the Flutter reducer; the exception and the no-message-when-cached path are
backend.

Operating system

Windows

.NET / Flutter versions

.NET 10.0.x (net10.0-windows), Flutter stable — exact patch versions to be confirmed by whoever
files this.

Logs / screenshots

info: Microsoft.AspNetCore.Hosting.Diagnostics[2]
      Request finished HTTP/1.1 POST http://localhost:5200/api/storage/abort-scan - 200 - application/json;+charset=utf-8 1004.8014ms
info: GSSystemAnalyzer.BackgroundWorkers.AutomationWorker[0]
      AutomationWorker: Skipping tick - active scan or nuke in progress
info: GSSystemAnalyzer.Engine.DiskScannerEngine[0]
      Scan f602d806-c02f-4e5f-a716-ee541f32edb6 was canceled by user.
info: GSSystemAnalyzer.Engine.DiskScannerEngine[0]
      Scan 5784742c-3318-4d06-ac33-1db82cea583f was canceled by user.
info: GSSystemAnalyzer.Engine.DiskScannerEngine[0]
      Scan 621c2d21-9d1a-470d-ad69-4868b56355b5 was canceled by user.

Debugger: System.OperationCanceledException: 'The operation was canceled.' —
GSSystemAnalyzer.dll!GSSystemAnalyzer.Engine.DiskScannerEngine.GetDirectorySize(...),
DiskScannerEngine.cs:323.

UI at the point of the stall: INITIALIZING..., SECTORS SCANNED: 12 / 13, 92.3%,
TARGET: Walking up the Engine.....

Pre-flight

  • I searched existing issues and this isn't a duplicate
  • I ran the backend as Administrator
  • The backend was running when I saw this

Proposed direction

Ordered cheapest-and-safest first. 1 to 3 are small and independently shippable; each one alone
visibly improves the symptom.

  1. Align the progress contract. Pick one spelling for the percentage field and one vocabulary
    for terminal statuses. Either rename the backend field to percentComplete or fix the frontend
    read at telemetry_service.dart:115; and either emit ABORTED from DiskScannerEngine.cs:309
    or add CANCELED to the list at telemetry_provider.dart:88-91. Issue Extract SignalR event names into named constants #148 ("Extract SignalR
    event names into named constants", open) is the natural home for preventing recurrence
    — the
    same class of defect, one level up.
  2. Include status in the per-directory payload at DiskScannerEngine.cs:277-284 (e.g.
    status = "SCANNING"), so the UI can leave INITIALIZING. Send completed/total in the
    INITIALIZING payload at :243 too, so the counters reset instead of persisting.
  3. Make the clear-on-finish branch actually clear. copyWith cannot express "set to null" as
    written (telemetry_provider.dart:42-57). Either add an explicit clearScanId flag, use a
    sentinel, or reset by constructing a fresh TelemetryState. Note this affects every copyWith
    field, not only currentScanId.
  4. Stop throwing for expected cancellation in the recursive walk — see Part 1.
  5. Emit a terminal message when totalNodes == 0 — see Part 4.
  6. Decide the refresh-press policy — reject, coalesce, or cancel-and-replace. This one is a
    product decision, not a mechanical fix; it should not be bundled with 1–5.

Constraints any fix must respect

  • Do not change the HTTP 499 contract at StorageController.cs:101-108. It is exercised by
    core/backend/GSSystemAnalyzer.Tests/Controllers/StorageControllerCancellationTests.cs, which
    also pins that a scan must call BeginScan() before returning (:88, :122, :156).
  • ScanProgress is broadcast to Clients.All with no per-client filtering, so the frontend's
    scanId gate is the only thing separating concurrent scans. Any fix to 2d must keep that gate
    working, not remove it.
  • The per-directory send at DiskScannerEngine.cs:277 is fire-and-forget (_ = ...SendAsync),
    unlike the awaited INITIALIZING/COMPLETED/CANCELED sends. Ordering between them is not
    guaranteed. If a fix starts relying on message order, this needs addressing first.
  • IsScanning is derived from the semaphore (DiskScannerEngine.cs:68,
    _scanLock.CurrentCount == 0) and AutomationWorker skips its tick based on it — visible in the
    log above. Changing the locking strategy in step 6 affects scheduled automation, not just the UI.
  • Check docs/notion/unit-test-checklist.md before proposing the tests — it is the merge gate
    for this repo. I have not quoted it here because I did not open it during this investigation.

Proposed acceptance criteria

  • After an abort, the UI reaches a terminal state and the progress figures reset.
  • During a scan the status advances past INITIALIZING and the percentage updates live.
  • A rescan immediately following an abort runs to completion and updates normally.
  • A fully-cached rescan (totalNodes == 0) still produces a terminal COMPLETED in the UI.
  • Aborting produces no OperationCanceledException break in the debugger from
    GetDirectorySize.
  • A regression test pins the backend's emitted status vocabulary against the frontend's
    terminal-status list, so CANCELED/ABORTED cannot drift apart again.
  • A regression test pins the progress field names, so percentageComplete/percentComplete
    cannot drift apart again.
  • Backend suite green, in particular StorageControllerCancellationTests and
    DiskScannerEngineTokenTests; flutter test green for the telemetry provider.

Prior art

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    backendC#bugSomething isn't workingcleanup & refactorCleanup unused method or function and refactoringfrontendFlutter/DartscannerStorage scanner

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions