Skip to content

Search: long-running IndexSpace traversal reuses an expired service token #3541

Description

@pntone

Search: long-running IndexSpace traversal reuses an expired service token

Describe the bug

A full-space reindex can outlive the service token acquired at the beginning of IndexSpace. The same authentication context is then reused for later gateway operations, causing the traversal to fail instead of refreshing its credentials.

In our original deployment, a force-rescan failed after 86,424 seconds (24 hours and 24 seconds) with PermissionDenied reporting an invalid core access token. OpenCloud and Tika remained healthy, with no container restarts or OOM events. Index writes and extraction activity stopped at that failure.

Setup and verification scope

  • Originally observed on OpenCloud 7.4.0, using Bleve and Tika for a large document archive. That build included a separate Bleve/zapx fix, but not this authentication refresh fix.
  • The original deployment ran in rootless Podman containers managed by systemd/Quadlet. The current deployment uses rootful Podman/Quadlet.
  • The original incident and diagnosis were recorded on 2026-08-19.
  • The same one-time authentication pattern is still present in the v8.0.0 source, checked on 2026-09-16.
  • Our current deployment uses OpenCloud 8.0.0, OpenSearch and Tika OCR with a local workaround. We have not repeated the full 24-hour failure on an unpatched v8.0.0 deployment.
  • This concerns the internal service/core token used by indexing, not the user's browser session or OIDC refresh token.

Steps to reproduce

  1. Use an unpatched search service and a space whose full traversal/content extraction takes longer than the internal service token lifetime.
  2. Start a forced reindex of that space with concurrency 1.
  3. Keep the indexing request alive beyond the token lifetime. A separate client-side timeout must not terminate the request first (see below).
  4. Allow a subsequent authenticated gateway operation to occur after the original token expires.

In the observed configuration, the token lifetime was 86,400 seconds. The token manager's vendored default is also 86,400 seconds.

For a faster regression test, use an injected clock and an authentication stub with distinguishable tokens. Advance the clock between the walker's initial Stat and a later ListContainer, and assert that the later request uses refreshed credentials.

Expected behavior

Indexing should refresh its internal authentication before expiry and continue traversing the space without requiring a restart or a global increase of token lifetime.

Actual behavior

The long-running traversal reused its initial service context and failed with PermissionDenied once the core token expired. This left the full reindex incomplete.

The elapsed time and error above come from our recorded incident report; they are not presented as a verbatim raw log excerpt.

Source analysis

In v8.0.0 services/search/pkg/search/service.go, IndexSpace:

  1. Calls getAuthContext(...) once before walking the tree (line 464).
  2. Passes that context into w.Walk(...) (line 505).
  3. Reuses it for s.engine.Search(...) in the callback (line 532).

There is no refresh in this traversal path. This explains why a healthy service can fail partway through a sufficiently long reindex.

Local workaround and regression test

We have a small local patch that caches an indexing-specific authentication context and refreshes it every 12 hours. It supplies the current context to the walker's Stat and ListContainer operations and the per-resource index lookup. It leaves the global token lifetime unchanged.

The accompanying test, TestIndexSpaceWalkerRefreshesServiceTokenDuringLongWalk, uses the real walker, mocked gateway calls and a controlled clock. It verifies that traversal crosses the refresh boundary and subsequent gateway calls receive the second token. The test passes on our patched v8.0.0 build.

The 12-hour interval is a workaround for our 24-hour token lifetime, not a proposed universal default. An upstream fix should account for the actual expiry/configuration, preserve cancellation and deadlines when refreshing authentication metadata, and propagate authentication-refresh failures.

I can provide the patch and regression test for review.

Separate CLI timeout issue

The v8.0.0 CLI also has a hard-coded 10-minute context timeout. That can prevent a long-running CLI reproduction from reaching token expiry. It is a separate issue: DeadlineExceeded after 10 minutes is not evidence of an expired token.

For our current background job, we added a local configurable --timeout option and disabled that deadline. This report is specifically about refreshing authentication during IndexSpace.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions