Skip to content

The open pull request list refreshes again on a busy repo, and says so when it couldn't - #111

Merged
webdevcody merged 2 commits into
mainfrom
fix-106-open-pr-list-stale
Sep 30, 2026
Merged

webdevcody merged 2 commits into
mainfrom
fix-106-open-pr-list-stale

Conversation

@webdevcody

Copy link
Copy Markdown
Contributor

Closes #106. On a busy repo the open pull request list stopped updating, because gh pr list asked GitHub for every CI check on every pull request and GitHub timed out every time. The list now asks for GitHub's single pass / fail / pending verdict per pull request, which comes back in about 2 seconds. When a refresh does fail, the PULL REQUESTS MODAL now says so.

Contents: 🐛 Symptom · 🔍 Cause · ✅ Fix · 📸 Before / After · 🔁 State · ⚠️ Risk · 🔧 Technical overview · 🧪 Proof · 📝 Notes

🐛 Symptom

The report came from a private repo with 83 open pull requests, averaging 31 checks each and up to 58. On nebula v0.40.2, v listed 72 pull requests. The 15 newest never showed up, and 4 that had merged or closed stayed listed. pr-cache/pull-requests.json got a fresh saved_at every few minutes, but its list never changed, and nothing on screen said it was stale.

🔍 Cause

The list ran gh pr list --json …,statusCheckRollup. gh fills that field with every check context on each pull request's head commit, which on that repo is thousands of entries. GitHub's GraphQL API answered with HTTP 504 after about 11 s on every call. nebula treats a failed lookup as a one-off: it keeps the last list, retires nothing, and backs off for up to 10 minutes. That is right for a flaky network, but this call never succeeded, so the list stayed on its last good answer indefinitely. The failure was silent because nothing recorded that the last ask had failed.

✅ Fix

  • One verdict per pull request, not every check. The list is now a single GraphQL query.
    • It asks for statusCheckRollup { state } on each pull request's last commit. That is the ✓ / ✗ / pending GitHub already shows on the pull request page.
    • The rows, their order (newest first) and the 100-row cap are the same as gh pr list gave.
    • Row badges still turn red: failing for failed checks, conflicts for merge conflicts.
  • A list that couldn't refresh says so. The PULL REQUESTS MODAL (v) shows couldn't refresh (^r retries) in the warn color on its own row under the filter.
    • The rows it kept stay listed and clickable underneath.
    • With no rows to keep, it says couldn't ask GitHub (^r retries) instead of the misleading no open pull requests.
    • While a retry is running, the title says refreshing… as before. The next answer clears the notice.
  • Unchanged.
    • A failed call still keeps the last list and backs off; one bad network round trip never blanks the rows.
    • The checkout's own pull request (gh pr view) and the reading pane's detail fetch still fold every check. Each asks about a single pull request, which never timed out.
    • The PR CACHE format is unchanged.

📸 Before / After

The same scene on both builds. The list answers once at boot, then every later ask fails the way GitHub's 504 did, and Ctrl+r asks again inside the modal. The footer's refreshing pull requests… is the Ctrl+r press's own flash in both shots.

Before (#106) — v0.41.0 After
The PULL REQUESTS MODAL on v0.41.0 after a failed refresh: three rows and no sign they are stale The same modal with this change: "couldn't refresh (^r retries)" under the filter, over the same three rows with their conflicts and failing badges

🔁 State

flowchart TD
  T["OPEN PRS beat / Ctrl+r"] --> Q{"which call?"}
  Q -- before --> OLD["gh pr list --json …,statusCheckRollup<br/>every check on every PR"]
  OLD -- "HTTP 504 after ~11 s, every time" --> MISS["answer: None"]
  MISS --> KEEP["keep last list, back off ≤ 10 min"]
  KEEP --> SILENT["modal looks current"]
  Q -- after --> NEW["gh api graphql LIST_QUERY<br/>statusCheckRollup { state }"]
  NEW -- "~2 s" --> ROWS["rows + health"]
  ROWS --> CLEAR["open_prs_failed cleared"]
  NEW -- "still fails (no gh, offline)" --> MARK["open_prs_failed marked"]
  MARK --> NOTE["modal: couldn't refresh (^r retries)"]
  classDef bad fill:#fdd,stroke:#c33,color:#000
  classDef good fill:#dfd,stroke:#3a3,color:#000
  class OLD,SILENT bad
  class NEW,NOTE good
Loading

⚠️ Risk

Verdict: 🟢 Low risk. The change only swaps which gh subcommand runs, uses a constant query, and adds one warning row to the modal.

Level Why
🔒 Security & production Low No new surface: the same gh the TUI already runs, with a constant query. {owner} / {repo} are filled in by gh from the checkout, the same way gh pr list resolves its repo, and no user text reaches the call. A repo gh can't resolve, or a GraphQL error, exits non-zero and reads as "couldn't ask", as before.
⚡ Performance Low Off every hot path: the call runs on a background task. It is much cheaper than before (reported 2.0–2.5 s vs a 504 at 11 s; 3.0 s for 62 rows on cli/cli). The draw adds one HashSet lookup.
🧩 Fit with the codebase Low open_prs_failed follows pr_detail_failed's pattern: a set of keys whose last ask failed, cleared by the next answer and pruned with the projects. The call goes through the same gh() helper and TIMEOUT, and the notice row is drawn with the modal's existing row helpers.

Rollback: git revert the merge brings gh pr list back. Nothing persisted changes shape (the OpenPr rows and the PR CACHE are untouched). The screenshots on pr-assets stay.

🔧 Technical overview

  • The line that mattered. crates/nebula-tui/src/pull_request.rs:
    • list now runs gh api graphql -F owner={owner} -F repo={repo} -F limit=100 -f query=LIST_QUERY. The query has gh pr list's own fields plus commits(last: 1) { … statusCheckRollup { state } }.
    • parse_list reads rows from data.repository.pullRequests.nodes.
    • health reads the rollup state when a node carries it (SUCCESS passes, FAILURE / ERROR fail, PENDING / EXPECTED are pending, null is absent). Otherwise it falls back to the old per-check fold, which gh pr view and the detail payload still use.
  • Remembering the failure.
    • crates/nebula-tui/src/event_loop.rs: note_open_prs_answer marks or clears App::open_prs_failed (declared in crates/nebula-tui/src/app.rs). Removed projects are pruned from it the same way as open_prs.
    • crates/nebula-tui/src/pr_modal.rs draws the notice row and shifts the rows' hit area under it, so clicks still land on the right row.
  • The screenshot harness. scripts/shot/bin/gh answers the GraphQL call from the same pr-list.json fixtures, converting each row's checks into a rollup state, so every existing PR scene renders as before. pr-list.ok-count makes it fail after N answers, and the new pr-list-stale scene uses that.
  • Why the rollup state and not a per-check fold. It is the verdict GitHub itself shows. Before relying on it, I compared it with the old fold on all 62 open pull requests on cli/cli: 0 disagreements.
  • Why not keep the per-check list and fetch health in a second call (the issue's option 2). The issue's own timings put the rollup query at 2.0–2.5 s against 1.5 s with no checks at all, so a second process and a merge-by-number step would buy about half a second.
  • Why a row, not a title suffix (the issue's option 3). At 100 columns the list column is 32 cells wide, and the title already cuts refreshing… to refres. The one thing this has to say must not be the part that gets clipped.

🧪 Proof

  • Regression tests (all new in this PR):
    • a_list_rows_checks_are_the_rollup_state and the_list_query_asks_for_the_rollup_state_alone in crates/nebula-tui/src/pull_request.rs. The second pins that the query asks for no check contexts.
    • a_list_that_could_not_be_refreshed_says_so in crates/nebula-tui/src/pr_modal.rs: the notice, the kept rows, the shifted hit area, refreshing… during a retry, couldn't ask GitHub with no rows, and the notice clearing.
    • the_open_pr_list_backs_off_when_empty_and_survives_a_failed_call in crates/nebula-tui/src/event_loop.rs now also checks that the failure mark is set and then cleared.
    • The existing list-parsing tests take the GraphQL answer shape.
  • Live. The real list() against cli/cli returned 62 rows in 3.0 s, with Passing 41 · Failing 5 · Absent 16 and 30 conflicting. The reporter's private repo was not available, so the 83-PR case is covered by the reporter's own timing of this query (all 83 in 2.0–2.5 s), not by a run here.
  • Gate. Run on origin/main (5938529) plus this change:
    • cargo fmt --check and make lint (clippy) are green.
    • cargo test --workspace --no-fail-fast: 1685 passed, 3 failed.
    • Two of the failures are the e2e_tui tests that also fail on untouched main (nebula_open_from_inside_a_session_raises_the_file_tabs, tui_drag_past_the_pane_top_autoscrolls_and_copies_the_run; their start_session waits on an old footer).
    • The third, branch_switch::tests::stash_and_switch_leaves_the_changes_in_a_named_stash, hit a transient run git: No such file or directory and passed 3 of 3 on rerun. It is in code this PR doesn't touch.

📝 Notes

  • There are two commits: the fix, and the SCREENSHOT HARNESS stub and scene that the fix makes necessary.
  • The branch starts at 331e836, which is already in main. It merges cleanly onto origin/main, and that merged tree is what the gate ran on.

🤖 Generated with Claude Code

webdevcody and others added 2 commits September 30, 2026 01:37
… request, so a busy repo's list refreshes again, and a list that couldn't refresh says so

- The list was `gh pr list --json ...,statusCheckRollup`, which asks for
  every check on every pull request's head commit. On a repo with 83 open
  pull requests and 30 to 60 checks each, GitHub answers that with an
  HTTP 504 every time. The call failed on every beat, the last good list
  stayed on screen, new pull requests never showed and merged ones never
  left.
- pull_request::list now runs one `gh api graphql` query with
  `gh pr list`'s own fields and order, and reads checks from the rollup's
  `state` on the last commit. That is GitHub's own pass / fail / pending
  word, computed server side. On cli/cli it matched the old per-check
  fold on all 62 open pull requests. `gh pr view` and the detail fetch
  still fold every check, since those ask about one pull request.
- A failed list lookup is now remembered per project
  (App::open_prs_failed). The PULL REQUESTS MODAL shows
  "couldn't refresh (^r retries)" on its own row under the filter, over
  the last rows that worked, or "couldn't ask GitHub" when there are
  none. It is a row, not a title suffix, so a narrow modal can't cut it
  off. The next answer clears it.
- Tests cover the GraphQL answer shape, the rollup word mapping, the
  query asking for no check contexts, the failure mark being set and
  cleared, and the modal notice with its hit area.
- Docs (how-it-works, keys, configuration) describe the query and the
  notice.

Closes #106

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…query, and a pr-list-stale scene shows a list that couldn't refresh

- nebula now asks for the open pull request list with `gh api graphql`
  instead of `gh pr list`. Without this, every PR scene's stand-in gh
  would fail that call and the grid and modal would render
  "couldn't ask GitHub".
- The stub hands back fixtures/pr-list.json, still in `gh pr list` rows,
  as the GraphQL answer. A row's statusCheckRollup checks become the
  rollup state on its last commit, so the existing fixtures and scenes
  keep working unchanged.
- fixtures/pr-list.ok-count holding N answers the first N list calls and
  fails every later one the way GitHub's 504 did.
- New scene pr-list-stale reuses pr-row-conflicts's fixtures with an
  ok-count of 1: `v`, then Ctrl+r, shows the modal saying
  "couldn't refresh (^r retries)" over the rows it kept.

Refs #106

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@webdevcody
webdevcody merged commit a8ac713 into main Sep 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Open PR list goes stale when gh pr list with statusCheckRollup times out on a busy repo

1 participant