Skip to content

SLURM job folders with per-attempt subfolders; SLURM results marked in History - #127

Merged
NCCU-Schultz-Lab merged 9 commits into
mainfrom
claude/output-naming-strategy-cvs33t
Sep 26, 2026
Merged

NCCU-Schultz-Lab merged 9 commits into
mainfrom
claude/output-naming-strategy-cvs33t

Conversation

@NCCU-Schultz-Lab

Copy link
Copy Markdown
Collaborator

Summary

SLURM job output is now laid out the way sbatch users expect. A failed job can be rerun in the same folder without overwriting anything, and every successful SLURM run still shows up in History, marked as a SLURM result. Local (in-kernel) runs are unchanged.

  • Job folders (JD.2)
    • Each submission gets a readable folder under the job root, named <formula>_<calc>_<method>_<basis>, e.g. H2O_opt_B3LYP_def2-SVP.
    • The same name is used as the SLURM --job-name, so squeue output matches the folder.
    • An optional Job name field on the Calculate tab (and quantui submit --job-name) overrides the default. A name already in use gets _2, _3, and so on.
  • Attempt subfolders (JD.3)
    • submit.slurm creates attempt-NN_job<SLURM id>/ every time it runs. The number is the highest existing one plus one, chosen under flock, and the script uses mkdir without -p, so it aborts rather than write into a folder that already exists.
    • The worker writes its outputs there, using a new --attempt-dir argument.
    • SLURM's own slurm-%j.out/.err files stay in the job folder, because SLURM opens them before the script runs and will not create a missing folder.
  • Resubmit (JD.4)
    • A new button on the Cluster Jobs tab runs the same submit.slurm again as a new attempt, then monitors it.
    • Checkpoints let the new attempt resume where the failed one stopped.
  • Hand-run reruns (JD.11)
    • A hand-run sbatch submit.slurm also creates a new attempt.
    • The worker updates the registry record only when its SLURM_JOB_ID is the job the record tracks, so a hand-run job cannot overwrite the record's status.
    • Pressing Refresh on the Cluster Jobs tab saves finished attempts that are not yet in History, hand-run ones included. Failed attempts stay out of History.
  • History (JD.5/JD.6)
    • Ingested results get additive execution_backend / slurm fields in result.json; no schema bump is needed.
    • History entries show 🖥 SLURM <job id>·a<attempt>, and the result card has a "Ran on" row showing the folder.
  • Job root (JD.1)
    • New System Settings field, compute.slurm_job_root. QUANTUI_STAGING_DIR still overrides it and locks the field.
    • When the job root is outside $HOME, Apptainer runs bind it automatically.
  • Shared files (JD.7)
    • result.molden and trajectory.xyz in attempt folders take the job name.
    • Ingest maps them back to the fixed names History expects.
    • result.json is now written atomically.
  • Fixes
    • Reconnecting to a finished job no longer saves a duplicate History entry. This bug predates the PR.
    • Removed the unused QUANTUI_RESULTS_DIR export from generated scripts (JD.9).
  • Back-compat (JD.8): records created before this change keep the staging/<request_id>/ layout and still list, reconnect and ingest. Resubmit on one of them explains why it can't rerun in place.

Testing

  • New tests/test_jobdirs.py:
    • Runs the attempt-setup block through real bash.
    • End to end: generates submit.slurm and runs it twice by hand, then checks that the first attempt's files are untouched and the registry record is unchanged.
    • Also covers naming, Resubmit, History provenance, hand-run ingest and deduplication, the job-root setting, and the job-name field.
  • Updated two existing tests that pinned the old worker command, and one stub dispatch signature.
  • pre-commit run (ruff + black) passes on all changed files.
  • Full pytest -m "not network" in a cloud container: 3402 passed, 18 skipped, 23 failed.
    • All 23 failures are NMR/Raman tests. They fail identically on main in that container because pyscf-properties would not build there, so they are environmental and unrelated to this branch.
    • CI installs the full [pyscf] extra, so CI is the real check for them.
  • Not yet verified on a real cluster. The NCShare manual pass (JD.10) is tracked in the planning repo.

Authorship

  • Claude (Opus 5.5): code edits, review, and conceptual discussion
  • Jonathan Schultz: overall vision, planning, review, and orchestration

🤖 Generated with Claude Code


Generated by Claude Code

NCCU-Schultz-Lab and others added 9 commits September 26, 2026 04:53
…late

The template exported QUANTUI_RESULTS_DIR=<staging>/results and created that
directory, but the batch worker never calls save_result() or list_results();
results reach History only when the app ingests the staging dir after the
job ends, using the app's own results dir. The export only left an empty
results/ folder in every staging dir. (M-JOBDIRS JD.9)

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.2/JD.3/JD.4 (backend)/JD.8.

- Each SLURM submission gets a readable job dir under the staging root,
  named <label>_<calc>_<method>_<basis> (or a sanitized --job-name), with
  _2, _3... for a new calculation that reuses a name. The same name is the
  SLURM --job-name (truncated to 40 chars) so squeue matches the dir.
- submit.slurm now creates its own attempt-NN_job<SLURM id>/ dir (numbered
  max+1 under a flock, mkdir without -p) and the worker writes into it via
  a new --attempt-dir flag. A hand-run `sbatch submit.slurm` therefore gets
  a fresh attempt and never touches an earlier attempt's files. SLURM's own
  slurm-%j.out/.err stay in the job dir because SLURM will not create a
  missing directory for them.
- JobRecord gains job_dir, attempts and ingested_attempts; staging_path
  resolves to the tracked job's attempt dir, so the monitor, ingest and
  terminal-state checks read the right attempt unchanged.
- SlurmBackend.resubmit() runs the job's submit.slurm again as a new
  attempt and points the record at the new SLURM job.
- The worker only updates the registry record when its SLURM_JOB_ID is the
  tracked one, so a hand-run attempt cannot flip the record's status.
- Apptainer runs bind the job dir when it lives outside $HOME.
- Legacy records (no job_dir) keep staging_root/<request_id>/ and behave as
  before.

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.5/JD.6.

- Ingest now passes provenance extras to save_result on all three paths
  (generic, frequency, reorganization energy): execution_backend="slurm"
  plus slurm.{job_id, attempt, job_dir, attempt_dir, request_id}. For a
  job-dir record the attempt number and job id come from the attempt dir
  name, so a hand-run attempt is labelled with the job that produced it.
- History list labels get a "🖥 SLURM <job>·a<attempt>" marker, next to
  the existing calibration marker.
- The saved-result card gets a "Ran on" row: SLURM batch, job id, attempt,
  and the attempt (or job) folder path.
- Additive: no schema bump. Local results and results ingested before this
  change have no provenance keys and render exactly as before.

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.11.

- A Cluster Jobs refresh now saves every finished attempt that is not yet
  in History: hand-run `sbatch submit.slurm` attempts (which the app never
  monitored) and tracked successes nobody reconnected to. The attempt the
  app is currently monitoring, or whose tracked job is still active, is
  left alone. result.json is written only on success, so failed attempts
  stay out of History.
- JobRecord.ingested_attempts records what has been saved; ingest_attempt()
  saves and records in one place. A hand-run attempt does not retarget the
  record's result_dir.
- Reconnecting (View progress) to a finished job whose result is already in
  History now shows it without saving a duplicate entry. Legacy records
  are covered too: a result_dir that already points at a History entry
  counts as ingested (this duplicate-on-reconnect predates M-JOBDIRS).

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.4 (UI).

Resubmit runs the selected finished or failed job's submit.slurm again as
a new attempt in the same job folder (SlurmBackend.resubmit), then opens
the Calculate tab monitoring the new attempt, the same as View progress.
Earlier attempts' files are kept, and checkpoints let the new attempt pick
up where the failed one stopped. It is refused while another calculation
is running or monitored, for jobs still active, and for jobs submitted
before per-job folders existed (with a message saying why).

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.1.

- New compute.slurm_job_root user setting (blank = ~/.quantui/staging).
  Precedence: QUANTUI_STAGING_DIR > setting > default, resolved in
  cluster_config.default_staging_root(), so the app, the CLI and the batch
  worker agree.
- System Settings shows a "SLURM job folder" field when SLURM is
  available. It accepts an absolute path (created if missing), shows an
  error for relative or unusable paths, and is locked with a note when
  QUANTUI_STAGING_DIR is set. A change applies to new submissions; jobs
  already in the registry keep their absolute folder paths.

Useful on clusters with small home quotas: point it at scratch or a
project directory (Apptainer runs bind it automatically, see the job-dir
commit).

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.2 (UI).

In SLURM mode the run panel shows an optional "Job name" field. It names
the job folder and the SLURM --job-name (sanitized: letters, digits, - and
_ kept, other characters become _). Blank keeps the default
<formula>_<calc>_<method>_<basis>. A name already in use gets _2, _3...,
and the field clears after a successful submit so the next job does not
reuse it by accident. The row is hidden in Local mode.

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.7.

- In an attempt folder the worker renames the files people copy out,
  result.molden and trajectory.xyz, to <job name>.molden and
  <job name>_trajectory.xyz, and records the mapping in
  result.json["artifact_names"]. Files the app reads back (result.json,
  orbitals.npz, progress.json, live.log) keep fixed names.
- Ingest maps the descriptive names back, so History entries keep the
  fixed names their loaders expect.
- write_worker_result is now atomic (tmp file + os.replace): a finished
  result.json marks a successful attempt, and a Cluster Jobs refresh may
  scan for it while the worker is still writing.
- Legacy staging runs (no --attempt-dir) are unchanged.

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS docs: job folder layout and rerun behavior in
apptainer/slurm/README.md (plus the QUANTUI_STAGING_DIR row and the
worker's --attempt-dir flag), the `quantui submit --job-name` help text,
and CHANGELOG entries under Unreleased.

Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@NCCU-Schultz-Lab
NCCU-Schultz-Lab merged commit a8a8ee5 into main Sep 26, 2026
6 checks passed
@NCCU-Schultz-Lab
NCCU-Schultz-Lab deleted the claude/output-naming-strategy-cvs33t branch September 26, 2026 16:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant