SLURM job folders with per-attempt subfolders; SLURM results marked in History - #127
Merged
Merged
Conversation
…late The template exported QUANTUI_RESULTS_DIR=<staging>/results and created that directory, but the batch worker never calls save_result() or list_results(); results reach History only when the app ingests the staging dir after the job ends, using the app's own results dir. The export only left an empty results/ folder in every staging dir. (M-JOBDIRS JD.9) Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.2/JD.3/JD.4 (backend)/JD.8. - Each SLURM submission gets a readable job dir under the staging root, named <label>_<calc>_<method>_<basis> (or a sanitized --job-name), with _2, _3... for a new calculation that reuses a name. The same name is the SLURM --job-name (truncated to 40 chars) so squeue matches the dir. - submit.slurm now creates its own attempt-NN_job<SLURM id>/ dir (numbered max+1 under a flock, mkdir without -p) and the worker writes into it via a new --attempt-dir flag. A hand-run `sbatch submit.slurm` therefore gets a fresh attempt and never touches an earlier attempt's files. SLURM's own slurm-%j.out/.err stay in the job dir because SLURM will not create a missing directory for them. - JobRecord gains job_dir, attempts and ingested_attempts; staging_path resolves to the tracked job's attempt dir, so the monitor, ingest and terminal-state checks read the right attempt unchanged. - SlurmBackend.resubmit() runs the job's submit.slurm again as a new attempt and points the record at the new SLURM job. - The worker only updates the registry record when its SLURM_JOB_ID is the tracked one, so a hand-run attempt cannot flip the record's status. - Apptainer runs bind the job dir when it lives outside $HOME. - Legacy records (no job_dir) keep staging_root/<request_id>/ and behave as before. Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.5/JD.6.
- Ingest now passes provenance extras to save_result on all three paths
(generic, frequency, reorganization energy): execution_backend="slurm"
plus slurm.{job_id, attempt, job_dir, attempt_dir, request_id}. For a
job-dir record the attempt number and job id come from the attempt dir
name, so a hand-run attempt is labelled with the job that produced it.
- History list labels get a "🖥 SLURM <job>·a<attempt>" marker, next to
the existing calibration marker.
- The saved-result card gets a "Ran on" row: SLURM batch, job id, attempt,
and the attempt (or job) folder path.
- Additive: no schema bump. Local results and results ingested before this
change have no provenance keys and render exactly as before.
Contributions:
- Claude (Opus 5.5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration
Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.11. - A Cluster Jobs refresh now saves every finished attempt that is not yet in History: hand-run `sbatch submit.slurm` attempts (which the app never monitored) and tracked successes nobody reconnected to. The attempt the app is currently monitoring, or whose tracked job is still active, is left alone. result.json is written only on success, so failed attempts stay out of History. - JobRecord.ingested_attempts records what has been saved; ingest_attempt() saves and records in one place. A hand-run attempt does not retarget the record's result_dir. - Reconnecting (View progress) to a finished job whose result is already in History now shows it without saving a duplicate entry. Legacy records are covered too: a result_dir that already points at a History entry counts as ingested (this duplicate-on-reconnect predates M-JOBDIRS). Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.4 (UI). Resubmit runs the selected finished or failed job's submit.slurm again as a new attempt in the same job folder (SlurmBackend.resubmit), then opens the Calculate tab monitoring the new attempt, the same as View progress. Earlier attempts' files are kept, and checkpoints let the new attempt pick up where the failed one stopped. It is refused while another calculation is running or monitored, for jobs still active, and for jobs submitted before per-job folders existed (with a message saying why). Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.1. - New compute.slurm_job_root user setting (blank = ~/.quantui/staging). Precedence: QUANTUI_STAGING_DIR > setting > default, resolved in cluster_config.default_staging_root(), so the app, the CLI and the batch worker agree. - System Settings shows a "SLURM job folder" field when SLURM is available. It accepts an absolute path (created if missing), shows an error for relative or unusable paths, and is locked with a note when QUANTUI_STAGING_DIR is set. A change applies to new submissions; jobs already in the registry keep their absolute folder paths. Useful on clusters with small home quotas: point it at scratch or a project directory (Apptainer runs bind it automatically, see the job-dir commit). Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.2 (UI). In SLURM mode the run panel shows an optional "Job name" field. It names the job folder and the SLURM --job-name (sanitized: letters, digits, - and _ kept, other characters become _). Blank keeps the default <formula>_<calc>_<method>_<basis>. A name already in use gets _2, _3..., and the field clears after a successful submit so the next job does not reuse it by accident. The row is hidden in Local mode. Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS JD.7. - In an attempt folder the worker renames the files people copy out, result.molden and trajectory.xyz, to <job name>.molden and <job name>_trajectory.xyz, and records the mapping in result.json["artifact_names"]. Files the app reads back (result.json, orbitals.npz, progress.json, live.log) keep fixed names. - Ingest maps the descriptive names back, so History entries keep the fixed names their loaders expect. - write_worker_result is now atomic (tmp file + os.replace): a finished result.json marks a successful attempt, and a Cluster Jobs refresh may scan for it while the worker is still writing. - Legacy staging runs (no --attempt-dir) are unchanged. Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
M-JOBDIRS docs: job folder layout and rerun behavior in apptainer/slurm/README.md (plus the QUANTUI_STAGING_DIR row and the worker's --attempt-dir flag), the `quantui submit --job-name` help text, and CHANGELOG entries under Unreleased. Contributions: - Claude (Opus 5.5): code edits, review, and conceptual discussion - Jonathan Schultz: overall vision, planning, review, and orchestration Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SLURM job output is now laid out the way
sbatchusers expect. A failed job can be rerun in the same folder without overwriting anything, and every successful SLURM run still shows up in History, marked as a SLURM result. Local (in-kernel) runs are unchanged.<formula>_<calc>_<method>_<basis>, e.g.H2O_opt_B3LYP_def2-SVP.--job-name, sosqueueoutput matches the folder.quantui submit --job-name) overrides the default. A name already in use gets_2,_3, and so on.submit.slurmcreatesattempt-NN_job<SLURM id>/every time it runs. The number is the highest existing one plus one, chosen underflock, and the script usesmkdirwithout-p, so it aborts rather than write into a folder that already exists.--attempt-dirargument.slurm-%j.out/.errfiles stay in the job folder, because SLURM opens them before the script runs and will not create a missing folder.submit.slurmagain as a new attempt, then monitors it.sbatch submit.slurmalso creates a new attempt.SLURM_JOB_IDis the job the record tracks, so a hand-run job cannot overwrite the record's status.execution_backend/slurmfields inresult.json; no schema bump is needed.🖥 SLURM <job id>·a<attempt>, and the result card has a "Ran on" row showing the folder.compute.slurm_job_root.QUANTUI_STAGING_DIRstill overrides it and locks the field.$HOME, Apptainer runs bind it automatically.result.moldenandtrajectory.xyzin attempt folders take the job name.result.jsonis now written atomically.QUANTUI_RESULTS_DIRexport from generated scripts (JD.9).staging/<request_id>/layout and still list, reconnect and ingest. Resubmit on one of them explains why it can't rerun in place.Testing
tests/test_jobdirs.py:bash.submit.slurmand runs it twice by hand, then checks that the first attempt's files are untouched and the registry record is unchanged.dispatchsignature.pre-commit run(ruff + black) passes on all changed files.pytest -m "not network"in a cloud container: 3402 passed, 18 skipped, 23 failed.mainin that container becausepyscf-propertieswould not build there, so they are environmental and unrelated to this branch.[pyscf]extra, so CI is the real check for them.Authorship
🤖 Generated with Claude Code
Generated by Claude Code