diff --git a/CHANGELOG.md b/CHANGELOG.md index 84fb7db..3e0a915 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,18 @@ and this project follows [Semantic Versioning](https://semver.org/spec/v2.0.0.ht ### Added +- **SLURM job folders with per-attempt subfolders (M-JOBDIRS)** โ€” each cluster + job gets a readable folder (`___`, or the + optional **Job name** on the Calculate tab) that doubles as the SLURM job + name. Every run of its `submit.slurm` creates a new `attempt-NN_job/` + subfolder, so **Resubmit** (new, Cluster Jobs tab) and a hand-run + `sbatch submit.slurm` never overwrite an earlier attempt. The job root is + configurable in **System Settings** (or `QUANTUI_STAGING_DIR`), for + clusters with small home quotas. +- **SLURM results marked in History** โ€” entries show + `๐Ÿ–ฅ SLURM ยทa` and the result card has a "Ran on" row with + the job folder. Hand-run attempts are added to History on the next Cluster + Jobs refresh; failed attempts stay in the job folder only. - **PyFock geometry and analysis phase** โ€” native-Windows PBE/def2 geometry optimization now uses PyFock's ASE calculator with analytical density-fitted gradients. Single-point and final optimized-geometry results retain orbital @@ -21,6 +33,13 @@ and this project follows [Semantic Versioning](https://semver.org/spec/v2.0.0.ht parity gate against PySCF. Unsupported hybrids, ions, open-shell systems, solvent, checkpoints, GPU, and orbital analysis are rejected before compute. +### Fixed + +- Reconnecting to a finished SLURM job no longer saves a second copy of its + result to History. +- Removed the unused `QUANTUI_RESULTS_DIR` export from generated SLURM + scripts (it only created an empty `results/` folder per job). + ## [0.8.2] - 2026-08-30 ### Added diff --git a/apptainer/slurm/README.md b/apptainer/slurm/README.md index 8975805..1e2b446 100644 --- a/apptainer/slurm/README.md +++ b/apptainer/slurm/README.md @@ -9,7 +9,37 @@ Reference scripts for operator-driven NCShare / Apptainer batch jobs. | `quantui-batch.sbatch` | Reference `sbatch` template (partition must be set) | | `quantui-gpu-test.sbatch` | GPU smoke test (operator use) | -QuantUI generates per-job scripts automatically under `~/.quantui/staging//submit.slurm` when students submit from the UI. +QuantUI generates per-job scripts automatically when students submit from the UI. + +## Job folders + +Each submission gets its own folder under the job root (default +`~/.quantui/staging`; set it in **System Settings โ†’ SLURM job folder** or with +`QUANTUI_STAGING_DIR`): + +``` +/H2O_opt_B3LYP_def2-SVP/ # job name = SLURM --job-name + request.json submit.slurm + slurm-812345.out slurm-812345.err # SLURM's own output, one pair per job + attempt-01_job812345/ # first run + live.log progress.json ... + attempt-02_job812399/ # a rerun, never overwrites attempt-01 + result.json orbitals.npz H2O_opt_B3LYP_def2-SVP.molden ... + latest -> attempt-02_job812399 +``` + +- The job name defaults to `___`; the optional + **Job name** field on the Calculate tab (or `quantui submit --job-name`) + overrides it. A name already in use gets `_2`, `_3`, .... +- `submit.slurm` creates a new `attempt-NN_job/` folder every time it + runs, so both **Resubmit** on the Cluster Jobs tab and a hand-run + `sbatch submit.slurm` from the job folder start a fresh attempt. Earlier + attempts are never overwritten, and checkpoints let a rerun resume. +- Every successful attempt appears in **History**, marked `๐Ÿ–ฅ SLURM ยทa`, + including hand-run ones (picked up on the next Cluster Jobs refresh). + Failed attempts stay in the job folder only. +- Jobs submitted before per-job folders existed keep their + `~/.quantui/staging//` layout and cannot be resubmitted in place. Status polling uses **`squeue`** for active jobs and **`sacct`** for terminal state, exit code, and cancel confirmation. Cluster Jobs **Remove** clears terminal registry rows without deleting staging logs. @@ -26,6 +56,7 @@ See the [NCShare SLURM batch runbook](https://github.com/The-Schultz-Lab/QuantUI | `QUANTUI_SLURM_CANCEL_CONFIRM_S` | `30` | Seconds to wait for `scancel` confirmation via `sacct` | | `QUANTUI_SLURM_PARTITION` | `common` | Default `#SBATCH` partition | | `QUANTUI_BATCH_IMAGE` | `~/quantui-gpu.sif` | Apptainer image for batch worker | +| `QUANTUI_STAGING_DIR` | *(unset โ€” System Settings value, else `~/.quantui/staging`)* | Job folder root. Overrides the System Settings value and locks that field. A root outside `$HOME` is bound into Apptainer automatically. | ## Operator runbook @@ -36,7 +67,11 @@ See the planning repo runbook: ## Worker entrypoint ```bash -python -m quantui.backends.worker --request /path/to/staging/request.json +python -m quantui.backends.worker --request /path/to/job/request.json \ + --attempt-dir /path/to/job/attempt-01_job812345 ``` +`submit.slurm` passes `--attempt-dir` itself. Without it, outputs go next to +`request.json` (the pre-job-folder layout). + Supported calc types: `single_point`, `geometry_opt`, `frequency`, `tddft`, `nmr`, `pes_scan`, `reorganization_energy`. diff --git a/quantui/app.py b/quantui/app.py index 13df928..67a8757 100644 --- a/quantui/app.py +++ b/quantui/app.py @@ -406,6 +406,9 @@ from quantui.app_runflow import ( update_scan_widgets as _run_update_scan_widgets, ) +from quantui.app_slurm import ( + on_slurm_job_root_changed as _slurm_on_job_root_changed, +) from quantui.app_slurm import ( on_slurm_jobs_cancel_clicked as _slurm_on_jobs_cancel_clicked, ) @@ -415,6 +418,9 @@ from quantui.app_slurm import ( on_slurm_jobs_remove_clicked as _slurm_on_jobs_remove_clicked, ) +from quantui.app_slurm import ( + on_slurm_jobs_resubmit_clicked as _slurm_on_jobs_resubmit_clicked, +) from quantui.app_slurm import ( on_slurm_jobs_view_clicked as _slurm_on_jobs_view_clicked, ) @@ -568,6 +574,7 @@ from quantui.app_xyz_input import ( on_xyz_fill_table as _xyz_on_fill_table, ) +from quantui.backends import cluster_config as _cluster_cfg from quantui.backends.dispatch import is_slurm_available from quantui.cancellation import CalcCancelled as _CalcCancelled @@ -1440,6 +1447,10 @@ class QuantUIApp: density_fit_enabled_cb: Any freq_parallel_enabled_cb: Any execution_backend_dd: Any + slurm_job_root_txt: Any + _slurm_job_name_txt: Any + _slurm_job_name_row: Any + slurm_job_root_note: Any quantum_engine_dd: Any quantum_engine_note: Any engine_capability_html: Any @@ -1491,6 +1502,7 @@ class QuantUIApp: _slurm_jobs_view_btn: Any _slurm_jobs_cancel_btn: Any _slurm_jobs_remove_btn: Any + _slurm_jobs_resubmit_btn: Any _slurm_jobs_status_html: Any slurm_jobs_tab_panel: Any _slurm_jobs_tab_index: int | None @@ -2124,6 +2136,9 @@ def _build_status_panel(self) -> None: execution_backend=self._user_settings.compute.execution_backend, slurm_available=is_slurm_available(), quantum_engine=self._user_settings.compute.quantum_engine, + slurm_job_root=self._user_settings.compute.slurm_job_root, + slurm_job_root_env_locked=_cluster_cfg.staging_root_env_configured(), + slurm_job_root_effective=str(_cluster_cfg.default_staging_root()), ) # โ”€โ”€ Welcome header โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ @@ -2464,6 +2479,9 @@ def _sync_root_tab_layout(self) -> None: self._slurm_jobs_tab_index = ( order.index("slurm_jobs") if "slurm_jobs" in order else None ) + self._slurm_job_name_row.layout.display = ( + "flex" if "slurm_jobs" in order else "none" + ) self._root_tab_order_cache = list(order) if "slurm_jobs" in order: @@ -2537,6 +2555,12 @@ def _wire_callbacks(self) -> None: self.execution_backend_dd.observe( self._safe_cb(self._on_execution_backend_changed), names="value" ) + self.slurm_job_root_txt.observe( + self._safe_cb( + lambda change: _slurm_on_job_root_changed(self, change["new"]) + ), + names="value", + ) self.quantum_engine_dd.observe( self._safe_cb(self._on_quantum_engine_changed), names="value" ) @@ -2650,6 +2674,9 @@ def _wire_callbacks(self) -> None: self._slurm_jobs_remove_btn.on_click( self._safe_cb(lambda _btn: _slurm_on_jobs_remove_clicked(self, _btn)) ) + self._slurm_jobs_resubmit_btn.on_click( + self._safe_cb(lambda _btn: _slurm_on_jobs_resubmit_clicked(self, _btn)) + ) self.cancel_btn.on_click(self._safe_cb(self._on_cancel)) self.basis_fix_btn.on_click(self._safe_cb(self._on_basis_fix)) self.charge_mult_suggest_btn.on_click( diff --git a/quantui/app_builders.py b/quantui/app_builders.py index 0c95c1c..183b2ec 100644 --- a/quantui/app_builders.py +++ b/quantui/app_builders.py @@ -128,6 +128,9 @@ def build_status_panel( execution_backend: str = "local", slurm_available: bool = False, quantum_engine: str = "auto", + slurm_job_root: str = "", + slurm_job_root_env_locked: bool = False, + slurm_job_root_effective: str = "", ) -> None: """Build the Status tab panel.""" cores, mem_gb = get_session_resources_fn() @@ -407,6 +410,44 @@ def _render_status(gpu_state: Any) -> str: "Use the Cluster Jobs tab to monitor and cancel runs." ) + # SLURM job folder root (M-JOBDIRS JD.1). Built unconditionally so the + # app can wire it; only shown when SLURM is available. + app.slurm_job_root_txt = widgets.Text( + value=slurm_job_root_effective if slurm_job_root_env_locked else slurm_job_root, + placeholder="~/.quantui/staging (default)", + continuous_update=False, + disabled=slurm_job_root_env_locked, + layout=layout_fn(width="420px"), + ) + if slurm_job_root_env_locked: + _root_note_text = ( + "Set by QUANTUI_STAGING_DIR in the container or " + "session environment." + ) + else: + _root_note_text = ( + "Each cluster job gets its own folder here, with one subfolder per " + "attempt. Leave blank for the default; on clusters with a small " + "home quota, use a scratch or project folder. Applies to new " + "submissions." + ) + app.slurm_job_root_note = widgets.HTML( + f'
' + f"{_root_note_text}
" + ) + slurm_root_rows: list[Any] = [] + if slurm_available: + slurm_root_rows = [ + widgets.HTML( + f'
SLURM job folder ' + f'' + "(persists across launches)
" + ), + app.slurm_job_root_txt, + app.slurm_job_root_note, + ] + fp_settings_rows: list[Any] = [ fp_toggle_label, app.freq_parallel_enabled_cb, @@ -431,6 +472,7 @@ def _render_status(gpu_state: Any) -> str: exec_backend_label, app.execution_backend_dd, exec_backend_note, + *slurm_root_rows, ], layout=layout_fn(margin="0 0 8px 0"), ) @@ -1485,6 +1527,31 @@ def build_shared_widgets( # native phase (first gradient / Hessian) visibly advances. app._run_elapsed_lbl = widgets.HTML(value="") + # Optional job folder / SLURM job name (M-JOBDIRS JD.2). Shown only in + # SLURM mode; blank uses ___. + app._slurm_job_name_txt = widgets.Text( + value="", + placeholder="default: ___", + description="Job name:", + style={"description_width": "80px"}, + layout=layout_fn(width="460px"), + tooltip=( + "Names this job's folder and its SLURM job. A name already in use " + "gets _2, _3, ... appended." + ), + ) + app._slurm_job_name_row = widgets.VBox( + [ + app._slurm_job_name_txt, + widgets.HTML( + f'
Optional. Letters, digits, - and _ are kept; ' + "other characters become _.
" + ), + ], + layout=layout_fn(display="none"), + ) + app._slurm_job_banner = widgets.HTML(value="", layout=layout_fn(display="none")) app._slurm_reconnect_btn = widgets.Button( description="View SLURM progress", @@ -1994,6 +2061,7 @@ def build_run_section(app: Any, *, layout_fn: Any) -> None: app.perf_estimate_html, app._resume_notice_html, app._resume_cb, + app._slurm_job_name_row, app._slurm_job_banner, widgets.HBox( [app._slurm_reconnect_btn], @@ -3456,6 +3524,15 @@ def build_slurm_jobs_tab(app: Any, *, layout_fn: Any) -> None: layout=layout_fn(width="120px"), tooltip="Cancel the selected active cluster job", ) + app._slurm_jobs_resubmit_btn = widgets.Button( + description="Resubmit", + icon="redo", + layout=layout_fn(width="120px"), + tooltip=( + "Run a finished or failed job again in the same job folder, as a " + "new attempt. Earlier attempts' files are kept." + ), + ) app._slurm_jobs_remove_btn = widgets.Button( description="Remove", icon="trash", @@ -3465,7 +3542,7 @@ def build_slurm_jobs_tab(app: Any, *, layout_fn: Any) -> None: app._slurm_jobs_status_html = widgets.HTML( value=( f'' - "Select a job and use View progress or Cancel." + "Select a job and use View progress, Cancel, or Resubmit." ) ) @@ -3485,6 +3562,7 @@ def build_slurm_jobs_tab(app: Any, *, layout_fn: Any) -> None: [ app._slurm_jobs_view_btn, app._slurm_jobs_cancel_btn, + app._slurm_jobs_resubmit_btn, app._slurm_jobs_remove_btn, ], layout=layout_fn(gap="8px", margin="6px 0"), diff --git a/quantui/app_formatters.py b/quantui/app_formatters.py index 5436b75..f635569 100644 --- a/quantui/app_formatters.py +++ b/quantui/app_formatters.py @@ -2,6 +2,7 @@ from __future__ import annotations +import html from pathlib import Path from typing import Any, Optional @@ -815,6 +816,49 @@ def format_reorg_result(r: Any) -> str: ) +def slurm_history_marker(data: dict[str, Any]) -> str: + """Short History-list marker for a SLURM result, e.g. ``๐Ÿ–ฅ SLURM 812345ยทa2 ``. + + Empty for local results and for results saved before SLURM provenance + was recorded (M-JOBDIRS JD.6). + """ + if data.get("execution_backend") != "slurm": + return "" + info = data.get("slurm") or {} + job_id = info.get("job_id") + attempt = info.get("attempt") + text = "๐Ÿ–ฅ SLURM" + if job_id: + text += f" {job_id}" + if attempt: + text += f"ยทa{attempt}" + return text + " " + + +def _slurm_provenance_row(data: dict[str, Any]) -> str: + """History-card row naming the SLURM job, attempt and job folder.""" + if data.get("execution_backend") != "slurm": + return "" + info = data.get("slurm") or {} + parts = ["SLURM batch"] + if info.get("job_id"): + parts.append(f"job {html.escape(str(info['job_id']))}") + if info.get("attempt"): + parts.append(f"attempt {int(info['attempt'])}") + value = " · ".join(parts) + folder = info.get("attempt_dir") or info.get("job_dir") + if folder: + value += ( + f'
{html.escape(str(folder))}' + ) + return ( + f'Ran on' + f'{value}' + ) + + def format_past_result(data: dict[str, Any], result_dir: Optional[Path] = None) -> str: """Format a saved result.json payload as an HTML result card.""" import base64 as _b64 @@ -912,6 +956,7 @@ def format_past_result(data: dict[str, Any], result_dir: Optional[Path] = None) # Shared 'extra' rows (correlation breakdown / solvent / device / dipole / # Mulliken) โ€” same builder as the live card so the two never drift. _extra = _result_extra_rows(lambda k, d=None: data.get(k, d)) + _extra += _slurm_provenance_row(data) # Embed thumbnail if saved _thumb_html = "" diff --git a/quantui/app_runflow.py b/quantui/app_runflow.py index 0589bea..77e0be1 100644 --- a/quantui/app_runflow.py +++ b/quantui/app_runflow.py @@ -2667,6 +2667,7 @@ def refresh_results_browser(app: Any) -> None: from quantui import list_results, load_result except ImportError: return + from quantui.app_formatters import slurm_history_marker from quantui.app_history import ( apply_history_filter, entry_date, @@ -2696,12 +2697,14 @@ def refresh_results_browser(app: Any) -> None: # comes from result.json's ``calibration_run_id`` extras field # written by the worker. calib_marker = "๐Ÿ”ง " if data.get("calibration_run_id") else "" + # SLURM batch results carry job id + attempt (M-JOBDIRS JD.6). + slurm_marker = slurm_history_marker(data) formula = data.get("formula", "?") method = data.get("method", "?") basis = data.get("basis", "?") label = ( f"{ts} ยท [{calc_badge}] " - f"{calib_marker}{formula} " + f"{calib_marker}{slurm_marker}{formula} " f"{method}/{basis}" ) entries.append( diff --git a/quantui/app_slurm.py b/quantui/app_slurm.py index c1294b9..5fdfa36 100644 --- a/quantui/app_slurm.py +++ b/quantui/app_slurm.py @@ -13,6 +13,7 @@ import logging import threading import time +from pathlib import Path from typing import Any from IPython.display import HTML @@ -169,8 +170,11 @@ def submit_slurm_run(app: Any) -> None: app.run_btn.disabled = True app.run_status.value = "Submitting to SLURMโ€ฆ" + name_widget = getattr(app, "_slurm_job_name_txt", None) + job_name = (str(name_widget.value).strip() or None) if name_widget else None + try: - request_id = backend.dispatch(request) + request_id = backend.dispatch(request, job_name=job_name) except SecurityError as exc: _reset_run_ui_after_submit_failure(app) app.run_status.value = str(exc) @@ -184,6 +188,10 @@ def submit_slurm_run(app: Any) -> None: _append_run_html(app, format_error_html(str(exc))) return + if name_widget is not None: + # A name belongs to one job; clear it so the next submit does not + # silently reuse it (it would get a _2 suffix). + name_widget.value = "" refresh_slurm_jobs_tab(app) _update_slurm_jobs_tab_title(app) @@ -194,7 +202,7 @@ def submit_slurm_run(app: Any) -> None: _append_run_stdout( app, f"\n๐Ÿ“ค Submitted batch job {slurm_id} (request {request_id})\n" - f"Staging directory: {record.staging_dir if record else '?'}\n", + f"Job directory: {(record.job_dir or record.staging_dir) if record else '?'}\n", ) threading.Thread( target=_monitor_slurm_job, @@ -298,6 +306,36 @@ def _slurm_job_error_cell(record: JobRecord) -> str: return f"{html.escape(code)}: {html.escape(message)}" +def ingest_new_slurm_attempts(app: Any) -> list[Path]: + """Save finished, not-yet-ingested attempts to History (M-JOBDIRS JD.11). + + Picks up hand-run ``sbatch submit.slurm`` attempts (which the app never + monitored) and tracked successes nobody reconnected to. The attempt the + app is monitoring right now is left to the monitor, and a tracked + attempt whose job is still active is skipped. + """ + from quantui.backends.slurm_ingest import ingest_attempt, uningested_attempts + + registry = ensure_job_registry(app) + monitored = getattr(app, "_slurm_active_request_id", None) + saved: list[Path] = [] + for record in list_slurm_jobs(app): + for attempt in uningested_attempts(record): + tracked = attempt == record.staging_path + if tracked and ( + record.request_id == monitored + or record.status.lower() in _ACTIVE_SLURM_STATUSES + ): + continue + try: + saved.append(ingest_attempt(registry, record, attempt)) + except Exception: # noqa: BLE001 โ€” one bad attempt must not block others + logger.exception("Failed to ingest SLURM attempt %s", attempt) + continue + record = registry.load(record.request_id) or record + return saved + + def refresh_slurm_jobs_tab(app: Any) -> None: """Re-render the Cluster Jobs tab from the on-disk registry.""" summary = getattr(app, "_slurm_jobs_summary_html", None) @@ -314,6 +352,21 @@ def refresh_slurm_jobs_tab(app: Any) -> None: except Exception: # noqa: BLE001 โ€” UI refresh must not crash the app logger.exception("Failed to refresh SLURM registry statuses") + new_saved = ingest_new_slurm_attempts(app) + if new_saved: + from quantui.app_runflow import refresh_results_browser + + try: + refresh_results_browser(app) + except Exception: # noqa: BLE001 + logger.exception("Failed to refresh History after SLURM ingest") + status = getattr(app, "_slurm_jobs_status_html", None) + if status is not None: + status.value = ( + f'' + f"Saved {len(new_saved)} finished cluster run(s) to History." + ) + records = list_slurm_jobs(app) active = active_slurm_job_count(app) limit = max_concurrent_slurm_jobs() @@ -447,6 +500,54 @@ def on_slurm_jobs_cancel_clicked(app: Any, _btn: Any = None) -> None: ) +def resubmit_slurm_job(app: Any, request_id: str) -> tuple[bool, str]: + """Resubmit a finished job as a new attempt in its job dir (M-JOBDIRS JD.4). + + Returns ``(ok, user message)``. On success the Calculate tab starts + monitoring the new attempt. + """ + if getattr(app, "_calc_running", False): + return False, ( + "A calculation is already running or being monitored. " + "Wait for it to finish before resubmitting." + ) + ensure_job_registry(app) + backend = slurm_backend_for_app(app) + try: + new_job_id = backend.resubmit(request_id) + except (ValueError, SecurityError) as exc: + return False, str(exc) + except RuntimeError as exc: + return False, f"Resubmit failed: {exc}" + record = app._job_registry.load(request_id) + folder = (record.job_dir if record else None) or "?" + return True, ( + f"Resubmitted as SLURM job {new_job_id}. Its attempt folder appears " + f"in {folder} when the job starts; earlier attempts are kept." + ) + + +def on_slurm_jobs_resubmit_clicked(app: Any, _btn: Any = None) -> None: + select = getattr(app, "_slurm_jobs_select", None) + request_id = select.value if select is not None else None + if not request_id: + return + ok, message = resubmit_slurm_job(app, request_id) + refresh_slurm_jobs_tab(app) + _update_slurm_jobs_tab_title(app) + status = getattr(app, "_slurm_jobs_status_html", None) + if status is not None: + color = _theme.css.TEXT_STRONG if ok else _theme.css.ACCENT_ERROR + status.value = ( + f'{html.escape(message)}' + ) + if ok: + attach_slurm_job(app, request_id) + go_to = getattr(app, "_go_to_calculate_tab", None) + if callable(go_to): + go_to() + + def on_slurm_jobs_remove_clicked(app: Any, _btn: Any = None) -> None: select = getattr(app, "_slurm_jobs_select", None) request_id = select.value if select is not None else None @@ -555,31 +656,31 @@ def _ingest_terminal_job(app: Any, record: Any) -> None: def _ingest_success(app: Any, record: Any) -> None: + # Reload: the record passed in may predate an ingest done elsewhere + # (a Cluster Jobs refresh, or an earlier reconnect). + record = app._job_registry.load(record.request_id) or record staging = record.staging_path result_path = staging / "result.json" - log_path = record.live_log_path if not result_path.exists(): app.run_status.value = "SLURM job finished but result.json is missing." return payload = json.loads(result_path.read_text(encoding="utf-8")) - log_text = ( - log_path.read_text(encoding="utf-8", errors="replace") - if log_path.exists() - else "" - ) try: from quantui.app_runflow import refresh_results_browser from quantui.backends.slurm_ingest import ( + already_ingested, completion_summary_html, - ingest_staging_success, + ingest_attempt, ) - saved_dir = ingest_staging_success(record, log_text) - app._job_registry.update_status( - record.request_id, "success", result_dir=str(saved_dir) - ) + if already_ingested(record, staging) and record.result_dir: + # Never save the same run to History twice (e.g. View progress + # on a finished job, or a refresh that already picked it up). + saved_dir = Path(record.result_dir) + else: + saved_dir = ingest_attempt(app._job_registry, record) app._last_result_dir = saved_dir refresh_results_browser(app) calc_type = payload.get("calc_type", record.calc_type) @@ -662,6 +763,46 @@ def on_slurm_reconnect_clicked(app: Any, _btn: Any = None) -> None: attach_slurm_job(app, request_id) +def on_slurm_job_root_changed(app: Any, value: str) -> None: + """Validate and persist the SLURM job folder root (M-JOBDIRS JD.1). + + Blank restores the default. A new root applies to new submissions; + existing jobs keep their absolute folder paths in the registry. + """ + note = getattr(app, "slurm_job_root_note", None) + + def _say(message: str, *, error: bool = False) -> None: + if note is None: + return + color = _theme.css.ACCENT_ERROR if error else _theme.css.TEXT_SUBTLE + note.value = ( + f'
' + f"{html.escape(message)}
" + ) + + if _cluster_cfg.staging_root_env_configured(): + return + raw = (value or "").strip() + if raw: + root = Path(raw).expanduser() + if not root.is_absolute(): + _say("Use a full path (for example /work//quantui-jobs).", error=True) + return + try: + root.mkdir(parents=True, exist_ok=True) + except OSError as exc: + _say(f"Cannot use that folder: {exc}", error=True) + return + if raw == app._user_settings.compute.slurm_job_root: + return + app._user_settings.compute.slurm_job_root = raw + app._user_settings.save() + # Rebuild the registry on next use so new jobs land under the new root. + app._job_registry = None + where = raw or str(_cluster_cfg.DEFAULT_STAGING_ROOT) + _say(f"New cluster jobs will be created in {where}.") + + def use_slurm_execution(app: Any) -> bool: pref = getattr(app._user_settings.compute, "execution_backend", "local") return pref == "slurm" and is_slurm_available() diff --git a/quantui/backends/cluster_config.py b/quantui/backends/cluster_config.py index 17ceaeb..e615e51 100644 --- a/quantui/backends/cluster_config.py +++ b/quantui/backends/cluster_config.py @@ -9,6 +9,7 @@ import logging import os +import shlex from pathlib import Path logger = logging.getLogger(__name__) @@ -117,11 +118,32 @@ def default_jobs_root() -> Path: return Path.home() / ".quantui" / "jobs" +DEFAULT_STAGING_ROOT = Path("~") / ".quantui" / "staging" + + +def staging_root_env_configured() -> bool: + """True when ``QUANTUI_STAGING_DIR`` pins the job root (UI shows it locked).""" + return bool(os.environ.get("QUANTUI_STAGING_DIR")) + + def default_staging_root() -> Path: + """Root for SLURM job folders. + + Precedence: ``QUANTUI_STAGING_DIR`` > the ``compute.slurm_job_root`` + user setting (M-JOBDIRS JD.1) > ``~/.quantui/staging``. + """ override = os.environ.get("QUANTUI_STAGING_DIR") if override: return Path(override).expanduser() - return Path.home() / ".quantui" / "staging" + try: + from quantui.user_settings import UserSettings + + configured = UserSettings.load().compute.slurm_job_root + except Exception: # noqa: BLE001 โ€” a broken settings file must not block jobs + configured = "" + if configured: + return Path(configured).expanduser() + return DEFAULT_STAGING_ROOT.expanduser() # Apptainer image for batch workers (NCShare-oriented default path). @@ -130,6 +152,40 @@ def default_staging_root() -> Path: os.path.expanduser("~/quantui-gpu.sif"), ) +# Shell block that gives every run of submit.slurm its own attempt dir +# (M-JOBDIRS). Numbering is max(existing NN) + 1 under a lock, and ``mkdir`` +# without ``-p`` makes the script fail rather than write into an existing +# directory โ€” an earlier attempt's files are never overwritten, whether the +# job came from QuantUI's Resubmit or a hand-run ``sbatch submit.slurm``. +# ``$JOB_DIR`` is set by the line build_attempt_setup() prepends. +_ATTEMPT_SETUP_BODY = r"""exec 9>"$JOB_DIR/.attempt.lock" +if command -v flock >/dev/null 2>&1; then flock 9; fi +attempt_n=0 +for d in "$JOB_DIR"/attempt-*; do + [ -d "$d" ] || continue + k="${d##*/attempt-}" + k="${k%%_*}" + case "$k" in + ''|*[!0-9]*) continue ;; + esac + k=$((10#$k)) + if [ "$k" -gt "$attempt_n" ]; then attempt_n=$k; fi +done +attempt_n=$((attempt_n + 1)) +ATTEMPT_DIR="$JOB_DIR/$(printf 'attempt-%02d_job%s' "$attempt_n" "${SLURM_JOB_ID:-manual$$}")" +mkdir "$ATTEMPT_DIR" +exec 9>&- +ln -sfn "${ATTEMPT_DIR##*/}" "$JOB_DIR/latest" 2>/dev/null || true +cd "$ATTEMPT_DIR" +echo "Attempt directory: $ATTEMPT_DIR" +""" + + +def build_attempt_setup(job_dir: str) -> str: + """Return the attempt-dir shell block for a job dir (shell-quoted).""" + return f"\nJOB_DIR={shlex.quote(job_dir)}\n{_ATTEMPT_SETUP_BODY}" + + # SLURM batch script template. ``{worker_command}`` is the full command line # run inside the allocation (Apptainer-wrapped when configured). SLURM_SCRIPT_TEMPLATE = """#!/bin/bash @@ -150,9 +206,7 @@ def default_staging_root() -> Path: echo "Working directory: $(pwd)" export OMP_NUM_THREADS="${{SLURM_CPUS_PER_TASK:-{cores}}}" -export QUANTUI_RESULTS_DIR="{results_dir}" -mkdir -p "$QUANTUI_RESULTS_DIR" - +{attempt_setup} {worker_command} echo "Job completed at: $(date -u +%Y-%m-%dT%H:%M:%SZ)" diff --git a/quantui/backends/registry.py b/quantui/backends/registry.py index 28285f6..f320d50 100644 --- a/quantui/backends/registry.py +++ b/quantui/backends/registry.py @@ -2,14 +2,24 @@ On-disk job registry for execution backends (M-CLUSTER2 CL2.1). Each submitted calculation gets ``/.json`` plus a -companion staging directory under ``staging_root//`` for live logs -and progress files while the job runs. +directory under ``staging_root`` for live logs, progress files and results. + +Two layouts exist: + +* **Job dir (M-JOBDIRS)** โ€” ``staging_root//`` holds ``request.json`` + and ``submit.slurm``; every run of that script creates its own + ``attempt-NN_job/`` subdirectory, so a retry never overwrites an + earlier attempt's files. :attr:`JobRecord.staging_path` resolves to the + attempt dir of the currently tracked SLURM job. +* **Legacy** โ€” ``staging_root//`` holds everything directly. + Records written before M-JOBDIRS have no ``job_dir`` and keep this behavior. """ from __future__ import annotations import json import logging +import re import time from dataclasses import asdict, dataclass, field from datetime import datetime, timezone @@ -25,10 +35,21 @@ _ACTIVE_STATUSES = frozenset({"queued", "pending", "running", "submitted"}) +_ATTEMPT_DIR_RE = re.compile(r"^attempt-(\d+)_job(.+)$") + + def _utc_now() -> str: return datetime.now(timezone.utc).replace(microsecond=0).isoformat() +def parse_attempt_dir_name(name: str) -> Optional[tuple[int, str]]: + """Return ``(attempt number, SLURM job id)`` for an attempt dir name, else None.""" + m = _ATTEMPT_DIR_RE.match(name) + if m is None: + return None + return int(m.group(1)), m.group(2) + + @dataclass class JobRecord: request_id: str @@ -43,6 +64,15 @@ class JobRecord: result_dir: Optional[str] = None resources: Dict[str, Any] = field(default_factory=dict) error: Optional[Dict[str, Any]] = None + # M-JOBDIRS โ€” None on legacy records (staging_dir holds everything). + job_dir: Optional[str] = None + # One entry per sbatch submission made through QuantUI: + # {"slurm_job_id", "submitted_at", "source"}. Hand-run sbatch attempts are + # not listed here; they are found on disk (see attempt_dirs()). + attempts: List[Dict[str, Any]] = field(default_factory=list) + # Names of attempt dirs (or, for legacy records, the staging dir) already + # saved to History, so a result is never ingested twice. + ingested_attempts: List[str] = field(default_factory=list) def to_dict(self) -> Dict[str, Any]: return asdict(self) @@ -62,15 +92,54 @@ def from_dict(cls, data: Dict[str, Any]) -> JobRecord: result_dir=data.get("result_dir"), resources=dict(data.get("resources") or {}), error=data.get("error"), + job_dir=data.get("job_dir"), + attempts=list(data.get("attempts") or []), + ingested_attempts=list(data.get("ingested_attempts") or []), ) @property def request_obj(self) -> CalculationRequest: return CalculationRequest.from_dict(self.request) + @property + def job_path(self) -> Optional[Path]: + return Path(self.job_dir) if self.job_dir else None + + def attempt_dirs(self) -> List[Path]: + """Attempt dirs on disk, oldest first (empty for legacy records).""" + job_path = self.job_path + if job_path is None or not job_path.is_dir(): + return [] + found = [] + for child in job_path.iterdir(): + parsed = parse_attempt_dir_name(child.name) + if parsed is not None and child.is_dir() and not child.is_symlink(): + found.append((parsed[0], child.name, child)) + return [path for _n, _name, path in sorted(found)] + + def attempt_dir_for(self, slurm_job_id: Optional[str]) -> Optional[Path]: + """The attempt dir created by SLURM job *slurm_job_id*, if it exists yet.""" + if not slurm_job_id: + return None + for path in self.attempt_dirs(): + parsed = parse_attempt_dir_name(path.name) + if parsed is not None and parsed[1] == str(slurm_job_id): + return path + return None + @property def staging_path(self) -> Path: - return Path(self.staging_dir) + """Where the tracked run writes ``live.log`` / ``result.json``. + + Legacy records: the staging dir. Job-dir records: the attempt dir of + the tracked SLURM job, or the job dir itself while that job has not + started yet (the attempt dir is created by the batch script). + """ + job_path = self.job_path + if job_path is None: + return Path(self.staging_dir) + attempt = self.attempt_dir_for(self.slurm_job_id) + return attempt if attempt is not None else job_path @property def live_log_path(self) -> Path: @@ -102,6 +171,22 @@ def staging_dir_for(self, request_id: str) -> Path: path.mkdir(parents=True, exist_ok=True) return path + def new_job_dir(self, name: str) -> Path: + """Create a fresh job dir named *name* (``_2``, ``_3``โ€ฆ on collision). + + Never reuses an existing directory: each new calculation gets its own. + """ + n = 1 + while True: + candidate = name if n == 1 else f"{name}_{n}" + path = safe_join(self.staging_root, candidate) + try: + path.mkdir(parents=True, exist_ok=False) + except FileExistsError: + n += 1 + continue + return path + def create( self, request: CalculationRequest, @@ -109,8 +194,19 @@ def create( *, resources: Optional[Dict[str, Any]] = None, status: str = "queued", + job_name: Optional[str] = None, ) -> JobRecord: - staging = self.staging_dir_for(request.request_id) + """Register *request*. + + With *job_name* the record gets a job dir (M-JOBDIRS layout); without + it, a legacy ``staging_root//`` dir. + """ + if job_name: + staging = self.new_job_dir(job_name) + job_dir: Optional[str] = str(staging) + else: + staging = self.staging_dir_for(request.request_id) + job_dir = None now = _utc_now() record = JobRecord( request_id=request.request_id, @@ -122,6 +218,7 @@ def create( created_at=now, updated_at=now, resources=dict(resources or {}), + job_dir=job_dir, ) self.save(record) return record @@ -184,6 +281,45 @@ def update_status( self.save(record) return record + def start_attempt( + self, request_id: str, slurm_job_id: str, *, source: str = "submit" + ) -> Optional[JobRecord]: + """Point *request_id* at a newly submitted SLURM job. + + Used for the first submission and for every Resubmit: the record + tracks the new job (status ``submitted``, previous error cleared) and + the submission is appended to :attr:`JobRecord.attempts`. + """ + record = self.load(request_id) + if record is None: + return None + record.status = "submitted" + record.slurm_job_id = slurm_job_id + record.error = None + record.attempts.append( + { + "slurm_job_id": slurm_job_id, + "submitted_at": _utc_now(), + "source": source, + } + ) + self.save(record) + return record + + def mark_ingested( + self, request_id: str, attempt_key: str, *, result_dir: Optional[str] = None + ) -> Optional[JobRecord]: + """Record that *attempt_key* has been saved to History.""" + record = self.load(request_id) + if record is None: + return None + if attempt_key not in record.ingested_attempts: + record.ingested_attempts.append(attempt_key) + if result_dir is not None: + record.result_dir = result_dir + self.save(record) + return record + def _submit_meta_path(self) -> Path: return self.jobs_root / ".slurm_submit_meta.json" diff --git a/quantui/backends/slurm.py b/quantui/backends/slurm.py index 5333745..de6d5b8 100644 --- a/quantui/backends/slurm.py +++ b/quantui/backends/slurm.py @@ -32,11 +32,14 @@ from .registry import JobRegistry from .slurm_errors import format_error_for_student from .slurm_utils import ( + SLURM_JOB_NAME_MAX_LEN, SlurmJobAccounting, + default_job_name, estimate_slurm_resources, parse_sacct_accounting, parse_sacct_states, parse_slurm_job_id, + sanitize_job_name, ) logger = logging.getLogger(__name__) @@ -118,27 +121,27 @@ def dispatch( if email is not None: resolved_events = validate_mail_events(mail_events) + # M-JOBDIRS: the job dir name doubles as the SLURM job name, so + # squeue output matches the directory on disk. + name = sanitize_job_name(job_name or "") or default_job_name(request) record = self.registry.create( request, self.backend_id, resources=resources, status="queued", + job_name=name, ) - staging = record.staging_path - request_path = staging / "request.json" + job_dir = record.job_path + assert job_dir is not None # create(job_name=...) always sets it + request_path = job_dir / "request.json" request_path.write_text(json.dumps(request.to_dict(), indent=2)) - label = ( - job_name or f"{request.molecule.get('label', 'quantui')}_{request.method}" - ) - label = "".join(c if c.isalnum() or c in "-_" else "_" for c in label)[:40] - - slurm_script = staging / "submit.slurm" + slurm_script = job_dir / "submit.slurm" self._write_slurm_script( slurm_script, - job_name=label, + job_name=job_dir.name[:SLURM_JOB_NAME_MAX_LEN], request_path=request_path, - staging_dir=staging, + job_dir=job_dir, resources=resources, depends_on=depends_on, email=email, @@ -161,15 +164,52 @@ def dispatch( ) raise - self.registry.update_status( - request.request_id, - "submitted", - slurm_job_id=slurm_job_id, - ) + self.registry.start_attempt(request.request_id, slurm_job_id, source="submit") self.registry.record_slurm_submit() return request.request_id - def _worker_command(self, request_path: Path, staging_dir: Path) -> str: + def resubmit(self, request_id: str) -> str: + """Run a finished job's ``submit.slurm`` again as a new attempt. + + The new attempt gets its own ``attempt-NN_job/`` dir inside the + same job dir (earlier attempts are untouched), and the record switches + to tracking the new SLURM job. Returns the new SLURM job id. + + Raises ``ValueError`` when the job cannot be resubmitted (unknown, + still active, or a legacy record without a job dir) and + ``RuntimeError`` when ``sbatch`` fails. + """ + record = self.registry.load(request_id) + if record is None: + raise ValueError(f"Job {request_id} is not in your registry.") + if record.status.lower() in _ACTIVE_RECORD: + raise ValueError( + f"Job {request_id} is still active ({record.status}). " + "Wait for it to finish or cancel it first." + ) + job_dir = record.job_path + script = job_dir / "submit.slurm" if job_dir is not None else None + if script is None or not script.exists(): + raise ValueError( + "This job was submitted before per-job folders existed, so it " + "cannot be resubmitted in place. Submit it again from the " + "Calculate tab." + ) + + active_slurm = sum( + 1 + for active in self.registry.list_active() + if active.backend_id == self.backend_id + ) + check_concurrent_job_limit(active_slurm) + check_submit_cooldown(self.registry.seconds_since_last_slurm_submit()) + + slurm_job_id = self._submit_to_slurm(script) + self.registry.start_attempt(request_id, slurm_job_id, source="resubmit") + self.registry.record_slurm_submit() + return slurm_job_id + + def _worker_command(self, request_path: Path, job_dir: Path) -> str: # AUDIT F21 โ€” every path here is inserted into shell text (this # string is embedded verbatim into the generated sbatch script), # not passed as an argv list, so an unquoted path containing a @@ -178,16 +218,25 @@ def _worker_command(self, request_path: Path, staging_dir: Path) -> str: # "--request /tmp/audit" plus a stray "folder/request.json" # argument. shlex.quote makes every one of these shell-safe # regardless of what it contains. + # + # "$ATTEMPT_DIR" is the one deliberate exception: it is a shell + # variable set by the attempt-setup block at run time, so it is + # double-quoted (expanded by the shell, never word-split) instead. py = shlex.quote(sys.executable) request_arg = shlex.quote(str(request_path)) - inner = f"{py} -m quantui.backends.worker --request {request_arg}" + inner = ( + f"{py} -m quantui.backends.worker --request {request_arg} " + '--attempt-dir "$ATTEMPT_DIR"' + ) if self.use_apptainer: image = shlex.quote(self.apptainer_image) - staging_arg = shlex.quote(str(staging_dir)) - return ( - f'apptainer exec --nv --bind "$HOME:$HOME" --pwd {staging_arg} ' - f"{image} {inner}" - ) + binds = '--bind "$HOME:$HOME"' + # A job root outside $HOME (e.g. scratch) must be bound too, or + # the container cannot see request.json or the attempt dir. + if not _is_within(job_dir, Path.home()): + job_arg = shlex.quote(str(job_dir)) + binds += f" --bind {job_arg}:{job_arg}" + return f'apptainer exec --nv {binds} --pwd "$ATTEMPT_DIR" {image} {inner}' return inner def _write_slurm_script( @@ -196,7 +245,7 @@ def _write_slurm_script( *, job_name: str, request_path: Path, - staging_dir: Path, + job_dir: Path, resources: Dict[str, int | str], depends_on: str | None, email: str | None, @@ -211,18 +260,20 @@ def _write_slurm_script( extra.append(f"#SBATCH --mail-type={events_str}") optional = "\n" + "\n".join(extra) if extra else "" - worker_command = self._worker_command(request_path, staging_dir) - results_dir = staging_dir / "results" + worker_command = self._worker_command(request_path, job_dir) content = cfg.SLURM_SCRIPT_TEMPLATE.format( job_name=job_name, partition=self.partition, cores=resources["cores"], memory=resources["memory_gb"], walltime=resources["walltime"], - output_file=str(staging_dir / "slurm-%j.out"), - error_file=str(staging_dir / "slurm-%j.err"), + # SLURM opens these before the script runs and will not create + # a missing directory, so they live in the job dir (already + # unique per job via %j), not in the attempt dir. + output_file=str(job_dir / "slurm-%j.out"), + error_file=str(job_dir / "slurm-%j.err"), optional_directives=optional, - results_dir=str(results_dir), + attempt_setup=cfg.build_attempt_setup(str(job_dir)), worker_command=worker_command, ) output_path.write_text(content) @@ -628,6 +679,14 @@ def _record_cancel_failure(self, record, *, code: str, message: str) -> None: ) +def _is_within(path: Path, parent: Path) -> bool: + try: + path.resolve().relative_to(parent.resolve()) + except (ValueError, OSError): + return False + return True + + def _map_slurm_status(slurm_status: str) -> str: upper = slurm_status.upper() if upper in ("PENDING", "CONFIGURING"): diff --git a/quantui/backends/slurm_ingest.py b/quantui/backends/slurm_ingest.py index bfb30b5..409e97f 100644 --- a/quantui/backends/slurm_ingest.py +++ b/quantui/backends/slurm_ingest.py @@ -11,7 +11,7 @@ from types import SimpleNamespace from typing import Any -from .registry import JobRecord +from .registry import JobRecord, JobRegistry, parse_attempt_dir_name from .worker_payload import molecule_from_dict # Sidecar files the batch worker may write for Analysis replay parity. @@ -25,9 +25,38 @@ ) -def _copy_staging_sidecars(staging: Path, saved_dir: Path) -> None: +def slurm_provenance(record: JobRecord, staging: Path) -> dict[str, Any]: + """``result.json`` extras that mark a History entry as a SLURM run. + + M-JOBDIRS JD.5: stored under ``execution_backend`` / ``slurm`` so the + History list and card can tell SLURM results from local ones. Additive โ€” + local results and results ingested before this simply lack the keys. + For a job-dir record, *staging* is the attempt dir and its name gives the + attempt number and the SLURM job id that produced it (which, for a + hand-run ``sbatch``, differs from the record's tracked job). + """ + parsed = parse_attempt_dir_name(staging.name) if record.job_dir else None + return { + "execution_backend": "slurm", + "slurm": { + "job_id": parsed[1] if parsed else record.slurm_job_id, + "attempt": parsed[0] if parsed else None, + "job_dir": record.job_dir or record.staging_dir, + "attempt_dir": str(staging) if parsed else None, + "request_id": record.request_id, + }, + } + + +def _copy_staging_sidecars( + staging: Path, saved_dir: Path, payload: dict[str, Any] | None = None +) -> None: + # Attempt dirs may hold shareable files under descriptive names + # (result.json["artifact_names"], M-JOBDIRS JD.7); History keeps the + # fixed names its loaders expect. + renamed = (payload or {}).get("artifact_names") or {} for name in _STAGING_SIDECAR_FILES: - src = staging / name + src = staging / renamed.get(name, name) if src.exists(): shutil.copy2(src, saved_dir / name) @@ -97,7 +126,11 @@ def _copy_trajectory(staging: Path, saved_dir: Path, payload: dict[str, Any]) -> def _ingest_frequency( - staging: Path, payload: dict[str, Any], record: JobRecord, log_text: str + staging: Path, + payload: dict[str, Any], + record: JobRecord, + log_text: str, + extras: dict[str, Any], ) -> Path: from quantui import save_result from quantui.results_storage import save_molden @@ -109,11 +142,12 @@ def _ingest_frequency( pyscf_log=log_text, calc_type="frequency", spectra=spectra, + extras=extras, ) ir = spectra.get("ir") or {} freqs = ir.get("frequencies_cm1") displacements = ir.get("displacements") - _copy_staging_sidecars(staging, saved_dir) + _copy_staging_sidecars(staging, saved_dir, payload) if not (saved_dir / "result.molden").exists() and freqs and displacements: mol_block = spectra.get("molecule") or {} atoms = mol_block.get("atoms") or [] @@ -135,7 +169,10 @@ def _ingest_frequency( def _ingest_reorganization_energy( - payload: dict[str, Any], record: JobRecord, log_text: str + payload: dict[str, Any], + record: JobRecord, + log_text: str, + extras: dict[str, Any], ) -> Path: from quantui import save_result @@ -180,6 +217,7 @@ def _ingest_reorganization_energy( pyscf_log=log_text, calc_type="reorganization_energy", spectra=payload.get("spectra") or {}, + extras=extras, ) _finalize_history_entry(saved_dir) return saved_dir @@ -193,14 +231,21 @@ def _ingest_with_sidecars( if calc_type := payload.get("calc_type"): if calc_type in ("geometry_opt", "pes_scan"): _copy_trajectory(staging, saved_dir, payload) - _copy_staging_sidecars(staging, saved_dir) + _copy_staging_sidecars(staging, saved_dir, payload) _finalize_history_entry(saved_dir) return saved_dir -def ingest_staging_success(record: JobRecord, log_text: str = "") -> Path: - """Read ``result.json`` from staging and save under ``results/``.""" - staging = record.staging_path +def ingest_staging_success( + record: JobRecord, log_text: str = "", *, attempt_dir: Path | None = None +) -> Path: + """Read ``result.json`` from staging and save under ``results/``. + + *attempt_dir* selects a specific attempt (used for hand-run attempts the + record does not track); by default the record's tracked staging path. + """ + staging = attempt_dir if attempt_dir is not None else record.staging_path + extras = slurm_provenance(record, staging) result_path = staging / "result.json" if not result_path.exists(): raise FileNotFoundError(f"Missing staging result: {result_path}") @@ -211,10 +256,10 @@ def ingest_staging_success(record: JobRecord, log_text: str = "") -> Path: from quantui import save_result if calc_type == "frequency": - return _ingest_frequency(staging, payload, record, log_text) + return _ingest_frequency(staging, payload, record, log_text, extras) if calc_type == "reorganization_energy": - return _ingest_reorganization_energy(payload, record, log_text) + return _ingest_reorganization_energy(payload, record, log_text, extras) result = _basic_result(payload, record) spectra = payload.get("spectra") @@ -223,11 +268,72 @@ def ingest_staging_success(record: JobRecord, log_text: str = "") -> Path: pyscf_log=log_text, calc_type=calc_type, spectra=spectra if spectra is not None else {}, + extras=extras, ) return _ingest_with_sidecars(staging, saved_dir, payload) +def already_ingested(record: JobRecord, staging: Path) -> bool: + """True when the run in *staging* has already been saved to History. + + Job-dir records list ingested attempt dirs by name. Legacy records + predate that list; for them, a ``result_dir`` that points at an existing + History entry (the worker sets it to the staging dir; the app replaces + it with the saved dir after ingest) means the tracked run was ingested. + """ + if staging.name in record.ingested_attempts: + return True + if record.job_dir is None and record.result_dir: + saved = Path(record.result_dir) + return saved != staging and (saved / "result.json").exists() + return False + + +def uningested_attempts(record: JobRecord) -> list[Path]: + """Finished attempt dirs of a job-dir record not yet saved to History. + + ``result.json`` is written only when a run succeeds, so its presence + marks a finished, successful attempt (M-JOBDIRS: successes only). + Includes hand-run ``sbatch submit.slurm`` attempts the record never + tracked (JD.11). Legacy records return ``[]``: their single run is + ingested by the monitor / reconnect path, as before. + """ + if record.job_dir is None: + return [] + return [ + d + for d in record.attempt_dirs() + if (d / "result.json").exists() and not already_ingested(record, d) + ] + + +def ingest_attempt( + registry: JobRegistry, record: JobRecord, attempt_dir: Path | None = None +) -> Path: + """Save one attempt to History and record it as ingested. + + *attempt_dir* defaults to the record's tracked staging path. The + record's ``result_dir`` is updated only for the tracked attempt, so it + keeps pointing at the History entry of the run the record reports on. + """ + staging = attempt_dir if attempt_dir is not None else record.staging_path + log_path = staging / "live.log" + log_text = ( + log_path.read_text(encoding="utf-8", errors="replace") + if log_path.exists() + else "" + ) + saved_dir = ingest_staging_success(record, log_text, attempt_dir=staging) + tracked = staging == record.staging_path + registry.mark_ingested( + record.request_id, + staging.name, + result_dir=str(saved_dir) if tracked else None, + ) + return saved_dir + + def completion_summary_html(saved_dir: Path, payload: dict[str, Any]) -> str: calc_type = payload.get("calc_type", "single_point") energy = float(payload.get("energy_hartree", float("nan"))) diff --git a/quantui/backends/slurm_utils.py b/quantui/backends/slurm_utils.py index 834eccd..9dd5a92 100644 --- a/quantui/backends/slurm_utils.py +++ b/quantui/backends/slurm_utils.py @@ -21,6 +21,51 @@ class SlurmJobAccounting: elapsed: str = "" +# Short calc-type tags for job dir names (M-JOBDIRS JD.2). +_CALC_TYPE_TAGS = { + "single_point": "sp", + "geometry_opt": "opt", + "frequency": "freq", + "tddft": "tddft", + "nmr": "nmr", + "pes_scan": "scan", + "reorganization_energy": "reorg", +} + +_JOB_NAME_MAX_LEN = 80 +# SLURM itself accepts longer names, but squeue's default format column is +# narrow; the job dir keeps the full name. +SLURM_JOB_NAME_MAX_LEN = 40 + + +def sanitize_job_name(name: str) -> str: + """Make *name* safe as a job dir name and SLURM ``--job-name``. + + Characters other than letters, digits, ``-`` and ``_`` become ``_``; + leading/trailing separators are stripped. Returns ``""`` when nothing + usable is left. + """ + cleaned = re.sub(r"[^A-Za-z0-9_-]+", "_", name.strip()) + cleaned = re.sub(r"_+", "_", cleaned).strip("_-") + return cleaned[:_JOB_NAME_MAX_LEN].rstrip("_-") + + +def default_job_name(request: CalculationRequest) -> str: + """``