Summary
The documented worker command — uv run celery -A scripts.run_worker:celery worker -l info — doesn't load host/.env, so Settings() falls back to the SQLite default and the model __table_args__ keep their per-module schema qualifiers (datasets.datasets_dataset etc.). When a task tries to execute against a real Postgres database (because the broker is reachable but the worker's settings disagree with the web process), every query fails with either:
sqlite3.OperationalError: unknown function: now() (when the worker is actually using SQLite), or
psycopg2.errors.UndefinedTable: relation "datasets.datasets_dataset" does not exist (when the worker is wired to Postgres but the model still has the schema-qualified table name from _default_provider).
Repro
- Set
SM_DATABASE_URL=postgresql+asyncpg://... in host/.env (which uvicorn picks up automatically through pydantic-settings via cwd).
- Start the worker the way
host/Makefile's `worker` target does: `uv run celery -A scripts.run_worker:celery worker -l info`.
- Enqueue a real task (e.g. `lacowiki_datasets.convert_dataset`).
- Worker fails to write `background_tasks_task_execution` rows because either the env var is missing or the schema provider doesn't match the dev migration.
The dev-only helper host/scripts/reset_admin_dev.py works around this with:
import simple_module_db.base as _smdb_base
_smdb_base._default_provider = lambda: _smdb_base.DatabaseProvider.SQLITE
…which is itself a sign the seam is too sharp for normal app code.
Expected
scripts/run_worker.py (or the build_celery factory it calls) should:
- Load
.env from the host directory before instantiating BackgroundTasksSettings/Settings.
- Apply the same provider/schema convention the web process applies, so model
__table_args__ line up with the actual database the worker reaches.
Workaround
Wrap the worker entrypoint:
# scripts/run_worker.py
from dotenv import load_dotenv
from pathlib import Path
load_dotenv(Path(__file__).resolve().parent.parent / ".env")
import simple_module_db.base as _smdb_base
_smdb_base._default_provider = lambda: _smdb_base.DatabaseProvider.SQLITE # match dev migrations
from background_tasks.celery_app import build_celery
from background_tasks.settings import BackgroundTasksSettings
celery = build_celery(BackgroundTasksSettings())
Acceptance
- A fresh
sm new --preset full host's worker command runs against the same database the web process uses, with no extra wiring in scripts/.
Summary
The documented worker command —
uv run celery -A scripts.run_worker:celery worker -l info— doesn't loadhost/.env, soSettings()falls back to the SQLite default and the model__table_args__keep their per-module schema qualifiers (datasets.datasets_datasetetc.). When a task tries to execute against a real Postgres database (because the broker is reachable but the worker's settings disagree with the web process), every query fails with either:sqlite3.OperationalError: unknown function: now()(when the worker is actually using SQLite), orpsycopg2.errors.UndefinedTable: relation "datasets.datasets_dataset" does not exist(when the worker is wired to Postgres but the model still has the schema-qualified table name from_default_provider).Repro
SM_DATABASE_URL=postgresql+asyncpg://...inhost/.env(whichuvicornpicks up automatically through pydantic-settings via cwd).host/Makefile's `worker` target does: `uv run celery -A scripts.run_worker:celery worker -l info`.The dev-only helper
host/scripts/reset_admin_dev.pyworks around this with:…which is itself a sign the seam is too sharp for normal app code.
Expected
scripts/run_worker.py(or thebuild_celeryfactory it calls) should:.envfrom the host directory before instantiatingBackgroundTasksSettings/Settings.__table_args__line up with the actual database the worker reaches.Workaround
Wrap the worker entrypoint:
Acceptance
sm new --preset fullhost's worker command runs against the same database the web process uses, with no extra wiring inscripts/.