A searchable database of AI/ML methods-paper metadata, live at observatory.dome-ml.org and one of the DOME-ML family of services alongside the DOME Registry. Over 800,000 publications have been screened and more than 360,000 classified as AI/ML methods papers, cross-linked to Europe PMC. Most classifications come from LLM processing, using a method validated against a hand-annotated expert benchmark before being scaled; over 6,000 records are human-curated or DOME Registry-confirmed, and the rest are not individually curator-reviewed.
A monorepo with two independent apps, following the *-ui / *-ws convention used across the
DOME-ML services.
| Part | Stack | Role |
|---|---|---|
observatory-ui/ |
Angular 22 (standalone components, signals), Bootstrap 5.3, Vitest | Static SPA. No runtime configuration at all — no env vars, no environment.ts. |
observatory-ws/ |
NestJS 11, Mongoose 8, Swagger, Jest | Read-only REST API. The only process in this repo that opens a database connection. See observatory-ws/README.md. |
schema/ |
Versioned JSON Schema + controlled vocabularies | Build-time input to both apps. Releases live in schema/releases/; see schema/README.md. |
There is no shared -core package between the two apps; overlapping types are duplicated on each
side rather than linked.
The SPA uses relative URLs only, so it must be served from a domain root with the API at /api on
the same origin.
Build and run the local Docker compose file:
docker compose -f docker-compose-local.yml up -d --build
# xdg-open http://localhost:8080/Build and deploy a production image to the production server:
docker compose -f docker-compose-build.yml build --push \
observatory-ws \
observatory-ui
docker -c $remote_context compose -f docker-compose-prod.yml up -d --pull always
# xdg-open https://observatory.dome-ml.org/Docker is the recommended way to run the service and the one this page documents, in two modes:
against the corpus database, or self-contained on a bundled 200-record sample. Both start the same
two containers with the same command; the only difference is what MONGODB_URI points at.
Running the apps directly with npm is also supported — see working on the code for development, and deploying without Docker for a deployment in that shape.
Both need a .env first: cp .env.example .env.
# Mode A -- against the corpus database. Set MONGODB_URI to the corpus host.
# That host is reachable only from the hosting institution's network, so you need
# the University of Padua VPN. It is shared out of band, not through this repository.
docker compose -f docker-compose-local.yml up --build
# Mode B -- self-contained, no VPN and no network access to anything.
# Set MONGODB_URI=mongodb://mongo:27017 in .env first. Starts a throwaway MongoDB
# and seeds it with the sample fixture.
docker compose -f docker-compose-local.yml --profile offline up --buildThen open http://localhost:8080. /api/* returns 502 for the first 40–50 seconds while the
backend warms its caches, which is expected.
# Same again without rebuilding the images -- much faster once built.
docker compose -f docker-compose-local.yml up
docker compose -f docker-compose-local.yml --profile offline up
# Run in the background, then follow the logs.
docker compose -f docker-compose-local.yml up --build -d
docker compose -f docker-compose-local.yml logs -f
# Stop and remove. Repeat the --profile offline flag, or the mongo containers are left behind.
docker compose -f docker-compose-local.yml down
docker compose -f docker-compose-local.yml --profile offline down -vPrerequisites. Docker with a Compose v2 CLI — docker compose, space-separated. Node is not
needed to run the service this way — only to work on the code, or to run it without Docker (see
CONTRIBUTING.md). Both images build from the repo root as
context, because each needs schema/, which sits outside its own directory.
| Service | Host port | Container port |
|---|---|---|
observatory-ui |
8080 | 80 — the only browser-facing origin |
observatory-ws |
3000 | 3000 — for curling the API directly; the browser never uses it |
Mode A still comes up if the database is unreachable: /api/health returns 200,
/api/health/ready returns 503, and the frontend serves.
Mode B is what makes the repository runnable on any machine. The offline profile adds a throwaway
MongoDB, then offline-database/seed.sh seeds it from the tracked
fixture (observatory-ui/fixtures/sample-records.json)
and stamps each record's record_modified and builds the same three indexes the real collection carries. It is a demo dataset, not a mirror of
production, and it is ephemeral: down discards it and the next up reseeds. It defaults to
mongo:7 for multi-architecture support while production runs MongoDB 4.2, so it is not a
version-parity environment; override with MONGO_IMAGE for closer parity.
Two things will break the stack quietly if changed:
- The service name
observatory-wsand container port3000are a hard contract.observatory-ui/nginx.confhardcodesobservatory-ws:3000as its proxy upstream. That pairing appears in three uncoupled places — the Compose service name,PORTin.env, and the nginx config — and nothing keeps them in sync. - A user-defined network is required. nginx resolves the backend at request time via Docker's
embedded DNS, which only resolves service names on a user-defined network, not the default
bridge.
API. Read-only REST under /api, browsable as Swagger at
/api/docs. Routes, limits, the whole-corpus export
loop and the environment variables are in observatory-ws/README.md.
.github/workflows/ci.yml runs lint, tests, builds, both Docker images
and schema validation on every pull request and every push to main. It reports on the commit; it
does not block a push. The same gates run locally in seconds — see
CONTRIBUTING.md.
Page views are counted by the lab's self-hosted Matomo (matomo.biocomputingup.it): cookieless,
IP-anonymised and hosted in the EU, so there is no consent banner and no Google Analytics. Browser
tracking is on. Its one switch is MATOMO_SITE_ID in
observatory-ui/src/app/core/analytics.config.ts,
and the privacy page reads the same constant, so what it says cannot drift from what the site does.
The disableCookies call must stay: without it the site would need a consent banner, versioned
consent state and a withdrawal path.
API usage tracking (observatory-ws/src/analytics/matomo.interceptor.ts) is built and switched
off. To turn it on, generate a Matomo auth token with tracking scope, set MATOMO_SITE_ID and
MATOMO_TOKEN in observatory-ws's deployment environment (never a tracked file), restart, then
curl an endpoint and confirm the hit reaches Matomo's real-time log with the caller's IP. The
token is not optional: without it Matomo attributes every call to the server's own IP. Switch
either half off by unsetting MATOMO_TOKEN and restarting, or setting MATOMO_SITE_ID to null
and redeploying. The privacy page's wording follows the browser switch only, so turning the API
half on means revising that page in the same change.
contact@dome-ml.org reaches the team, for anything at all — questions about the data, collaborations, or a request to correct or remove a record.
dome-support@googlegroups.com is the Google Group, for release announcements and discussion between people using the corpus. Join from the group's page, or by email to dome-support+subscribe@googlegroups.com without a Google account.
For anything worth a public trail, open an issue. There is a short form for each kind of report, so you are not guessing what we need:
| Template | Use it for |
|---|---|
| A record is wrong or missing | A paper misclassified, wrong metadata, or absent from the corpus. |
| Search isn't finding what I expect | A search returning nothing, too little, or the wrong things. |
| Something on the site is broken | A page that errors or renders wrong. |
| A question or suggestion | Everything else. |
The forms live in .github/ISSUE_TEMPLATE/. Blank issues stay enabled —
the forms are there to save people guesswork, not to refuse anything that does not fit one of four
shapes.
The Support page carries the same routes plus an
FAQ covering the behaviours that surprise people most: authors are indexed as surname plus initials,
search starts from the AI/ML positives rather than the whole screened corpus, words match whole
(with their forms), and a word-beginning search (neuro*) takes a slower path through the corpus.
observatory-ws/README.md— the API reference, environment variables and database expectations.schema/README.md— the record schema and controlled vocabularies, and how releases are versioned.- Issues — what is still open,
across both repositories. The recurring jobs themselves — refreshing the corpus and archiving
each release to Zenodo — run from the sister repository's
refresh-cycleskill. CONTRIBUTING.md— how to propose a change, and the npm dev workflow, local gates and code conventions. ·CODE_OF_CONDUCT.mdAGENTS.md— working conventions, and a record of the things that have gone wrong before. Read it before changing code, whether you are a person or an AI coding agent.CITATION.cff— how to cite this.LICENSE.md— CC BY 4.0 on the classification/enrichment layer this project adds. The underlying bibliographic metadata and abstracts come largely from Europe PMC and keep their own terms; full text is linked, never hosted.