Skip to content

Repository files navigation

DOME Observatory

CI

A searchable database of AI/ML methods-paper metadata, live at observatory.dome-ml.org and one of the DOME-ML family of services alongside the DOME Registry. Over 800,000 publications have been screened and more than 360,000 classified as AI/ML methods papers, cross-linked to Europe PMC. Most classifications come from LLM processing, using a method validated against a hand-annotated expert benchmark before being scaled; over 6,000 records are human-curated or DOME Registry-confirmed, and the rest are not individually curator-reviewed.

Architecture

A monorepo with two independent apps, following the *-ui / *-ws convention used across the DOME-ML services.

Part Stack Role
observatory-ui/ Angular 22 (standalone components, signals), Bootstrap 5.3, Vitest Static SPA. No runtime configuration at all — no env vars, no environment.ts.
observatory-ws/ NestJS 11, Mongoose 8, Swagger, Jest Read-only REST API. The only process in this repo that opens a database connection. See observatory-ws/README.md.
schema/ Versioned JSON Schema + controlled vocabularies Build-time input to both apps. Releases live in schema/releases/; see schema/README.md.

There is no shared -core package between the two apps; overlapping types are duplicated on each side rather than linked.

The SPA uses relative URLs only, so it must be served from a domain root with the API at /api on the same origin.

Quick usage instructions

Local development

Build and run the local Docker compose file:

docker compose -f docker-compose-local.yml up -d --build
# xdg-open http://localhost:8080/

Production deployment

Build and deploy a production image to the production server:

docker compose -f docker-compose-build.yml build --push \
  observatory-ws \
  observatory-ui
docker -c $remote_context compose -f docker-compose-prod.yml up -d --pull always
# xdg-open https://observatory.dome-ml.org/

Running it locally

Docker is the recommended way to run the service and the one this page documents, in two modes: against the corpus database, or self-contained on a bundled 200-record sample. Both start the same two containers with the same command; the only difference is what MONGODB_URI points at.

Running the apps directly with npm is also supported — see working on the code for development, and deploying without Docker for a deployment in that shape.

Both need a .env first: cp .env.example .env.

# Mode A -- against the corpus database. Set MONGODB_URI to the corpus host.
# That host is reachable only from the hosting institution's network, so you need
# the University of Padua VPN. It is shared out of band, not through this repository.
docker compose -f docker-compose-local.yml up --build

# Mode B -- self-contained, no VPN and no network access to anything.
# Set MONGODB_URI=mongodb://mongo:27017 in .env first. Starts a throwaway MongoDB
# and seeds it with the sample fixture.
docker compose -f docker-compose-local.yml --profile offline up --build

Then open http://localhost:8080. /api/* returns 502 for the first 40–50 seconds while the backend warms its caches, which is expected.

# Same again without rebuilding the images -- much faster once built.
docker compose -f docker-compose-local.yml up
docker compose -f docker-compose-local.yml --profile offline up

# Run in the background, then follow the logs.
docker compose -f docker-compose-local.yml up --build -d
docker compose -f docker-compose-local.yml logs -f

# Stop and remove. Repeat the --profile offline flag, or the mongo containers are left behind.
docker compose -f docker-compose-local.yml down
docker compose -f docker-compose-local.yml --profile offline down -v

Prerequisites. Docker with a Compose v2 CLI — docker compose, space-separated. Node is not needed to run the service this way — only to work on the code, or to run it without Docker (see CONTRIBUTING.md). Both images build from the repo root as context, because each needs schema/, which sits outside its own directory.

Service Host port Container port
observatory-ui 8080 80 — the only browser-facing origin
observatory-ws 3000 3000 — for curling the API directly; the browser never uses it

Mode A still comes up if the database is unreachable: /api/health returns 200, /api/health/ready returns 503, and the frontend serves.

Mode B is what makes the repository runnable on any machine. The offline profile adds a throwaway MongoDB, then offline-database/seed.sh seeds it from the tracked fixture (observatory-ui/fixtures/sample-records.json) and stamps each record's record_modified and builds the same three indexes the real collection carries. It is a demo dataset, not a mirror of production, and it is ephemeral: down discards it and the next up reseeds. It defaults to mongo:7 for multi-architecture support while production runs MongoDB 4.2, so it is not a version-parity environment; override with MONGO_IMAGE for closer parity.

Two things will break the stack quietly if changed:

  • The service name observatory-ws and container port 3000 are a hard contract. observatory-ui/nginx.conf hardcodes observatory-ws:3000 as its proxy upstream. That pairing appears in three uncoupled places — the Compose service name, PORT in .env, and the nginx config — and nothing keeps them in sync.
  • A user-defined network is required. nginx resolves the backend at request time via Docker's embedded DNS, which only resolves service names on a user-defined network, not the default bridge.

API. Read-only REST under /api, browsable as Swagger at /api/docs. Routes, limits, the whole-corpus export loop and the environment variables are in observatory-ws/README.md.

Tests and CI

.github/workflows/ci.yml runs lint, tests, builds, both Docker images and schema validation on every pull request and every push to main. It reports on the commit; it does not block a push. The same gates run locally in seconds — see CONTRIBUTING.md.

Analytics

Page views are counted by the lab's self-hosted Matomo (matomo.biocomputingup.it): cookieless, IP-anonymised and hosted in the EU, so there is no consent banner and no Google Analytics. Browser tracking is on. Its one switch is MATOMO_SITE_ID in observatory-ui/src/app/core/analytics.config.ts, and the privacy page reads the same constant, so what it says cannot drift from what the site does. The disableCookies call must stay: without it the site would need a consent banner, versioned consent state and a withdrawal path.

API usage tracking (observatory-ws/src/analytics/matomo.interceptor.ts) is built and switched off. To turn it on, generate a Matomo auth token with tracking scope, set MATOMO_SITE_ID and MATOMO_TOKEN in observatory-ws's deployment environment (never a tracked file), restart, then curl an endpoint and confirm the hit reaches Matomo's real-time log with the caller's IP. The token is not optional: without it Matomo attributes every call to the server's own IP. Switch either half off by unsetting MATOMO_TOKEN and restarting, or setting MATOMO_SITE_ID to null and redeploying. The privacy page's wording follows the browser switch only, so turning the API half on means revising that page in the same change.

Support

contact@dome-ml.org reaches the team, for anything at all — questions about the data, collaborations, or a request to correct or remove a record.

dome-support@googlegroups.com is the Google Group, for release announcements and discussion between people using the corpus. Join from the group's page, or by email to dome-support+subscribe@googlegroups.com without a Google account.

For anything worth a public trail, open an issue. There is a short form for each kind of report, so you are not guessing what we need:

Template Use it for
A record is wrong or missing A paper misclassified, wrong metadata, or absent from the corpus.
Search isn't finding what I expect A search returning nothing, too little, or the wrong things.
Something on the site is broken A page that errors or renders wrong.
A question or suggestion Everything else.

The forms live in .github/ISSUE_TEMPLATE/. Blank issues stay enabled — the forms are there to save people guesswork, not to refuse anything that does not fit one of four shapes.

The Support page carries the same routes plus an FAQ covering the behaviours that surprise people most: authors are indexed as surname plus initials, search starts from the AI/ML positives rather than the whole screened corpus, words match whole (with their forms), and a word-beginning search (neuro*) takes a slower path through the corpus.

Key links and further information

  • observatory-ws/README.md — the API reference, environment variables and database expectations.
  • schema/README.md — the record schema and controlled vocabularies, and how releases are versioned.
  • Issues — what is still open, across both repositories. The recurring jobs themselves — refreshing the corpus and archiving each release to Zenodo — run from the sister repository's refresh-cycle skill.
  • CONTRIBUTING.md — how to propose a change, and the npm dev workflow, local gates and code conventions. · CODE_OF_CONDUCT.md
  • AGENTS.md — working conventions, and a record of the things that have gone wrong before. Read it before changing code, whether you are a person or an AI coding agent.
  • CITATION.cff — how to cite this.
  • LICENSE.md — CC BY 4.0 on the classification/enrichment layer this project adds. The underlying bibliographic metadata and abstracts come largely from Europe PMC and keep their own terms; full text is linked, never hosted.

About

DOME Observatory

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages