diff --git a/content/api/cli.md b/content/api/cli.md index d91078d2..13b318ed 100644 --- a/content/api/cli.md +++ b/content/api/cli.md @@ -58,9 +58,11 @@ In these commands: ## Publishing dataset -{{< refItem name="ota dataset [--file ]" description="Export the versions dataset into a ZIP file. The dataset title is defined in the configuration." example="npx ota dataset --file dataset.zip" />}} +{{< refItem name="ota dataset [--file ]" description="Export the versions dataset into a ZIP file, stored in the directory defined by `dataset.storagePath` in the configuration (`./data/datasets` by default) alongside a `metadata.json` file describing it. The file name defaults to the dataset title, defined in the configuration, followed by the current date. Only the latest dataset is kept: previous archives in that directory are deleted." example="npx ota dataset --file dataset.zip" />}} -To export the dataset into a ZIP file and publish it to configured platforms (GitHub releases, GitLab releases, and/or data.gouv.fr): +The latest dataset is exposed by the [Collection API]({{< relref "api/collection" >}}). + +To also publish the dataset to configured platforms (GitHub releases, GitLab releases, and/or data.gouv.fr): {{< refItem name="ota dataset --publish [--file ]" description="Export and publish dataset to all configured platforms" example="npx ota dataset --publish" />}} @@ -74,11 +76,11 @@ These environment variables can be defined in a [`.env` file]({{< relref "collec > **Note**: If both GitHub and GitLab tokens are configured, GitHub takes precedence. data.gouv.fr can be used alongside either GitHub or GitLab. -To export, publish the dataset and remove the local copy that was created after it has been uploaded: +To export, and optionally publish, the dataset on the schedule defined by `dataset.publishingSchedule` in the configuration: -{{< refItem name="ota dataset --publish --remove-local-copy [--file ]" description="Export, publish dataset and remove local copy after upload" example="npx ota dataset --publish --remove-local-copy" />}} +{{< refItem name="ota dataset --schedule [--publish] [--file ]" description="Export, and optionally publish, the dataset on the schedule defined in the configuration" example="npx ota dataset --schedule --publish" />}} -{{< refItem name="ota dataset --schedule [--file ]" description="Schedule export, publishing and local copy removal" example="npx ota dataset --schedule --publish --remove-local-copy" />}} +> **Note**: The `--remove-local-copy` option was removed in engine v16, as the local copy is now the reference dataset served by the Collection API. Remove it from the `dataset:schedule` script of the collection `package.json`, otherwise the command fails with an unknown option error. ## Exposing the collection API diff --git a/content/api/collection.md b/content/api/collection.md index 5d052343..6d7a70b4 100644 --- a/content/api/collection.md +++ b/content/api/collection.md @@ -7,7 +7,7 @@ weight: 2 As Open Terms Archive is decentralised, each instance embarks its own API. The documentation relevant to the specific version of the engine on that instance is provided on that instance itself. -The Collection API exposes data over HTTP: versioning data in JSON format and [Atom feeds]({{< relref "/terms/how-to/be-notified" >}}) for subscription. Its [OpenAPI](https://swagger.io/specification/) specification can be found at `http://localhost:///docs`. +The Collection API exposes data over HTTP: versioning data in JSON format, [Atom feeds]({{< relref "/terms/how-to/be-notified" >}}) for subscription, and the metadata and archive of the latest [dataset]({{< relref "/api/cli#publishing-dataset" >}}) generated on the instance. Its [OpenAPI](https://swagger.io/specification/) specification can be found at `http://localhost:///docs`. That endpoint exposes both the OpenAPI specification if the requested `Content-Type` is JSON, and a Swagger UI for visual and interactive documentation otherwise. diff --git a/content/collections/how-to/publish-to-datagouv.md b/content/collections/how-to/publish-to-datagouv.md index 74e4f16e..c2b5aa90 100644 --- a/content/collections/how-to/publish-to-datagouv.md +++ b/content/collections/how-to/publish-to-datagouv.md @@ -68,7 +68,7 @@ In your collection's configuration file (e.g., `config/production.json`), add th } ``` -### 3. Configure for testing (optional) +### 2. Configure for testing (optional) If you want to test with the demo environment first, add `useDemo`: @@ -84,7 +84,7 @@ If you want to test with the demo environment first, add `useDemo`: } ``` -### 4. Set the API key +### 3. Set the API key Create a `.env` file at the root of your collection repository (if it doesn't already exist) and add your data.gouv.fr API key: @@ -107,10 +107,10 @@ This will create and publish a dataset to data.gouv.fr. Check the output to veri To automatically publish datasets on a schedule, use the `--schedule` flag: ```bash -npx ota dataset --schedule --publish --remove-local-copy +npx ota dataset --schedule --publish ``` -This will publish datasets according to the schedule defined in your configuration (by default, every Monday at 8:30 AM). +This will publish datasets according to the schedule defined in your configuration (by default, every Monday at 8:30 AM). The latest generated dataset is also kept locally and served by the [Collection API]({{< relref "api/collection" >}}). ## Publishing to multiple platforms diff --git a/content/collections/reference/configuration.md b/content/collections/reference/configuration.md index 9f7a6e21..39bd209b 100644 --- a/content/collections/reference/configuration.md +++ b/content/collections/reference/configuration.md @@ -166,7 +166,7 @@ The reporter section manages how issues are reported when terms content is inacc ### Dataset -The dataset section configures how datasets are published. Datasets can be published to GitHub releases, GitLab releases, and/or data.gouv.fr. If both GitHub and GitLab tokens are configured, GitHub takes precedence. +The dataset section configures how datasets are generated, stored and published. The latest generated dataset is stored locally and exposed through the [Collection API]({{< relref "api/collection" >}}). It can also be published to GitHub releases, GitLab releases, and/or data.gouv.fr. If both GitHub and GitLab tokens are configured, GitHub takes precedence. {{< refItem name="dataset.title" @@ -191,6 +191,13 @@ The dataset section configures how datasets are published. Datasets can be publi default="30 8 * * MON" />}} +{{< refItem + name="dataset.storagePath" + type="string" + description="Path to the directory where the latest generated dataset archive is stored, alongside a `metadata.json` file describing it. This directory is managed by the engine, which deletes any `.zip` archive it does not reference at each generation, so it must not be shared with other files. Its content is exposed by the Collection API." + default="./data/datasets" +/>}} + #### data.gouv.fr publishing The data.gouv.fr section configures publishing to the French government's open data platform. Either `datasetId` or `organizationIdOrSlug` must be configured. diff --git a/content/deployment/reference/server-specifications.md b/content/deployment/reference/server-specifications.md index 4c2ed1e7..13d9974c 100644 --- a/content/deployment/reference/server-specifications.md +++ b/content/deployment/reference/server-specifications.md @@ -22,6 +22,8 @@ Storage usage typically grows at a rate of 0.5 MB per tracked terms per month o - Services with large legal teams and heavy website layouts: ~1 MB per terms per month - Smaller services: ~0.1-0.3 MB per terms per month +If datasets are generated with `ota dataset`, the latest archive is kept in the directory defined by [`dataset.storagePath`]({{< relref "collections/reference/configuration#ref-dataset-storagepath" >}}). Plan enough space for two archives, as the new one is written before the previous one is deleted. + ### Examples - Tracking 5 very large social media platforms on their 5 most common terms types (such as Terms of Service, Privacy Policy, Trackers Policy, Developer Agreement, Community Guidelines) would require approximately 300 MB of additional storage per year. diff --git a/content/terms/how-to/be-notified.md b/content/terms/how-to/be-notified.md index 38ba3143..910658fc 100644 --- a/content/terms/how-to/be-notified.md +++ b/content/terms/how-to/be-notified.md @@ -23,6 +23,8 @@ Subscribe to changes across the entire collection: `{collection-api-endpoint}/fe Subscribe to changes for all terms of a specific service: `{collection-api-endpoint}/feed/{serviceId}` +The service ID is case sensitive: it is the name of the service declaration file without the extension. + > For example, for all terms of GitHub in the Demo collection: `http://162.19.74.224/collection-api/v1/feed/GitHub` ## For one terms type of a service diff --git a/content/terms/how-to/manage-custom-terms-type.md b/content/terms/how-to/manage-custom-terms-type.md index ff80b0f7..25a95d42 100644 --- a/content/terms/how-to/manage-custom-terms-type.md +++ b/content/terms/how-to/manage-custom-terms-type.md @@ -46,7 +46,7 @@ If you need a faster solution for production use, you can fork the terms-types r ```json { "dependencies": { - "@opentermsarchive/engine": "^14.0.0", + "@opentermsarchive/engine": "^16.0.0", "@opentermsarchive/terms-types": "/terms-types#main" } }