Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions content/api/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,9 +58,11 @@ In these commands:

## Publishing dataset

{{< refItem name="ota dataset [--file <filename>]" description="Export the versions dataset into a ZIP file. The dataset title is defined in the configuration." example="npx ota dataset --file dataset.zip" />}}
{{< refItem name="ota dataset [--file <filename>]" description="Export the versions dataset into a ZIP file, stored in the directory defined by `dataset.storagePath` in the configuration (`./data/datasets` by default) alongside a `metadata.json` file describing it. The file name defaults to the dataset title, defined in the configuration, followed by the current date. Only the latest dataset is kept: previous archives in that directory are deleted." example="npx ota dataset --file dataset.zip" />}}

To export the dataset into a ZIP file and publish it to configured platforms (GitHub releases, GitLab releases, and/or data.gouv.fr):
The latest dataset is exposed by the [Collection API]({{< relref "api/collection" >}}).

To also publish the dataset to configured platforms (GitHub releases, GitLab releases, and/or data.gouv.fr):

{{< refItem name="ota dataset --publish [--file <filename>]" description="Export and publish dataset to all configured platforms" example="npx ota dataset --publish" />}}

Expand All @@ -74,11 +76,11 @@ These environment variables can be defined in a [`.env` file]({{< relref "collec

> **Note**: If both GitHub and GitLab tokens are configured, GitHub takes precedence. data.gouv.fr can be used alongside either GitHub or GitLab.

To export, publish the dataset and remove the local copy that was created after it has been uploaded:
To export, and optionally publish, the dataset on the schedule defined by `dataset.publishingSchedule` in the configuration:

{{< refItem name="ota dataset --publish --remove-local-copy [--file <filename>]" description="Export, publish dataset and remove local copy after upload" example="npx ota dataset --publish --remove-local-copy" />}}
{{< refItem name="ota dataset --schedule [--publish] [--file <filename>]" description="Export, and optionally publish, the dataset on the schedule defined in the configuration" example="npx ota dataset --schedule --publish" />}}

{{< refItem name="ota dataset --schedule [--file <filename>]" description="Schedule export, publishing and local copy removal" example="npx ota dataset --schedule --publish --remove-local-copy" />}}
> **Note**: The `--remove-local-copy` option was removed in engine v16, as the local copy is now the reference dataset served by the Collection API. Remove it from the `dataset:schedule` script of the collection `package.json`, otherwise the command fails with an unknown option error.

## Exposing the collection API

Expand Down
2 changes: 1 addition & 1 deletion content/api/collection.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ weight: 2

As Open Terms Archive is decentralised, each instance embarks its own API. The documentation relevant to the specific version of the engine on that instance is provided on that instance itself.

The Collection API exposes data over HTTP: versioning data in JSON format and [Atom feeds]({{< relref "/terms/how-to/be-notified" >}}) for subscription. Its [OpenAPI](https://swagger.io/specification/) specification can be found at `http://localhost:<port>/<basePath>/<API version>/docs`.
The Collection API exposes data over HTTP: versioning data in JSON format, [Atom feeds]({{< relref "/terms/how-to/be-notified" >}}) for subscription, and the metadata and archive of the latest [dataset]({{< relref "/api/cli#publishing-dataset" >}}) generated on the instance. Its [OpenAPI](https://swagger.io/specification/) specification can be found at `http://localhost:<port>/<basePath>/<API version>/docs`.

That endpoint exposes both the OpenAPI specification if the requested `Content-Type` is JSON, and a Swagger UI for visual and interactive documentation otherwise.

Expand Down
8 changes: 4 additions & 4 deletions content/collections/how-to/publish-to-datagouv.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ In your collection's configuration file (e.g., `config/production.json`), add th
}
```

### 3. Configure for testing (optional)
### 2. Configure for testing (optional)

If you want to test with the demo environment first, add `useDemo`:

Expand All @@ -84,7 +84,7 @@ If you want to test with the demo environment first, add `useDemo`:
}
```

### 4. Set the API key
### 3. Set the API key

Create a `.env` file at the root of your collection repository (if it doesn't already exist) and add your data.gouv.fr API key:

Expand All @@ -107,10 +107,10 @@ This will create and publish a dataset to data.gouv.fr. Check the output to veri
To automatically publish datasets on a schedule, use the `--schedule` flag:

```bash
npx ota dataset --schedule --publish --remove-local-copy
npx ota dataset --schedule --publish
```

This will publish datasets according to the schedule defined in your configuration (by default, every Monday at 8:30 AM).
This will publish datasets according to the schedule defined in your configuration (by default, every Monday at 8:30 AM). The latest generated dataset is also kept locally and served by the [Collection API]({{< relref "api/collection" >}}).

## Publishing to multiple platforms

Expand Down
9 changes: 8 additions & 1 deletion content/collections/reference/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,7 +166,7 @@ The reporter section manages how issues are reported when terms content is inacc

### Dataset

The dataset section configures how datasets are published. Datasets can be published to GitHub releases, GitLab releases, and/or data.gouv.fr. If both GitHub and GitLab tokens are configured, GitHub takes precedence.
The dataset section configures how datasets are generated, stored and published. The latest generated dataset is stored locally and exposed through the [Collection API]({{< relref "api/collection" >}}). It can also be published to GitHub releases, GitLab releases, and/or data.gouv.fr. If both GitHub and GitLab tokens are configured, GitHub takes precedence.

{{< refItem
name="dataset.title"
Expand All @@ -191,6 +191,13 @@ The dataset section configures how datasets are published. Datasets can be publi
default="30 8 * * MON"
/>}}

{{< refItem
name="dataset.storagePath"
type="string"
description="Path to the directory where the latest generated dataset archive is stored, alongside a `metadata.json` file describing it. This directory is managed by the engine, which deletes any `.zip` archive it does not reference at each generation, so it must not be shared with other files. Its content is exposed by the Collection API."
default="./data/datasets"
/>}}

#### data.gouv.fr publishing

The data.gouv.fr section configures publishing to the French government's open data platform. Either `datasetId` or `organizationIdOrSlug` must be configured.
Expand Down
2 changes: 2 additions & 0 deletions content/deployment/reference/server-specifications.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,8 @@ Storage usage typically grows at a rate of 0.5 MB per tracked terms per month o
- Services with large legal teams and heavy website layouts: ~1 MB per terms per month
- Smaller services: ~0.1-0.3 MB per terms per month

If datasets are generated with `ota dataset`, the latest archive is kept in the directory defined by [`dataset.storagePath`]({{< relref "collections/reference/configuration#ref-dataset-storagepath" >}}). Plan enough space for two archives, as the new one is written before the previous one is deleted.

### Examples

- Tracking 5 very large social media platforms on their 5 most common terms types (such as Terms of Service, Privacy Policy, Trackers Policy, Developer Agreement, Community Guidelines) would require approximately 300 MB of additional storage per year.
Expand Down
2 changes: 2 additions & 0 deletions content/terms/how-to/be-notified.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,8 @@ Subscribe to changes across the entire collection: `{collection-api-endpoint}/fe

Subscribe to changes for all terms of a specific service: `{collection-api-endpoint}/feed/{serviceId}`

The service ID is case sensitive: it is the name of the service declaration file without the extension.

> For example, for all terms of GitHub in the Demo collection: `http://162.19.74.224/collection-api/v1/feed/GitHub`

## For one terms type of a service
Expand Down
2 changes: 1 addition & 1 deletion content/terms/how-to/manage-custom-terms-type.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ If you need a faster solution for production use, you can fork the terms-types r
```json
{
"dependencies": {
"@opentermsarchive/engine": "^14.0.0",
"@opentermsarchive/engine": "^16.0.0",
"@opentermsarchive/terms-types": "<your-organization-or-username>/terms-types#main"
}
}
Expand Down
Loading