This repository contains OpenGeoMetadata Aardvark records for the U.S. Geological Survey's Historical Topographic Map Collection (HTMC): scans of all scales and editions of the topographic maps USGS published from 1884 to 2006, about 186,000 maps.
The records are generated from USGS's public inventory of the collection, along with index maps of each state's maps at each scale. This repository is not an official USGS product.
This metadata is provided under the CC0 v1.0 License, which allows for free use and redistribution without restrictions. The maps themselves are in the U.S. public domain.
metadata-aardvark/
htmc/
usgs-htmc.json the collection record
314/
usgs-htmc-314262.json one record per map, named by its id
...
index-maps/
usgs-htmc-index-az-48000.json one record per index map
...
index-maps/
htmc/
usgs-htmc-index-az-48000.geojson the index map itself
...
withdrawn.json records that have left the repository
Each map's id is usgs-htmc- followed by its USGS scan ID. Records are grouped in directories by the thousands of their scan ID (usgs-htmc-314262 is in 314/), which keeps every directory under GitHub's 1,000-entry listing limit. A record's path follows from its id, so there is no layers.json.
- Version: OGM Aardvark, with no custom fields.
- Other formats: each record links the map's original FGDC metadata, which USGS hosts; it isn't copied here.
- Updates: a GitHub Actions workflow rebuilds every record daily from USGS's nightly inventory. Files only change when USGS's metadata for a map does, so an unchanged inventory produces no commit.
- Validation: the tests in
test/run before every harvest.
USGS publishes an inventory of every map in the collection nightly, at historicaltopo.zip. The zip holds:
historicaltopo.csv: one row per map, 43 columns.historicaltopo_readme.txtin the same zip defines them.historicaltopo.txt: a report giving the number of maps and the date.
| Aardvark field | Inventory columns | Notes |
|---|---|---|
id |
scan_id |
usgs-htmc-314262 |
dct_title_s |
map_scale, map_name, primary_state, date_on_map, imprint_year |
USGS's own title, as ScienceBase and The National Map show it, plus (printed 1918) when the printing year differs from the edition year. Many editions survive in several printings. |
dct_alternative_sm |
map_name |
The printed map name |
dct_description_sm |
date_on_map, map_scale, grid_size, publishers, the production years, print_releases, variations, sheet size, projection_printed, hdatum_printed |
Written out as two paragraphs |
dct_creator_sm, dct_publisher_sm |
publishers |
Library of Congress name headings, e.g. Geological Survey (U.S.). Unknown is left out. |
dct_issued_s, dct_temporal_sm, gbl_indexYear_im |
date_on_map |
|
gbl_dateRange_drsim |
all the year columns | From the survey to this copy's printing |
dct_spatial_sm |
state_list, county_list, primary_state |
See Places |
dct_language_sm |
languages |
eng, spa |
dct_identifier_sm |
sciencebase_url, product_inventory_uuid |
|
dcat_keyword_sm |
grid_size |
e.g. Historical Topographic Maps, 7.5 x 7.5 minute |
locn_geometry |
geom_wkt |
|
dcat_bbox, dcat_centroid |
westbc, eastbc, northbc, southbc |
|
pcdm_memberOf_sm |
Always usgs-htmc, the collection record |
|
gbl_mdModified_dt |
metadata_date |
|
dct_references_s |
geotiff_url, product_url, metadata_url, sciencebase_url, thumbnail_url |
COG, GeoPDF and GeoTIFF downloads, FGDC metadata, ScienceBase landing page, thumbnail |
The remaining fields are the same for every map:
gbl_resourceClass_smisMaps.gbl_resourceType_smisTopographic maps.dct_format_sisGeoTIFF.dct_rights_smcarries USGS's public-domain statement.dct_license_smis the Creative Commons Public Domain Mark.
dct_spatial_sm lists every state, territory, province or Mexican state in state_list. For a map within a single U.S. state or territory, it also lists that map's counties, qualified by the state: Coconino County, Arizona. The inventory gives bare county names, so their type is inferred:
- Louisiana: parishes.
- Alaska: census areas (the inventory's
(CA)suffix), or bare names for boroughs and municipalities. - Independent cities: the
(city)suffix, listed as the city itself. - Territories: municipios and districts, without a suffix.
Maps covering several states get no counties, because the inventory doesn't say which state each county is in.
Some maps were scanned more than once, from different printed copies, and the inventory describes the copies identically. Their rows match in every column except the ones identifying the copy:
scan_idandproduct_inventory_uuid;- the sheet size;
- the file names, sizes and URLs;
metadata_date.
The harvester uses the latest scan ID to generate the record, and the others are skipped. If USGS later adds a newer scan of a map that already has a record, the newer scan replaces it, and the old record is withdrawn as a duplicate.
There is an OpenIndexMaps index map for each scale of map covering each U.S. state or territory. Each is a GeoJSON file of map outlines, and has a record of its own, e.g. usgs-htmc-index-ca-24000 for California's 1:24,000-scale maps.
- Maps: every map with a record is outlined in the index map of each state or territory in its
state_list, so maps along a border are in both states' index maps. Canada and Mexico get no index maps. - Sheets: index maps aren't created for series that have only one sheet.
- Editions: each map has its own outline, as the specification recommends, so a sheet's editions and printings are stacked on the same outline. They're in date order, so viewers draw the latest on top.
- Properties: each outline's
labelis the map's name and edition year, plus the printing year when it differs:Bright Angel 1906 (printed 1918). It also has the map'stitle,datePub,dateSurvey,dateReprnt,scale,publisherandavailable, and its record's id asrecId.websiteUrl,downloadandthumbUrllink to its ScienceBase page, GeoTIFF and thumbnail.
When a record leaves the repository, it is deleted and logged in withdrawn.json.
- Duplicate: if the map is still in the inventory but a later scan now describes it identically, the entry's reason is
duplicate, withis_replaced_bynaming the later scan's record. - Quality: if the map is still in the inventory but no longer has a GeoTIFF, the reason is
quality. A twin that has a GeoTIFF is named inis_replaced_by. - Superseded: if the map's product (
product_inventory_uuid) reappears under a new scan ID, USGS rescanned it. The entry's reason issuperseded, withis_replaced_bynaming the new record. - Upstream-removed: otherwise, the reason is
upstream-removed. - Index maps: an index map is withdrawn, also as
upstream-removed, once fewer than two of its sheets remain. - Republished: entries are only removed when their map or index map comes back.
The harvester stops without changing anything in two cases:
- The CSV's row count disagrees with the report, as a truncated download would cause.
- The run would remove more than 2% of all records. Set
FORCE=1if this is intentional.
The harvester needs Ruby 4.0 (the csv and minitest gems ship with it) and unzip.
ruby harvester.rbThe harvester downloads the inventory to tmp/, rebuilds every record and index map, and prints what changed. It takes about a minute. These environment variables change how it runs:
SOURCE=path/to/historicaltopo.zipuses a local copy of the inventory, or a directory holding its CSV and report, instead of downloading it.DRY_RUN=1reports what would change without writing anything.FORCE=1allows a run that would withdraw more than 2% of the records.INDEX_MAP_URL=http://localhost:8000/index-maps/htmcchanges where index map records say their GeoJSON is, to try them in a local GeoBlacklight before the files are on GitHub. Don't commit records written this way.
To run the tests:
for test in test/*_test.rb; do ruby "$test"; doneFor problems with the records, open an issue in this repository. Every record is regenerated from USGS's inventory daily, so edits to the files would be overwritten:
- How fields are mapped: change
mapper.rbandplaces.rb, orindex_map.rbfor how maps are grouped into index maps. - The inventory's own data: corrections have to be made by USGS.