Following up on the upstream query. Here is exactly what the app runs now, after the fixes in #55
Question: What facilities are upstream from PFOS samples? Block A = Facilities (no filters, no region), Block C = Samples, PFOS, York County ME.
Block C is Constrained
What we run now
Shared prefixes for every query below:
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX owl: <http://www.w3.org/2002/07/owl#>
PREFIX xsd: <http://www.w3.org/2001/XMLSchema#>
PREFIX geo: <http://www.opengis.net/ont/geosparql#>
PREFIX spatial: <http://purl.org/spatialai/spatial/spatial-full#>
PREFIX kwg-ont: <http://stko-kwg.geog.ucsb.edu/lod/ontology/>
PREFIX kwgr: <http://stko-kwg.geog.ucsb.edu/lod/resource/>
PREFIX hyf: <https://www.opengis.net/def/schema/hy_features/hyf/>
PREFIX coso: <http://w3id.org/coso/v1/contaminoso#>
PREFIX fio: <http://w3id.org/fio/v1/fio#>
PREFIX naics: <http://w3id.org/fio/v1/naics#>
PREFIX qudt: <http://qudt.org/schema/qudt/>
PREFIX nhdplusv2: <http://nhdplusv2.spatialai.org/v1/nhdplusv2#>
Step 1: find the samples (200 rows, 14s)
SELECT DISTINCT (?spC AS ?iri) WHERE {
?s2target rdf:type kwg-ont:S2Cell_Level13 .
?s2target spatial:connectedTo kwgr:administrativeRegion.USA.23031 .
?spC rdf:type coso:SamplePoint ;
spatial:connectedTo ?s2target .
?observationC coso:observedAtSamplePoint ?spC .
?observationC coso:ofDSSToxSubstance ?substanceC .
VALUES ?substanceC { <http://w3id.org/DSSTox/v1/DTXSID3031864> }
?s2target spatial:connectedTo ?ds_flowline .
?upstream_flowline hyf:downstreamFlowPathTC ?ds_flowline .
?upstream_flowline rdf:type hyf:HY_FlowPath ;
spatial:connectedTo ?s2neighbor .
?s2anchor kwg-ont:sfTouches | owl:sameAs ?s2neighbor .
?s2anchor rdf:type kwg-ont:S2Cell_Level13 .
?s2anchor kwg-ont:sfContains ?facilityA .
?facilityA fio:ofIndustry ?industryCodeA .
?industryCodeA a naics:NAICS-IndustryCode .
}
Step 2: find the facilities (1,497 rows, 15s)
Identical body, different projection:
SELECT DISTINCT (?facilityA AS ?iri) WHERE {
# ... same WHERE clause as step 1, verbatim ...
}
Step 3: draw the rivers (2,516 flowlines, 7s)
SELECT DISTINCT ?flowline ?flowlineWKT ?fl_type ?streamName WHERE {
{
SELECT DISTINCT ?flowline WHERE {
{
SELECT DISTINCT ?s2cellus WHERE {
VALUES ?facilityA { /* the 1,497 facility IRIs resolved by step 2 */ }
?s2anchor rdf:type kwg-ont:S2Cell_Level13 .
?s2anchor kwg-ont:sfContains ?facilityA .
?facilityA fio:ofIndustry ?industryCodeA .
?industryCodeA a naics:NAICS-IndustryCode .
?s2anchor kwg-ont:sfTouches | owl:sameAs ?s2cellus .
}
}
?upstream_flowline rdf:type hyf:HY_FlowPath ;
spatial:connectedTo ?s2cellus ;
hyf:downstreamFlowPathTC ?flowline .
}
}
{
SELECT DISTINCT ?flowline WHERE {
{
SELECT DISTINCT ?s2celltarget WHERE {
VALUES ?spC { /* the 200 sample-point IRIs resolved by step 1 */ }
?s2celltarget rdf:type kwg-ont:S2Cell_Level13 .
?spC rdf:type coso:SamplePoint ;
spatial:connectedTo ?s2celltarget .
}
}
?_flTarget spatial:connectedTo ?s2celltarget .
?flowline hyf:downstreamFlowPathTC ?_flTarget .
}
}
?flowline geo:hasGeometry/geo:asWKT ?flowlineWKT ;
nhdplusv2:hasFTYPE ?fl_type .
OPTIONAL { ?flowline rdfs:label ?streamName }
}
Steps 4 to 6 (hydration and boundaries, unchanged by any of this)
SELECT
(COUNT(DISTINCT ?observation) as ?resultCount)
(COUNT(DISTINCT ?sample) as ?sampleCount)
(MAX(?numericResult) as ?max)
(GROUP_CONCAT(DISTINCT ?substance; separator="; ") as ?substances)
(GROUP_CONCAT(DISTINCT ?matTypeLabel; separator="; ") as ?materials)
(SAMPLE(?spName) as ?samplePointName)
?sp ?spWKT ?s2cell
WHERE {
VALUES ?sp { /* the 200 sample-point IRIs resolved by step 1 */ }
?sp spatial:connectedTo ?s2cell ;
geo:hasGeometry/geo:asWKT ?spWKT .
OPTIONAL { ?sp rdfs:label ?spName }
?s2cell rdf:type kwg-ont:S2Cell_Level13 .
?observation rdf:type coso:ContaminantObservation ;
coso:observedAtSamplePoint ?sp ;
coso:ofDSSToxSubstance ?substance ;
coso:analyzedSample ?sample ;
coso:hasResult ?result .
?sample rdfs:label ?sampleLabel ;
coso:sampleOfMaterialType ?matType .
?matType rdfs:label ?matTypeLabel .
OPTIONAL { ?result qudt:quantityValue/qudt:numericValue ?numericResult }
OPTIONAL { ?result qudt:quantityValue ?_qv .
?_qv rdf:type coso:NonDetectQuantityValue .
BIND(true AS ?nonDetect) }
BIND(COALESCE(STR(?numericResult),
IF(BOUND(?nonDetect), "non-detect", "non-quantified")) AS ?result_value)
VALUES ?substance { <http://w3id.org/DSSTox/v1/DTXSID3031864> }
} GROUP BY ?sp ?spWKT ?s2cell
@@@HYDRATE_ANCHOR_BY_IRI|federation@@@
SELECT DISTINCT ?facility ?facWKT ?facilityName ?industryCode ?industryName ?s2cell WHERE {
VALUES ?facility { /* the 1,497 facility IRIs resolved by step 2 */ }
?s2cell kwg-ont:sfContains ?facility ;
rdf:type kwg-ont:S2Cell_Level13 .
?facility fio:ofIndustry ?industryCode ;
geo:hasGeometry/geo:asWKT ?facWKT ;
rdfs:label ?facilityName .
?industryCode a naics:NAICS-IndustryCode ;
rdfs:label ?industryName .
}
@@@GET_REGION_BOUNDARIES|spatialkg@@@
SELECT ?region (SAMPLE(?_name) AS ?regionName) (SAMPLE(?_wkt) AS ?regionWKT) WHERE {
?region kwg-ont:administrativePartOf kwgr:administrativeRegion.USA.23 ;
rdfs:label ?_name ;
geo:hasGeometry/geo:asWKT ?_wkt .
FILTER(STRSTARTS(STR(?region), STR(kwgr:)))
} GROUP BY ?region
Your query
SELECT DISTINCT (?spC AS ?iri) WHERE {
# start at samplepoint (limited to a specific region)
?s2target spatial:connectedTo kwgr:administrativeRegion.USA.23031 .
?spC rdf:type coso:SamplePoint ;
spatial:connectedTo ?s2target .
?observationC rdf:type coso:ContaminantObservation ;
coso:observedAtSamplePoint ?spC ;
coso:ofDSSToxSubstance ?substanceC ;
coso:hasResult ?resultC .
?resultC coso:measurementUnit ?unitC .
OPTIONAL { ?resultC qudt:quantityValue/qudt:numericValue ?numericResultC }
OPTIONAL { ?resultC qudt:quantityValue ?_qvC .}
VALUES ?substanceC { <http://w3id.org/DSSTox/v1/DTXSID3031864> }
# trace upstream
?s2target kwg-ont:sfTouches | owl:sameAs ?s2neighbor .
?s2neighbor spatial:connectedTo ?ds_flowline .
?upstream_flowline rdf:type hyf:HY_FlowPath ;
hyf:downstreamFlowPathTC ?ds_flowline .
# find facilities in S2 cells anywhere upstream
?s2anchor rdf:type kwg-ont:S2Cell_Level13;
spatial:connectedTo ?upstream_flowline ;
spatial:connectedTo ?facilityA .
?facilityA fio:ofIndustry ?industryCodeA .
?industryCodeA a naics:NAICS-IndustryCode .
}
Few notes on the comparison
**1. We are starting from the sample side now. Before that, it started from every facility in the graph and ran out of memory.
Note - The app now leads with whichever side the question narrows (a region, a filter, or a pinned IRI list) when the other side is wide open. Here, that is the samples. Ran a few test, this was more optimal.
2. The Merrimack was not coming from this query. I filtered your query's own ?upstream_flowline set for Merrimack, Winnipesaukee, Suncook, and Presumpscot and got zero hits. It came from the separate query that draws the river layer (step 3 above), which used to follow the rivers downstream to the end of the network without ever referencing the samples. That is fixed: 122 Merrimack segments down to 0, 755 flowlines removed, 0 added.
One open question, for you rather than the code
To connect a sample or a facility to a river there has to be a river in its S2 cell (~1km). Often there isn't, so both our queries widen by one ring of neighbouring cells. Your query widens the sample side. Ours widens the facility side.
My question - Should we widen the query on Sample or Facility? or Both?
Few Measurements -
For York County that matters:
| Side |
Total |
Has a river in its own cell |
Lost if not widened |
| PFOS sample points |
361 |
282 (78%) |
78 (22%) |
| Facilities |
1,639 |
1,343 (82%) |
289 (18%) |
|
Sample points |
Facilities |
| Widen samples (yours) |
324 |
1,060 |
| Widen facilities (ours) |
200 |
1,497 |
| Union of both |
330 |
1,507 |
The 130 sample points we currently drop include 82 with measured PFOS detections, up to 2,200 ng/L. Widening both sides in one query fails (2.6 GB allocation failure; tried two ways), so the proposal is to run both and union them.
Following up on the upstream query. Here is exactly what the app runs now, after the fixes in #55
Question: What facilities are upstream from PFOS samples? Block A = Facilities (no filters, no region), Block C = Samples, PFOS, York County ME.
Block C is Constrained
What we run now
Shared prefixes for every query below:
Step 1: find the samples (200 rows, 14s)
Step 2: find the facilities (1,497 rows, 15s)
Identical body, different projection:
Step 3: draw the rivers (2,516 flowlines, 7s)
Steps 4 to 6 (hydration and boundaries, unchanged by any of this)
Your query
Few notes on the comparison
**1. We are starting from the sample side now. Before that, it started from every facility in the graph and ran out of memory.
Note - The app now leads with whichever side the question narrows (a region, a filter, or a pinned IRI list) when the other side is wide open. Here, that is the samples. Ran a few test, this was more optimal.
2. The Merrimack was not coming from this query. I filtered your query's own
?upstream_flowlineset for Merrimack, Winnipesaukee, Suncook, and Presumpscot and got zero hits. It came from the separate query that draws the river layer (step 3 above), which used to follow the rivers downstream to the end of the network without ever referencing the samples. That is fixed: 122 Merrimack segments down to 0, 755 flowlines removed, 0 added.One open question, for you rather than the code
To connect a sample or a facility to a river there has to be a river in its S2 cell (~1km). Often there isn't, so both our queries widen by one ring of neighbouring cells. Your query widens the sample side. Ours widens the facility side.
My question - Should we widen the query on Sample or Facility? or Both?
Few Measurements -
For York County that matters:
The 130 sample points we currently drop include 82 with measured PFOS detections, up to 2,200 ng/L. Widening both sides in one query fails (2.6 GB allocation failure; tried two ways), so the proposal is to run both and union them.