Skip to content

Upstream query: the SPARQL we run now, and one open question on cell widening #57

Description

@prayaslashkari

Following up on the upstream query. Here is exactly what the app runs now, after the fixes in #55

Question: What facilities are upstream from PFOS samples? Block A = Facilities (no filters, no region), Block C = Samples, PFOS, York County ME.

Block C is Constrained

What we run now

Shared prefixes for every query below:

PREFIX rdf:       <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
PREFIX rdfs:      <http://www.w3.org/2000/01/rdf-schema#>
PREFIX owl:       <http://www.w3.org/2002/07/owl#>
PREFIX xsd:       <http://www.w3.org/2001/XMLSchema#>
PREFIX geo:       <http://www.opengis.net/ont/geosparql#>
PREFIX spatial:   <http://purl.org/spatialai/spatial/spatial-full#>
PREFIX kwg-ont:   <http://stko-kwg.geog.ucsb.edu/lod/ontology/>
PREFIX kwgr:      <http://stko-kwg.geog.ucsb.edu/lod/resource/>
PREFIX hyf:       <https://www.opengis.net/def/schema/hy_features/hyf/>
PREFIX coso:      <http://w3id.org/coso/v1/contaminoso#>
PREFIX fio:       <http://w3id.org/fio/v1/fio#>
PREFIX naics:     <http://w3id.org/fio/v1/naics#>
PREFIX qudt:      <http://qudt.org/schema/qudt/>
PREFIX nhdplusv2: <http://nhdplusv2.spatialai.org/v1/nhdplusv2#>

Step 1: find the samples (200 rows, 14s)

    SELECT DISTINCT (?spC AS ?iri) WHERE {
      ?s2target rdf:type kwg-ont:S2Cell_Level13 .
      ?s2target spatial:connectedTo kwgr:administrativeRegion.USA.23031 .
      ?spC rdf:type coso:SamplePoint ;
                spatial:connectedTo ?s2target .
      ?observationC coso:observedAtSamplePoint ?spC .
      ?observationC coso:ofDSSToxSubstance ?substanceC .
      VALUES ?substanceC { <http://w3id.org/DSSTox/v1/DTXSID3031864> }
      ?s2target spatial:connectedTo ?ds_flowline .
      ?upstream_flowline hyf:downstreamFlowPathTC ?ds_flowline .
      ?upstream_flowline rdf:type hyf:HY_FlowPath ;
                spatial:connectedTo ?s2neighbor .
      ?s2anchor kwg-ont:sfTouches | owl:sameAs ?s2neighbor .
      ?s2anchor rdf:type kwg-ont:S2Cell_Level13 .
      ?s2anchor kwg-ont:sfContains ?facilityA .
      ?facilityA fio:ofIndustry ?industryCodeA .
      ?industryCodeA a naics:NAICS-IndustryCode .
    }

Step 2: find the facilities (1,497 rows, 15s)

Identical body, different projection:

SELECT DISTINCT (?facilityA AS ?iri) WHERE {
  # ... same WHERE clause as step 1, verbatim ...
}

Step 3: draw the rivers (2,516 flowlines, 7s)

    SELECT DISTINCT ?flowline ?flowlineWKT ?fl_type ?streamName WHERE {
      {
        SELECT DISTINCT ?flowline WHERE {
          {
        SELECT DISTINCT ?s2cellus WHERE {
          VALUES ?facilityA { /* the 1,497 facility IRIs resolved by step 2 */ }
          ?s2anchor rdf:type kwg-ont:S2Cell_Level13 .
          ?s2anchor kwg-ont:sfContains ?facilityA .
      ?facilityA fio:ofIndustry ?industryCodeA .
      ?industryCodeA a naics:NAICS-IndustryCode .
          ?s2anchor kwg-ont:sfTouches | owl:sameAs ?s2cellus .
        }
      }
          ?upstream_flowline rdf:type hyf:HY_FlowPath ;
            spatial:connectedTo ?s2cellus ;
            hyf:downstreamFlowPathTC ?flowline .
        }
      }
      {
        SELECT DISTINCT ?flowline WHERE {
          {
            SELECT DISTINCT ?s2celltarget WHERE {
              VALUES ?spC { /* the 200 sample-point IRIs resolved by step 1 */ }
              ?s2celltarget rdf:type kwg-ont:S2Cell_Level13 .
              ?spC rdf:type coso:SamplePoint ;
                spatial:connectedTo ?s2celltarget .
            }
          }
          ?_flTarget spatial:connectedTo ?s2celltarget .
          ?flowline hyf:downstreamFlowPathTC ?_flTarget .
        }
      }
      ?flowline geo:hasGeometry/geo:asWKT ?flowlineWKT ;
                nhdplusv2:hasFTYPE ?fl_type .
      OPTIONAL { ?flowline rdfs:label ?streamName }
    }
Steps 4 to 6 (hydration and boundaries, unchanged by any of this)
    SELECT
      (COUNT(DISTINCT ?observation) as ?resultCount)
      (COUNT(DISTINCT ?sample) as ?sampleCount)
      (MAX(?numericResult) as ?max)
      (GROUP_CONCAT(DISTINCT ?substance; separator="; ") as ?substances)
      (GROUP_CONCAT(DISTINCT ?matTypeLabel; separator="; ") as ?materials)
      (SAMPLE(?spName) as ?samplePointName)
      ?sp ?spWKT ?s2cell
    WHERE {
      VALUES ?sp { /* the 200 sample-point IRIs resolved by step 1 */ }
      ?sp spatial:connectedTo ?s2cell ;
          geo:hasGeometry/geo:asWKT ?spWKT .
      OPTIONAL { ?sp rdfs:label ?spName }
      ?s2cell rdf:type kwg-ont:S2Cell_Level13 .
      ?observation rdf:type coso:ContaminantObservation ;
          coso:observedAtSamplePoint ?sp ;
          coso:ofDSSToxSubstance ?substance ;
          coso:analyzedSample ?sample ;
          coso:hasResult ?result .
      ?sample rdfs:label ?sampleLabel ;
          coso:sampleOfMaterialType ?matType .
      ?matType rdfs:label ?matTypeLabel .
      OPTIONAL { ?result qudt:quantityValue/qudt:numericValue ?numericResult }
      OPTIONAL { ?result qudt:quantityValue ?_qv .
                 ?_qv rdf:type coso:NonDetectQuantityValue .
                 BIND(true AS ?nonDetect) }
      BIND(COALESCE(STR(?numericResult),
             IF(BOUND(?nonDetect), "non-detect", "non-quantified")) AS ?result_value)
      VALUES ?substance { <http://w3id.org/DSSTox/v1/DTXSID3031864> }
    } GROUP BY ?sp ?spWKT ?s2cell
@@@HYDRATE_ANCHOR_BY_IRI|federation@@@
    SELECT DISTINCT ?facility ?facWKT ?facilityName ?industryCode ?industryName ?s2cell WHERE {
      VALUES ?facility { /* the 1,497 facility IRIs resolved by step 2 */ }
      ?s2cell kwg-ont:sfContains ?facility ;
              rdf:type kwg-ont:S2Cell_Level13 .
      ?facility fio:ofIndustry ?industryCode ;
                geo:hasGeometry/geo:asWKT ?facWKT ;
                rdfs:label ?facilityName .
      ?industryCode a naics:NAICS-IndustryCode ;
                    rdfs:label ?industryName .
    }
@@@GET_REGION_BOUNDARIES|spatialkg@@@
    SELECT ?region (SAMPLE(?_name) AS ?regionName) (SAMPLE(?_wkt) AS ?regionWKT) WHERE {
      ?region kwg-ont:administrativePartOf kwgr:administrativeRegion.USA.23 ;
              rdfs:label ?_name ;
              geo:hasGeometry/geo:asWKT ?_wkt .
      FILTER(STRSTARTS(STR(?region), STR(kwgr:)))
    } GROUP BY ?region

Your query

    SELECT DISTINCT (?spC AS ?iri) WHERE {
# start at samplepoint (limited to a specific region)
      ?s2target spatial:connectedTo kwgr:administrativeRegion.USA.23031 .
      ?spC rdf:type coso:SamplePoint ;
                spatial:connectedTo ?s2target .
      ?observationC rdf:type coso:ContaminantObservation ;
          coso:observedAtSamplePoint ?spC ;
          coso:ofDSSToxSubstance ?substanceC ;
          coso:hasResult ?resultC .
      ?resultC coso:measurementUnit ?unitC .
      OPTIONAL { ?resultC qudt:quantityValue/qudt:numericValue ?numericResultC }
      OPTIONAL { ?resultC qudt:quantityValue ?_qvC .}
      VALUES ?substanceC { <http://w3id.org/DSSTox/v1/DTXSID3031864> }

# trace upstream
      ?s2target kwg-ont:sfTouches | owl:sameAs ?s2neighbor .
        ?s2neighbor spatial:connectedTo ?ds_flowline .
      ?upstream_flowline rdf:type hyf:HY_FlowPath ;
               hyf:downstreamFlowPathTC ?ds_flowline .
# find facilities in S2 cells anywhere upstream
      ?s2anchor rdf:type kwg-ont:S2Cell_Level13;
             spatial:connectedTo ?upstream_flowline ;
             spatial:connectedTo ?facilityA .
      ?facilityA fio:ofIndustry ?industryCodeA .
      ?industryCodeA a naics:NAICS-IndustryCode .
    }

Few notes on the comparison

**1. We are starting from the sample side now. Before that, it started from every facility in the graph and ran out of memory.

Note - The app now leads with whichever side the question narrows (a region, a filter, or a pinned IRI list) when the other side is wide open. Here, that is the samples. Ran a few test, this was more optimal.

2. The Merrimack was not coming from this query. I filtered your query's own ?upstream_flowline set for Merrimack, Winnipesaukee, Suncook, and Presumpscot and got zero hits. It came from the separate query that draws the river layer (step 3 above), which used to follow the rivers downstream to the end of the network without ever referencing the samples. That is fixed: 122 Merrimack segments down to 0, 755 flowlines removed, 0 added.

One open question, for you rather than the code

To connect a sample or a facility to a river there has to be a river in its S2 cell (~1km). Often there isn't, so both our queries widen by one ring of neighbouring cells. Your query widens the sample side. Ours widens the facility side.

My question - Should we widen the query on Sample or Facility? or Both?

Few Measurements -

For York County that matters:

Side Total Has a river in its own cell Lost if not widened
PFOS sample points 361 282 (78%) 78 (22%)
Facilities 1,639 1,343 (82%) 289 (18%)
Sample points Facilities
Widen samples (yours) 324 1,060
Widen facilities (ours) 200 1,497
Union of both 330 1,507

The 130 sample points we currently drop include 82 with measured PFOS detections, up to 2,200 ng/L. Widening both sides in one query fails (2.6 GB allocation failure; tried two ways), so the proposal is to run both and union them.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions