Skip to content

Development -> Main, Sync - #60

Merged
prayaslashkari merged 88 commits into
mainfrom
development
Sep 25, 2026
Merged

prayaslashkari merged 88 commits into
mainfrom
development

Conversation

@prayaslashkari

Copy link
Copy Markdown
Collaborator

No description provided.

prayaslashkari and others added 30 commits August 25, 2026 16:00
Recreates UC1-CQ2c ("what streams are downstream at most N km from
facilities of industry X?") from David's notebook, which needed two
things the query builder could not express.

Streams as an answerable entity. Flowlines previously reached the map
only as a decorative layer traced from the anchors; hyf:HY_FlowPath was
not selectable in either block. Adds 'streams' to EntityType with an
FTYPE filter, IRI hydration, and map wiring. When the target is a
flowline the ?s2target hop is replaced with a direct bind — that hop
would otherwise match any flowline sharing a cell with the answer — and
the supporting stream layer is skipped so flowlines aren't drawn twice.

Cumulative distance cutoff. downstream/upstream traces were the full
transitive closure. Adds an optional maxDistanceKm that sums
nhdplusv2:hasFlowPathLength over the segments between seed and
candidate and filters on the total. Unset, the emitted SPARQL is
byte-identical to before, so existing questions are unaffected.

The bound is not just a filter: "samples within 30 km downstream of NH
airports" returns 27 sample points in ~15s, where the unbounded form
times out against the federation gateway.

Two things worth knowing for review:

The notebook and this implementation do not agree, and the notebook is
wrong. Its outer block re-joins facilities on schema1:address as a
required triple, a predicate only 13 of 144 NH airport facilities carry,
so it silently discards 91% of its own anchor set. Drop that triple and
it returns 1,547 flowlines where we return 1,605 on the same seeding —
we are a superset. Address stays OPTIONAL here.

Neighbour-cell expansion is kept, per discussion: we seed from the
facility's S2 cell and its 8 neighbours where the notebook uses the
facility's own cell only. That is the app-wide convention and accounts
for the rest of the difference (2,757 vs 1,605 for NH airports at 30km).

scripts/flow-distance-check.mjs runs the real planner against the live
endpoints and asserts the step plan, the superset relation, that the
bound actually excludes flowlines, and that every hydrated stream has
drawable geometry.
The "+1" from David's second UC1-CQ2c notebook. A distance budget runs
out at whatever segment happens to fit, which is an artifact of how
NHDPlus split the river rather than a real feature — the drawn path
stops mid-channel. Extending one segment past the boundary means the
flowpath visibly crosses the threshold instead of ending at it, which
is the notebook's stated intent: the total may deliberately exceed the
limit.

Always on when a cutoff is set. A checkbox for a 1%-at-30km difference
is a control nobody would understand, and "within 30 km" already reads
as approximate.

Implemented as a zero-or-one property path, which yields the endpoint
and its immediate neighbour in one triple. A UNION says the same thing,
but QLever — which the notebooks run against — returns unbound results
for MIN() over a variable bound inside a UNION, verified on a two-row
test case with no data involved. The path form works on both hosts.
Same reason `?a = ?b` rather than sameTerm(): QLever has not
implemented sameTerm.

Fringe segments cannot inherit their parent's distance, since that
would report a sub-threshold number for a flowline outside the
threshold. They carry the parent's path plus the parent's own length —
the distance to where the fringe segment begins, which never
understates. A flowline reachable as both a valid endpoint and a fringe
keeps the smaller value, since MIN runs over both.

Effect scales inversely with the cutoff, as the fringe is a larger
share of a smaller answer (NH airports):

    5 km    1,533 -> 1,669   (+8.9%)
   10 km    2,019 -> 2,096   (+3.8%)
   30 km    2,757 -> 2,784   (+1.0%)

Verified against apps.okn.us: 2,784 at 30 km, 1,669 at 5 km, 3,490
unbounded (unchanged), and flowline distances spanning 0.01–32.23 km
with the fringe correctly reporting past the threshold. FRINK is
returning 503 across all five endpoints right now, so the committed
check script has not been re-run against it — the two hosts were
verified to return identical result sets earlier in this work.
…led substances

The Substance dropdown returned nothing. buildDiscoverSubstancesQuery required
`?substance dcterms:alternative ?_label`, and that predicate has zero triples on
substances, on sawgraph and federation alike. Because it was required rather
than OPTIONAL the query returned 0 rows instead of returning substances without
labels. Substance names live on rdfs:label.

Switching the predicate alone still hid 32 of the 101 substances, since only 69
carry rdfs:label. Those 32 account for 47,642 observations, 5.0% of everything
with a substance link. So both label patterns are now OPTIONAL, with a fallback
to the DTXSID from the URI in useSubstances, mirroring how useMaterialTypes
already handles a missing label.

Verified live: unfiltered 0 -> 69 -> 101 substances, Maine 0 -> 69 -> 79.

The likely origin of the wrong predicate is the facility table in
docs/SCHEMA.md, where dcterms:alternative is a legitimate facility predicate:
3,824,195 triples in fiokg, 0 of them on comptox:ChemicalEntity.
…ounts

Both select components already computed the option count but only rendered it
while a search term was typed, so opening a dropdown showed a bare "Select all".
It now always shows the number of options the checkbox would tick, for example
"Select all (79)". The number still excludes disabled options and narrows with
an active search, so it keeps matching what the checkbox actually does.

Counts are comma formatted everywhere, on the Select all row and on individual
options: "PFOA (20,120)" rather than "PFOA (20120)".

Note that dropdown search matches on the label text, which contains the count,
so typing 20120 no longer matches "PFOA (20,120)".
…ging log

The wiki page and SCHEMA.md both described the substance query wrongly, and the
wiki's stated remedy would not have worked.

Dropdowns Substance: query blocks updated to the OPTIONAL rdfs:label form,
numbers refreshed (Maine returns 79, not 71; PFOA 20120, not 15217), the fixed
"Known issue" callout removed, and the "dcterms:alternative only exists on
federation" diagnosis replaced. That claim was wrong: the predicate returns zero
rows on federation too, so both variants were broken, not just the unfiltered
one. Added a troubleshooting entry for bare DTXSIDs, and supplied the screenshot
the page has referenced since it was written.

Dropdowns: documented the Select all count once, as shared FlatSelect behaviour.

SCHEMA.md had no substance or chemical inventory at all. Added a Substance
Labels section with the coverage that matters (rdfs:label 69 of 101,
skos:altLabel 25, dcterms:alternative none), a comptox:ChemicalEntity class
count, and an explicit warning that the identically named facility predicate is
a different thing.

DEBUGGING.md records the root cause and the reusable lesson: a required label
pattern does not degrade, it deletes.
Week 36 entry 2 said dcterms:alternative "exists only on federation, which is
why the region-scoped variant works and returns 71 substances for Maine", and
proposed pointing the no-region query at federation. Measured again, that
predicate returns zero rows on federation as well, so the region-scoped variant
was also broken and the proposed fix would not have worked. Corrected in place
with a dated note rather than a silent edit.

Week 37 covers the query fix, the 32 recovered substances, the Select all
counts, and this documentation pass.
32 substances have no rdfs:label of their own and were showing as bare DTXSIDs.
Every one of them is the target of a comptox:sameAsDSSToxSubstance link from the
source-data parameter it was matched to, and that parameter carries a name on
rdfs:label. Added it as a fourth step in the chain:

  skos:altLabel -> rdfs:label -> parameter's rdfs:label -> DTXSID

Verified live: 101 of 101 named unfiltered, 79 of 79 for Maine, no DTXSID rows
and no duplicate labels. The DTXSID step is now unreachable and stays as a
safety net.

Surveyed all 39 predicates a substance carries first. Only two name it; the rest
are identifiers (hasCASRN, hasInChIKey, all inside the already-labelled 69),
provenance URIs, or physicochemical properties. owl:sameAs is self-referential.

The parameter also carries acronyms covering 68 rather than 25, deliberately not
adopted: the _A suffix marks the acid as distinct from the anion, so PFOS_A is
Perfluorooctanesulfonic acid and PFOS is Perfluorooctanesulfonate, two different
DTXSIDs. Borrowing them collapsed 10 pairs of distinct substances into identical
rows. The full chemical names keep them apart.

Three QLever behaviours shaped the query, all documented in DEBUGGING.md:
COALESCE(SAMPLE(x), SAMPLE(y)) crashes on federation; SAMPLE over a nested
OPTIONAL can return unbound even when rows bind it, which silently cost
coverage; MIN is deterministic where SAMPLE is not.
SCHEMA.md: replaced the substance coverage table with a field-by-field verdict
across all 39 predicates, the regeneration query, the four-step label chain, and
why the parameter's skos:altLabel must not be used as an acronym (_A marks acid
vs anion, _L and _BR mark isomers). Also records the parameter-name quality
problems and the known upstream PFECHS misalignment.

Dropdowns Substance: the new OPTIONAL in both query blocks with its gloss, a
note on MIN over SAMPLE, and three rewritten troubleshooting entries covering
bare DTXSIDs, full names where an acronym was expected, and the non-PFAS entry.

DEBUGGING.md: the three QLever aggregate behaviours. The silent one, where
SAMPLE returns unbound despite bound rows in the group, is the one worth
remembering.
Maine EGAD instance data has moved to http://w3id.org/sawgraph/v2/me-egad-data#.
Verified live: 833,201 observations, 63,744 samples and 10,331 sample points on
v2, zero on v1.

Nothing broke, because no query constructs a me-egad-data IRI; they all reach
that data by type and predicate. The prefix was declared but never referenced,
and it now points at a namespace with no triples, which is the same trap that
dcterms:alternative was. Removing it rather than repointing it: if a query ever
needs it, it should add the correct namespace deliberately.

Controlled vocabulary is unaffected, staying in the root v1/me-egad# namespace.
…ation

Controlled vocabulary now lives in the root namespaces on both endpoints, which
retires three claims on this page.

The "URIs change with the region" quirk is gone: sawgraph and federation once
named groundwater me-egad# against me-egad-data#, so a selection made before
picking a state stopped matching after. Verified 2026-09-09, all six fallback
material types return identical counts on both endpoints. Its troubleshooting
entry is rewritten as history rather than current behaviour.

Species are under v1/us-wqp#biologicalTaxon.*, not us-wqp-data# as written.
…alue

coso:measurementValue became multi-valued. A non-detect result now carries two
literals, "non-detect" and "non-quantified", so binding it duplicated every
non-detect row: of 1,041,155 distinct results only 317,676 have a single value.

It is also string-typed, which makes MAX(?result_value) return "non-quantified"
rather than a number whenever a non-detect reaches the aggregate.

Replaced the binding at all four sites with resultValueClauses(), which derives
one value per result from qudt:quantityValue: the number if there is one, else
"non-detect", else "non-quantified". Measured over the whole graph that is
1,041,433 rows for 1,041,420 results, down from roughly 723,000 duplicates; the
13 stragglers carry two numeric values and are absorbed by the GROUP BY.

The two aggregates now use MAX(?numericResult). qudt:numericValue is xsd:double
and orders numerically, verified returning 229.0 / 4.87 / 39.9 where the old
form returned a string.

No user-visible change yet: coso:measurementUnit is still required everywhere,
and only detects have one, so non-detects never reach these queries. This is the
correctness groundwork for the two commits that follow.
The Include non-detects checkbox had been inert in both directions. It keyed off
qudt:enumeratedValue, which has zero triples: contaminoso a533442 (2026-04-20)
removed it in favour of typing the value node instead.

?enumDetected therefore never bound, so:
  range + include  ->  FILTER(numeric || BOUND(?enumDetected))  dropped non-detects
                       even when asked to include them
  no range + exclude -> FILTER(!BOUND(?enumDetected))           was always true,
                       so the exclusion did nothing

Now keyed off ?nonDetect from resultValueClauses, which tests
?qv rdf:type coso:NonDetectQuantityValue (723,410 instances). Chose the narrow
class over NonQuantifiedQuantityValue (814,291) because non-detect is a strict
subset; the other 90,881 are unquantified for other reasons.

Exclude now means "keep only rows with a real number", which also drops that
non-quantified cohort. The redundant OPTIONALs and the xsd:decimal(?result_value)
COALESCE leg go away, since resultValueClauses already binds those.

Still no user-visible change: the required unit join keeps non-detects out of
the query entirely. The next commit addresses that.
coso:measurementUnit exists on exactly 228,124 results, which is precisely the
set that has a numeric value. Requiring it therefore excluded every non-detect
from every sample query: about 70% of all results, 722,206 of 1,041,420.

That join was only ever needed by the ng/L concentration-range filter, so it is
now emitted only when a range is set (needsUnitJoin), and the unit test moved
onto the numeric branch of the filter where it belongs.

Verified live on one sample point:
  include non-detects   28 observations   max=202.0
  exclude non-detects   16 observations   max=202.0
  range 10-100 exclude   3 observations   max=19.4
  range 10-100 include  15 observations   max=19.4
Warm-cache timings 2.5-3.5s; the 23s first run was cold cache.

Two places deliberately keep the join, both documented in the code:

The sample-detail queries, because making the unit optional there produces
reproducible 502s and >120s hangs on a healthy endpoint. It buys nothing today
since resultTransformer.ts discards non-numeric values anyway.

The fused pipeline query, because admitting non-detects roughly triples the
matched rows and the endpoint OOMs on statewide runs. Verified: the prebuilt
Maine query fails outright with them included, and completes in 32s with 29 map
features without.

Consequence to be aware of: a sample point can now report more observations on
the map than the popup lists, because the popup path is still detect-only.
Surfacing non-detects in the popup is a separate, deferred decision.
Every material type IRI we ship pointed at v1/me-egad-data#, which matches
zero triples. Controlled vocabulary was moved out of the -data namespace into
the root while instance data moved the other way, to v2/me-egad-data#. Source
of truth is pfas-kg, datasets/maine/egad/controlledVocab/sample_type.ttl,
where the prefix is v1/me-egad# and 47 terms are defined. Confirmed by
Katrina's note on the reload: "all controlled vocabulary terms are in the root
namespaces".

Verified live, all four combinations for groundwater:

  v1/me-egad#sampleMaterialType.GW        46,110 triples
  v1/me-egad-data#sampleMaterialType.GW        0
  v2/me-egad#sampleMaterialType.GW             0
  v2/me-egad-data#sampleMaterialType.GW        0

The failure was silent. VALUES with a dead IRI is valid SPARQL, so the
endpoint answers 200 with no rows and the app cannot tell that apart from
"no data here". The PFHpA Cumberland County card ran every step, reported
success and drew an empty map.

Also corrects the sludge code. The fallback listed .SO, which appears nowhere
in the vocabulary; the real code is .SU. All six shipped IRIs now resolve:
GROUNDWATER 609,496, SOIL 80,386, SURFACE WATER 26,424, LEACHATE 6,126,
DRINKING WATER 2,858, SLUDGE 850.

Scoped to the sample side, the card's own filters now return 65 sample points
and 171 observations in 2.3s, against 0 before. The full card still fails on
the downstream hydrology join, which times out and OOMs on federation. That
is pre-existing and unrelated: the same query fails identically with the old
dead IRI, and the endpoint is currently reporting 88 MB available.
Entries 5 to 8 were written locally on 2026-09-09 and never committed. Restored
them, then corrected the material entry, which described a fix that had not
landed in the repo until ef0c40b.

Changes to that entry:

- Dated 2026-09-11, not 09-09, and moved last so the file stays chronological.
  Remaining entries renumbered, and the two cross references between them
  repointed.
- Replaced "went from 0 to 49 sample points and 268 facilities" with what is
  measurable today: 65 sample points and 171 observations, scoped to the sample
  side. The old figure was a whole-pipeline result and cannot be reproduced,
  because the card now fails on the downstream hydrology join.
- Records that the card is still broken for a second, unrelated reason. The
  same query times out with the old dead IRI too, so the namespace fix neither
  caused nor worsened it.
- Labels the sludge count, which was a bare "334". It is a sample count, now
  330 samples / 850 observations.
- Adds the source of truth, pfas-kg's sample_type.ttl, and the four-namespace
  comparison behind the fix.
- Notes that QLever returns HTTP 429 for a query timeout, which reads like rate
  limiting and is not. That cost some time to work out.

The audit entry claimed 7 of 8 prebuilt queries passed. True on the day only
because the material fix was applied in the working tree; anyone checking out
that commit range would have seen the PFHpA card fail. Said so plainly.
The Material dropdown already grouped its 171 options into Water, Biota, Solid
Material, Air and Other, but the headings were plain divs. No checkbox, no
arrow, so no way to say "all water samples" and no way to fold a group away.
That hurts because the distribution is lopsided: 131 of the 171 are individual
fish species, all under Biota.

Points the field at HierarchicalSelect, the tree the Industry dropdown already
uses, rather than FlatSelect.

No SPARQL changes. The groups are a display concept the graph knows nothing
about, so a heading carries a synthetic __group: code that is expanded back to
its children's real material-type URIs before anything is stored. The query the
pipeline sends is identical to ticking every child by hand. Counts roll up
correctly because MIN(?bucketPrio) already forces each material type into
exactly one group — verified for Maine, where Water reads 556,269 and its
thirteen children sum to 556,269.

Generalising the tree component for a second caller:

- buildTree takes an optional explicit parent, since material-type URIs have no
  prefix convention to derive one from. Prefix trimming stays as the fallback,
  so Industry is untouched.
- When a caller supplies explicit parents, its ordering is kept rather than
  sorted by code length. Without this the discovery query's ORDER BY DESC(?num)
  was thrown away and Biota's 131 species came out in opaque URI order. Caught
  by the tree check, not by reading the code.
- labelOnly keeps the code out of labels, chips and search, because "3253 -
  Agricultural Chemical Manufacturing" is right for NAICS and a URI is not.
  Search matching on the code would otherwise make "w3id" match all 171.
- industries prop renamed to items; one call site.

Also adds materialTypeLabels, so the generated question reads "GROUNDWATER
samples" rather than "GW samples" and "Lepomis macrochirus" rather than "11868".

Verified in the running app against Maine: four tickable groups with Air
correctly absent, Water expanding to thirteen count-ordered children, ticking
the heading yielding one chip instead of thirteen, and partial selection showing
indeterminate.
…traps

The Material page described a control that no longer exists and printed two
SPARQL queries that never matched the code.

Wiki:

- Dropdowns Material. Rewrites the wiring for HierarchicalSelect, adds a section
  on where the five groups come from (MIN(?prio) over the four direct
  coso:MaterialSample subclasses, 171 material types producing 178 type-to-bucket
  links, seven overlaps resolved by MIN), and a group-count table for unfiltered,
  Maine and Cumberland. Spells out that ticking a group changes nothing about the
  SPARQL.
- Both SPARQL blocks on that page were missing the bucket block and ?bucketPrio
  entirely, so copying either gave ungrouped results. Both now match what the
  code sends, and both were run verbatim against the live endpoints: 171 rows and
  25 rows respectively.
- Five new "if it looks wrong" entries: the fallback now shows as a single Other
  heading because those six entries carry no group; group counts are observations
  while results are sample points; a selection that drops out of the option list
  after a region change stops being visible while still filtering; 14 samples
  carry no material type and vanish under any selection; and a dead IRI returns
  an empty map with no error.
- Dropdowns index claimed HierarchicalSelect was "used only by Industry" and that
  everything used FlatSelect "except one". Rewritten around why each control is
  chosen and how the two build their trees differently.
- Industry: buildTree's description was half the story now that an explicit
  parent can win over prefix trimming. Industry Counts: rollupCounts is shared,
  and summing children only holds when they partition the parent.

DEBUGGING.md, two entries:

- A dead IRI fails silently. VALUES with an IRI that has no triples is valid
  SPARQL, returns 200 with zero rows, and is indistinguishable from an honest
  empty result. Same family as the dcterms:alternative entry above it.
- QLever reports a query timeout as HTTP 429, which reads like rate limiting and
  is not; the body carries the real reason. A 500 with "Tried to allocate X MB"
  is the same class. Available memory was seen swinging from 370 MB to 88 MB in
  one afternoon, so the same query passes and fails minutes apart.

SCHEMA.md: records the August 2026 namespace split, which was written down
nowhere. Instance data moved to v2/<source>-data#, controlled vocabulary went the
other way into the v1 roots. One WQP sample carries both, which is the clearest
illustration. Adds the four-way IRI comparison and the material-type inventory:
171 values across four unreconciled vocabularies, plus the 14 samples with none.

Also fixes three deep links that were already broken before this change, found by
resolving every #L reference against source: County pointed at a blank line and a
stray ");", Industry Counts at "staleTime: Infinity", Industry at an import. And
corrects a stale example claiming sampleMaterialType.WW has no label — all 171
carry labels and .WW is WASTE WATER.
The August 2026 reload split the sawgraph namespaces in two directions:
instance data moved to v2/<source>-data#, controlled vocabulary came out
of -data into the root v1/<source>#. ef0c40b fixed the IRIs we ship, but
not the ones users had already saved.

Saved questions (localStorage) and published workflows (server DB) both
store a full AnalysisQuestion, vocabulary IRIs included. Any of those
created before the reload still pins the dead form, and a VALUES clause
listing a dead IRI is valid SPARQL: the endpoint answers 200 with zero
rows, so the map renders empty and looks like an honest "no data here".

Rewrite on read instead of migrating the database. queryStore.loadQuestion
is the single funnel for saved, published, prebuilt and new questions, so
the migration goes there. Stringify/replace/parse rather than walking the
tree, which also catches the label maps, where IRIs are object keys.

Verified live on both sawgraph and federation: the split is complete and
nothing is dual-published. Instance data is 100% v2 (63,744 Maine +
44,868 WQP samples), vocabulary is 100% root v1 (86 me-egad + 85 us-wqp
material types). So every stored -data# vocabulary IRI is dead, and no v2
instance IRI may be rewritten.
QLever kills any query at 30s and reports it as HTTP 429 "Too Many Requests",
which reads as rate limiting but is a timeout. Our heaviest queries sat at or
past that wall, so a large part of the query space failed — a full sweep of the
156 shapes the editor can build found 37 failing outright and 32 more running
>15s against the 30s limit (docs/QUERY-MATRIX.md).

Root cause is unbounded work per request: a single query carries the whole
river-network closure or a statewide spatial join, so cost scales with how much
data exists in the state rather than with what the user asked to see.

Execution now runs a query whole and, only if the engine refuses, splits the
scope into slices that fit and merges the results:

- scope.ts picks a split axis by probing each side with LIMIT 201 (not COUNT —
  counting Maine's unfiltered facilities takes 11.0s). Anchor entities, else
  target entities, else counties. The right axis flips between states: Maine has
  4,528 samples / 2,976 facilities, Illinois has 78 / 22,574, and county slicing
  that works in Maine dies in Cook County at 3.3GB.
- executor.ts merges slices with dedupe, reports per-chunk progress, returns
  partial results rather than discarding a run when some slices fail, and caps
  each step at 120s so a step whose slices all fail cannot subdivide forever.
- Aggregate queries may only be sliced on their GROUP BY key, so per-entity
  counts stay correct.

Also drops GET_SAMPLE_DETAILS from the pipeline. It fetched every observation
for every sample up front — 18-37MB per run — to fill popups that open one at a
time; a single sample point is ~30KB, now fetched on popup open. Region
boundaries are cached per session (measured 25.5s cold for data that never
changes), and the nine distinct SPARQL failures are classified and explained to
the user instead of sharing one "Something went wrong" card.

Deletes MAX_INLINE_SAMPLE_IRIS and the inline-vs-re-derive fork, along with
buildFusedSampleAggregateQuery and buildFusedSampleDetailsQuery: slices are
small enough to inline, so neither horn of that dilemma applies.

Measured against the live endpoints:
- ME samples downstream of airports: 21.2s -> 16.1s warm, same 1,172 -> 999 rows
- samples near samples [ME]: 429 timeout -> 86.3s
- IL samples downstream of airports: 500 OOM -> 218s (one flowline slice partial)
- ME samples downstream of 83,038 wells: 429 -> works via 16 county slices

Two shapes still fail: facilities/wells downstream of facilities with no filter
or region on the second block. That block means "every facility in the graph",
all 1,506,326 of them, traced through the national river network. They now fail
in ~2.5 min with guidance naming the second block, rather than at 30s with a
generic message — measurement confirmed that narrowing the *first* block does
not help (still OOM at 2.6GB) while constraining the second does.
Verification sweep of all 124 shapes through the real engine found three
defects in what the previous commit shipped.

No client-side timeout. executeSparql waited indefinitely, so a stalled endpoint
blocked single requests for 959s and 1,918s in the sweep. The per-step budget
could not intervene because it is checked between slices, and — worse — a hang
never becomes a classified error, so the query never split. Three shapes
recorded as "still failing" were only hanging; all three work once a stall is
reported as a splittable timeout at 60s. facilities upstream facilities now
returns 2,813 rows in 224s, complete.

Two budget rules, both wrong in opposite directions. Budgeting elapsed time
killed runs that were succeeding: wells near(4) had 13 of 16 county slices
working and was cut off at 120s for being big. Budgeting only wasted time then
broke the opposite case: Illinois downstream has slices that fail slowly before
the ones that succeed, so it gave up before reaching the answer. Both are
replaced by one rule — every slice is attempted, only the 300s step total is
bounded. Illinois now completes in 248s, wells near(4) in 283s with 69,127 rows
and nothing missing.

Failed and skipped slices were reported as one number, so the UI blamed size for
work that had simply run out of budget. They now carry separate counts and
separate advice: narrow the question, versus just run it again against a warmer
cache.

Sweep result: of the 37 shapes failing at baseline, 34 now work and none
regressed. The 3 that remain all leave the second block unconstrained — the
largest being "every facility in the graph", all 1,506,326 of them. Small
questions are unaffected: one county near-1-mile is 0.8s, and wells near 2 miles
returns 45,280 rows in 18.8s. All 8 dashboard queries pass, three of them
materially faster now that the 18-37MB detail step is gone.

Documented in docs/QUERY-MATRIX.md Part 7, raw data in docs/query-matrix-after.csv.
Three documentation gaps, none of which the code change closed on its own.

DEBUGGING.md is the file CLAUDE.md points at first for pipeline and SPARQL
issues, and it had nothing about any of this — no mention of 429, timeout or
OOM. The single most useful fact, that a 429 from these endpoints is a 30s query
timeout and not rate limiting, was only in a plan document. Added an entry in the
existing format: symptom, root cause, the table of all nine failure responses and
what each really means (including 413/502 surfacing in the browser as a CORS
error because the response omits ACAO), and prevention notes — read the exception
body before assuming rate limiting, an unfiltered second block means every entity
in the graph, and never benchmark against Maine alone.

Moved the execution plan from drafts/ to active/ per docs/plans/README.md, now
that Phases 1 and 2 are implemented and verified, and fixed the references to it.

Added the week's changelog: bounded execution, on-demand popup measurements,
explained failures, and the query matrix itself.

Also refreshed the team email draft, which still claimed upstream and multi-hop
questions were broken — Phase 2 fixed both. It now carries the verified numbers:
37 failing shapes down to 3, no regressions.
Opening a shared link re-ran the entire SPARQL pipeline in the visitor's
browser — 26-218s, sometimes failing outright — to rebuild an answer that does
not change. That is the flow behind the original "Something went wrong" report.

The plan called for moving execution to the server. Exploration found a cheaper
answer: the publisher's browser has already computed the result, and it is
sitting in the store when they click Publish. Publishing now hands it over,
authenticated by the editToken the server already issues, and every later
visitor gets it in one request.

Measured on a dashboard question, against a local Postgres and the live
endpoints: 37ms cached versus 6,604ms live on a warm graph engine — 178x — with
identical row counts either way (57 samples, 8 facilities, 16 boundaries). A
1.4MB wire payload stores as 86KB gzipped, and the internal IRI lists the map
never renders are trimmed before it is sent.

Only trusted paths write. The publisher writes their own result; a
maintainer-run script (scripts/warm-cache.mts) writes the dashboard questions
with a token that never ships in the browser bundle. Reads are public. An open
write endpoint was considered and rejected rather than mitigated: this tool
renders PFAS readings near named facilities, and anyone able to write the cache
could place fabricated contamination on a real map identically for every
visitor, which no amount of framework escaping defends against.

A cached map says so, with the date it was computed and a Re-run live button —
this tool's output ends up in reports, and a stored answer must never be
mistaken for a fresh one. Any miss or error falls through to running locally
exactly as before, so the cache can only make a run faster.

Server-side execution stays deferred. The read path is identical either way, so
it can be swapped in later without touching the client, and avoiding it means no
engine bundling and no Railway build-root change.

Two bugs the verification caught:

- The global 64kb JSON parser rejected result uploads before the route-scoped
  parser ran. Route-scoped body limits are useless if something global already
  read the body; index.ts now skips the two upload paths and every other route
  keeps its 64kb guard (verified: an oversized publish POST still 413s).
- Canonicalisation left empty objects behind, so {industryCodes: []} hashed
  differently from a question with no filter at all — a filter set and then
  cleared would have missed its own cache entry forever. scripts/check-cache-key.mts
  covers this and 14 other cases, since a key bug is silent in both directions.

Also documents the four Railway services in the README (nothing in the repo
recorded where the app runs) and drops the last reference to the Render API,
which now returns 503.
The dashboard questions are cached by a maintainer-run script, which only helps
if running it is easy. Actions -> Warm result cache -> Run workflow now does it:
choose development or production, optionally filter to one question by title,
and read what was cached in the job summary.

Concurrency is capped at one run per environment — these queries are heavy on a
shared graph engine and two runs would compete for it. The job fails with an
explicit message when the environment's token is missing, rather than uploading
nothing and reporting success.

No schedule attached. A cached entry goes stale only when the graph reloads or
its 30-day TTL lapses, both of which we control; the workflow header carries the
three lines that would add a weekly cron if that changes.

tsx moves from an on-demand npx download to a pinned devDependency so CI resolves
it deterministically, and the repo scripts get npm aliases (warm-cache,
check-cache-key, query-matrix).
development had moved ahead of main with fixes to the same sample templates this
branch restructured, so the three conflicts needed care rather than a default
resolution.

src/engine/templates/fusedQueries.ts — development replaced the
coso:measurementValue join with resultValueClauses() so non-detects stop being
dropped, and switched the aggregate's MAX to ?numericResult. This branch had
deleted buildFusedSampleAggregateQuery and buildFusedSampleDetailsQuery outright,
because re-deriving the spatial trace inside a hydrate query is what timed out.
Kept the deletion: their fix is not lost, because the same resultValueClauses()
call already applies on the path that replaced those functions —
buildSampleRetrievalByIriQuery and buildSampleDetailByIriQuery in
templates/downstreamSamples.ts — and in bindEntityInCell, where it composes
cleanly inside the new sampleObservations guard. A note in the file records this
so the next reader does not think the fix was dropped.

docs/DEBUGGING.md — development already documented the 429-is-a-timeout finding
on 2026-09-11, including the observation that trimming unused bindings helps.
Kept their entry and trimmed this branch's to what it adds: the sweep numbers, the
bindEntityInCell root cause, and three traps not recorded anywhere else (an
unfiltered second block means every entity in the graph; 413/502 surfaces as a
CORS error; ?timeout=120s returns 403).

Verified after merging, against the live endpoints: both the near and downstream
airport questions still succeed, with target and anchor counts unchanged
(590/45 and 1172/190). Hydrated sample counts rose — 515 to 568, and 999 to 1120
— which is development's fix working: non-detect samples are no longer hidden by
the measurementValue join.
perf(engine): bound query work by splitting on failure
The accent was two families that disagreed. --color-primary, --color-primary-300
and --color-primary-50 were never defined, so a dozen rules fell through to
their hex fallbacks and rendered Tailwind violet next to Chakra blue. Fill the
ramp, alias indigo onto primary, strip all 17 var(--token, #hex) fallbacks, and
map the remaining bare hexes onto gray and status tokens.

Map layer and relationship tag colors are untouched: they encode data
categories, not chrome.

Also switch the system font stack to Inter, with tabular figures so result
counts stop jittering as they update.
Replaces the 'Welcome to Sawgraph!' heading with a band that says what the tool
does: an interface for researchers to explore the knowledge graph without
writing SPARQL. Two ways in, the existing How it works tour and a new analysis.

Dismissable for the visit, deliberately not persisted.

The header also gains a link to the SAWGraph project site, which had an empty
right side on the dashboard. (Its CSS rides along with the palette commit; both
live in App.css.)
Extends the existing docs rule to every string the app shows: tour steps,
glossary definitions, the Maine wells description, the max-distance error
suggestion, progress and partial-result messages, the empty community state.
Recast into commas, colons, parens or new sentences rather than swapped
mechanically. Code comments are left alone.
Swaps Carto light_all and dark_all for Esri World Light/Dark Gray Canvas, with
the matching reference overlay as a second layer and maxNativeZoom 16.
The thumbnail shows the real answer, not a generic map: scripts/build-thumbnails
runs each prebuilt question through the pipeline and draws the result points and
flowlines over an Esri basemap, one committed SVG per question in public/thumbs.
Offline and manual, the way the query matrix is; re-run it when a prebuilt
question changes shape.

The basemap is requested in EPSG:4326 so plotting a point is a linear map from
degrees to pixels, and inlined as JPEG, which is what takes a thumbnail from
250KB to 130KB. Region boundaries are dropped, they are noise at 132px wide.

A card with no generated thumbnail falls back to a plain basemap of its region
rather than a broken image.
fix: upstream questions traced the wrong way, plus a dashboard and theme pass
CommunityAnalysisCard shares .query-card, so putting display:flex there for the
prebuilt thumbnail laid its title, author and tags out in a row. Move the row
layout to .query-card-with-thumb, which only the prebuilt cards use.
fix(dashboard): thumbnail row layout leaked onto the community cards
Both surfaced on "What facilities are upstream from PFOS samples in York and
Cumberland counties (Maine)?", which failed at step 1 with 500 out-of-memory,
sliced into its two counties, and failed both slices.

Join order. buildFusedWhereBody always wrote the anchor side first, and an
upstream question puts block A there. With no region and no industry filter on
block A the query started from every facility in the graph and traced the
national flowline network to find samples in two Maine counties. QLever does
not search join orders exhaustively at these body sizes, so it follows the
query text: writing the constrained side first is the whole fix, same triples
and same variables in a different order. Measured across all 72 hydrology
shape/config pairs: 5 rescued, 0 regressions, 0 rows lost, faster in 50 of 66
comparable runs. Distance-bounded traces and IRI-pinned slices keep the
anchor-first order, the first because boundedTrace embeds the seed twice and
was never measured reordered, the second because a pin already made the anchor
the small side.

Observation joins. bindEntityInCell emitted the whole observation chain
whenever any sample filter was set, including coso:measurementUnit, which a
non-detect does not have. A substance-only question with "include non-detects"
ticked therefore dropped every non-detect: 774 of 1,688 Cumberland PFOS
observations, turning 23 sample points and 13 facilities into 32 and 15 once
only the joins an active filter reads are emitted. The same required join hid
2,860 of 3,334 popup rows in buildSampleDetailByIriQuery; moving it after
resultValueClauses() with the symbol lookup nested inside costs what it did
before. buildEntityProbeQuery now uses the same lean bind, since the probe list
becomes the chunk membership and a narrowing join there removes entities from
every slice.

scripts/check-query-joins.mts asserts both rules on the emitted SPARQL (265
checks) and runs in CI. docs/DEBUGGING.md records the measurements and the two
designs that were tried and rejected, including the sub-SELECT that rescued 5
shapes and broke 7.
GET_FLOWLINE_GEOMETRIES traced a one-sided transitive closure: it seeded from
the cells of every resolved anchor, widened each by sfTouches, and followed the
network to its end. Nothing in the query referenced the targets, so an anchor
sitting near a drainage divide pulled in the whole of the neighbouring basin.

On "What facilities are upstream from PFOS samples in York County, ME" that
drew 122 segments of the Merrimack and 21 of the Winnipesaukee, ~130km away in
a basin no York sample drains from, plus the Presumpscot, the Suncook and the
Powwow. Nothing looked broken. The map just had extra rivers on it, and they
were real rivers.

The closure is now intersected with the flowlines that reach the resolved
targets. Measured live with the real answer sets (1,497 facilities, 200 sample
points): 3,271 flowlines down to 2,516, in the same 7s. 755 removed, 0 added,
so the fix can only take away flowlines that were never on a path from a
facility to a sample. Of those, 684 are in a different basin and 71 lie below
a sample, which is the one visible behaviour change: the drawn river now stops
at the sample instead of running on to the sea.

Written as an intersection of two DISTINCT closures. The direct membership
test, `?flowline hyf:downstreamFlowPathTC? ?_flTarget`, leaves ?flowline
unbound on the left and QLever times out joining on it (31s, "Join on
?flowline"). The reach is a bare TC rather than reflexive: `TC?` was measured
and recovers 0 flowlines while costing 8s.

The bounded branch gets the same intersection. At 25km it went 3,034 -> 2,516,
matching the unbounded set, because connectivity dominates the distance cap at
that range; at 5km the cap still bites (2,314), so the two constraints compose.
The shipped bounded query still drew a Merrimack segment through its cap.

scripts/check-flowline-scope.mts asserts both ends are bound for all 100 shapes
that draw the layer, bounded and unbounded, and runs in CI.
fix(engine): bound the flowline layer at both ends
The one-cell sfTouches expansion that connects an entity to the river network
is applied to exactly one side of a hydrology trace, and which side is an
artefact of how buildFusedWhereBody orders its blocks rather than a modelling
decision. Measured in York County: 22% of PFOS sample points and 18% of
facilities have no flowline in their own S2 cell, so whichever side is not
widened silently loses about a fifth of its entities.

On "What facilities are upstream from PFOS samples?" scoped to York County,
widening the sample side instead of the facility side gives 324 sample points
and 1,060 facilities against the shipped 200 and 1,497. The union of both is
330 and 1,507, and recovers 130 sample points of which 82 carry quantified PFOS
detections up to 2,200 ng/L. Points with a confirmed detection go from 132 to
214.

The plan proposes running both variants and unioning them, because widening
both sides inside one query does not run: two formulations, flat and with the
cell set materialised first, both fail with "tried to allocate 2.6 GB". Added
cost is 2-3s rather than double, since the second variant caches at 1.3s while
the shipped form does not.

Drafted rather than active: it touches all 72 hydrology shapes, moves counts
upward across them, and is the second executor change queued against the same
code path as the retry hook in the 2026-09-17 plan. The two should be designed
together so there is one mechanism rather than two similar ones.

Also records two QLever measurement traps hit while producing the figures: a
nested OPTIONAL containing a BIND returned a count of 0 alongside a non-zero
MAX over the same variable, and an OPTIONAL used for predicate coverage
under-reported silently. Every figure was re-derived with flat COUNT queries.
Brings in docs/plans/drafts/2026-09-19-symmetric-cell-widening.md, which
documents the follow-up work that falls out of the two fixes on this branch.
No code change.
Cleanup pass over the two fixes on this branch. No intended behaviour change
except where noted; the York question emits a byte-identical body for steps 1
and 2 and the same 2,516 flowlines for step 3, verified live.

sideIsConstrained was a third copy of a predicate the repo already had twice,
in scope.ts as hasBlockFilters and in sparqlErrors.ts as isNarrowed. scope.ts
asks the identical question to pick a chunking axis, so the two drifting apart
means the executor slices on an axis the template already led with. scope.ts
imports fusedQueries, so the shared predicate lands here as blockIsFiltered and
scope.ts uses it. sparqlErrors keeps its own, which is this plus the region
check.

That copy also had a defect the original did not: a generic "any field is
non-null" scan counts includeNondetects: true, which is the UI's default
checked state and narrows nothing. A samples block with only that set was read
as constrained and reordered the whole body. sampleJoinsNeeded already draws
the line correctly, so samples delegate to it. Verified: that block now leads
with ?s2anchor, includeNondetects: false and a substance filter still lead with
?s2target.

buildFusedFlowlineQuery re-derived the anchor block over IRIs its own VALUES
list had already pinned, which is exactly what targetReachClause eight lines
below refuses to do. Same bare-block treatment now. For a facilities anchor
that drops a redundant industry VALUES; for a samples anchor it drops the
entire observation chain, including the coso:measurementUnit join this branch
exists to keep out of queries that never read it.

Also: pinValues instead of a fourth hand-rolled VALUES block (it guards the
empty list, the inline form emitted "VALUES ?spC {  }"); needsUnitJoin instead
of restating the range predicate the comment already pointed at; a dead branch
in check-query-joins where both arms returned 'other'; and the reflexive-reach
assertion in check-flowline-scope, which matched the downstream spelling only
and left the upstream shape unguarded by a check whose whole point is the TC?
regression.

Two stale comments corrected: the sampleObservations doc named hydrate queries
that were deleted in 66f3fda, and buildFusedWhereBody still described the
hasSampleFilters override the same commit removed.

Deferred to the plan docs rather than done here, both recorded as tasks:
emitting the body from an ordered fragment list instead of two hand-maintained
renderings (worth doing with the widen option, which would otherwise make it
four), and bounding targetReachClause by region when the target list is too
large to inline. Also added a task to measure the distance-bounded shape
reordered, since !maxDistanceKm is the one exclusion in the gate that is a "we
did not look" rather than a measured "it was worse".
…join rule

The two fixes on this branch were recorded as bugs but not as behaviour. Someone
reading the app had no way to learn what the "Include non-detects" checkbox does,
and the wiki's only mention was three sentences saying it exists.

Wiki: new page "Samples Non Detects", linked from the sidebar, Home and the
Dropdowns index. What a non-detect is (it is a measurement, not a gap: the record
carries detection limits instead of a value, which is why it has no
coso:measurementUnit), that the box is ticked by default, and that expanding
"+ Add Filters" writes nothing and leaves the SPARQL byte-identical.

The behaviour the page exists for: the checkbox filters at two levels, not one.
York County, all substances, 451 sample points and 23,530 observations ticked
against 367 and 7,110 unticked. The 367 surviving points hold 21,822 observations
between them, so unticking also strips 14,712 non-detect rows from popups that
stay on the map. Non-detects are 70% of observations in York and 77% across
Maine, so this is most of the data rather than a trim. Also the four-way point
classification, and why 14 York points have no measurements at all: they are all
EGAD soil sampling locations whose results are not loaded, while every well at
the same site has 24 to 30 observations.

DEBUGGING.md: a 2026-09-19 entry for the stream layer drawing the Merrimack, with
the one-sided closure that caused it, the intersection that fixed it (755
flowlines removed, 0 added), and the two forms measured and rejected. Records
that a distance cap is not a substitute, since the shipped bounded query drew a
Merrimack segment through its own 25km cap. Also the includeNondetects: true
defect from the cleanup pass, where the UI's default ticked state was read as a
narrowed side and reordered the whole body.

ARCHITECTURE.md: a new Key Design Decision on why each filter declares its own
joins, since "every required triple is also a filter" is the principle behind
both bugs and the thing a newcomer needs before writing a template. Pattern 3
gains a note that a closure needs both ends bound.

Corrects an explanation I had wrong in three places. `TC?` recovering 0 extra
flowlines is not because the segment arrives via another target; it is because
hyf:downstreamFlowPathTC is already reflexive, which ARCHITECTURE.md Pattern 3
had said all along. Verified: X TC X holds for 2,104 of 2,104 York County
flowlines, 434,501 self-pairs graph-wide. The measurement and the conclusion are
unchanged, only the reason. One consequence worth having right: the segment a
target sits on is drawn, and what the intersection excludes is the river below it.
…tream layer

Entry 13 covers the join order and the unit join fixed on 2026-09-17, including
that dashboard counts will move up when it deploys and that figures taken off
the app before that date understate PFAS presence.

Entry 14 covers the stream layer drawing the Merrimack, the includeNondetects
defect found while reviewing it, the new wiki page, and the cell-widening gap
that is measured but deliberately not fixed here.
The dot sat at the top of the content box rather than on the text beside it, and
the running step carried its own padding, so a step visibly shifted when it
started. Padding moves onto .timeline-content so every step has it, and the
marker gains a matching top offset.
`maxDistanceKm` appeared zero times in the harness. M4, which the header called
a distance sweep, sweeps the `near` relationship's hop distance in miles; the
hydrology relationship's flow-distance bound in km was swept by nothing, and the
only bounded shape the matrix ever ran was one dashboard prebuilt in M5.

Worse, `mk` attached the region to block A unconditionally. For an upstream
question the planner maps block A to the anchor, so a region on A always left
the anchor as the constrained side: the target-first join order was unreachable
from this harness, and so was every bounded question that leaves block A wide
open, which is the shape users build.

`mk` now takes `{ km, regionOn }`, and MD runs four shapes bounded and unbounded
with the region on each block. It measures both discovery steps rather than step
0 like M1-M4, because the two projections share a WHERE clause but not a fate:
at 30 km the samples projection answered in 25s while the facilities projection
ran out of memory at 32s, and a step-0-only phase would have called that working.

The baseline (32 rows, 11 pass) shows the rule cleanly: a bounded query works
when the region sits on the block that becomes the anchor and fails when it sits
on the target. Upstream makes block A the anchor, downstream makes block C the
anchor, and the mirror pair `facilities upstream(30km) samples [on A]` and
`samples downstream(30km) facilities [on C]` return identical counts, 2,399 and
2,523. Wells is the exception and fails on all four bounded combinations for its
own reasons, as in W38 entry 12.
The constrained-side-first reorder skipped distance-bounded queries, with a
comment saying that shape had not been measured either way. Measured now: every
bounded shape whose constrained side was the target failed, and the bound was
the thing breaking them. "Facilities upstream from PFOS samples in York" with
block A left wide open, at 30km: 429 in query planning at 30s on both
projections. Unbounded, the same question answers.

boundedTrace embeds its seed inside an aggregate that sums path lengths over two
closure hops, so seeding the wide-open side sums them across the national
flowline graph: 1,506,326 facilities. Seeded from the 132 samples instead, the
same question answers in 20s and 3s, returning 132 samples and 918 facilities.

The reorder is the same rule the unbounded path already followed, so the guard
loses one clause. boundedTrace gains a seed side, and the trace, the GROUP BY
and the "+1" fringe all follow it: the fringe still extends away from the seed,
which means the physical end that gets the extra segment now follows the seed
side. Checked across all 1,152 hydrology shape and config pairs: 144 queries
change text, exactly the bounded ones whose constrained side is the target, and
the other 1,008 are byte-identical.

Correctness, before trusting the speed. Where both forms run, they return
identical IRI sets on both projections. Against the unbounded answer the bounded
one is a strict subset, 918 of 1,494 facilities with 0 added. Across 5, 10, 30
and 50km the answers nest, each bound's set contained in the next.

Two checks had to change, and both were rebuilt rather than relaxed.
check-query-joins asserted that bounded traces keep the anchor-first order,
which was the defect; it now asserts the same constrained-side rule as every
other shape, and separately that the aggregate seeds from that side too, since a
query can read target-first while the cost stays where it was.

check-trace-direction is the more delicate one. It kept
`?_flEnd hyf:downstreamFlowPathTC ?_flMid` on a blacklist of reversed traces,
and a target-seeded block writes that exact triple while tracing correctly. That
string cannot decide it any more, so queries naming both ends are now checked by
reachability: is there a directed path from the anchor's flowline to the
target's. Verified by mutation - reintroducing the 2026-09-16 direction bug
(traceDirection derived from the relationship name) still fails the check, and a
half-applied reorder that moves the text but leaves the aggregate seeding from
the anchor fails the new one.
The script existed since 2026-09-17 but was never committed, so package.json on
feat/demo-questions-prewarmed registered `check-query-snapshots` against a file
nobody else had and that branch's CI never ran it. Committing it here, with the
grid extended to cover what this week showed it was blind to.

Why it earns a place next to the six rule-checks: those assert properties of a
query, so they see a shape only if someone thought to assert something about it.
This one notices when a change moves a shape nobody was thinking about, which is
the failure mode of a builder shared by every relationship. Three of the bugs
fixed this week were invisible in behaviour and obvious in the emitted SPARQL.

The grid gains two axes, both of which hid something real:

Region placement. It put the region on block A always. For a downstream
question that makes the target the constrained side, so the target-first join
order was recorded; for an upstream question it makes the anchor constrained, so
that path never was, and neither was the question behind #55 and #57 - block A
wide open, block C scoped to a county.

The flow-distance bound. A bounded trace is a different query rather than a
filtered one: it seeds an aggregate. The grid had two bounded shapes, both with
the region on block A. Measured on the fix in ec221d9: the old grid reports one
moved shape, `VARIANT: downstream maxDistanceKm=10`. The new one reports 74, and
buckets them as bounded 72, region on A 36, region on C 36.

6 x 6 x 3 x 2 unbounded, 6 x 6 x 2 x 2 bounded, plus 15 variants for hop counts,
per-type filter blocks and the two questions that have actually broken in the
app. 375 shapes.

Two things keep the file readable rather than enormous. The PREFIX preamble is
byte-identical in every query and was 66% of it, so it is stripped and recorded
once; a change to the prefix list still moves exactly one entry. And query
bodies are written once each in a table at the end and referenced by hash: 375
shapes reference 2,078 step queries but only 603 distinct ones, the
region-boundary query alone being identical in 373 of them. Together those take
the snapshot from 1.7MB to 688KB and turn a one-line change to a shared query
from 373 lines of diff into one.

Known limit, worth stating rather than discovering later: it drives everything
through planPipeline, so it covers shapes a user can ask for and is blind to
builder arguments the planner never produces. A reordering of the anchor-seeded
`upstream` branch of boundedTrace passes all seven checks, because the planner
always passes 'downstream'. The guarantee is "no reachable shape moved
unintentionally", not "no emitted SPARQL moved".
Four MD sweeps. The result is 2 shapes fixed, 0 regressed, and it took three
extra runs to be able to say that honestly.

The first sweep (002e529, committed earlier) was cold, and the sweep taken after
the fix was warm, so comparing them credited the fix with three flips. One of
those three is `facilities upstream samples [ME on A] step1`, which is unbounded
and anchor-constrained: the reorder cannot reach it, and the generated SPARQL is
byte-identical at 1,950 characters either side of ec221d9. It times out cold and
answers in 15s warm. That row is the cache, not the code. The cold sweep now
carries a .note saying so.

Re-running the pre-fix engine warm (e84a11f's src, current harness) gives the
comparison that means something:

  pre-fix warm   12/32 pass
  post-fix warm  14/32 pass, identical in three consecutive runs

  FIXED  facilities upstream(30km) samples [ME on C] step1   oom -> 3213 rows, 17s
  FIXED  samples downstream(30km) facilities [ME on A] step1 oom -> 3213 rows, 15s
  REGRESSED  none
  row-count drift where both pass  none

Both fixed rows are exactly the condition the reorder covers: bounded, with the
constrained side on the target. All 8 bounded rows whose constrained side is the
anchor are unchanged, which is the guard working.

Three things the sweep says that the spot-check could not.

Run-to-run noise is zero here. Three sweeps at identical code disagree on no row
and report row counts identical to the unit, so a single flip in this file is
signal, provided both halves are warm.

Step 0 of both fixed shapes still fails. The fix lands on step 1 both times.
These are statewide Maine with no filters, where the samples projection is the
far heavier half; the York County question this started from passes both steps
in 20s and 3s. Question size, not the reorder, but it means "bounded upstream
questions work now" is false as a general claim.

`facilities upstream(30km) streams [ME on C]` did not start passing. It changed
failure mode, timeout to out-of-memory on step 0 and to a sort-estimate timeout
on step 1. Worth knowing before anyone reads the two fixes as covering the
bounded target-constrained case generally.

Wells is unchanged on all four bounded combinations, as it was before, for its
own reasons (W38 entry 12).
QUERY-MATRIX.md gains Part 10: what the flow-distance bound was doing, the
one-variable isolation (block A wide open gives a 4.3 GB allocation failure,
block A scoped to York County gives 1,429 rows in 4s), the warm-against-warm
before and after, and what the fix does not do.

DEBUGGING.md gains a 2026-09-20 entry with the root cause, the reachability
tripwire that replaced a string blacklist a correct query now trips, and the two
measurement traps this cost: comparing a cold sweep against a warm one, and
measuring an endpoint while your own pipeline run is hammering it.

W38 entries 15 and 16. Entry 15 states the limits in the entry rather than
leaving them to be found: step 0 of the statewide shapes still fails, both
rescued rows are step 1, the streams shape changed failure mode rather than
passing, and wells is untouched.
Fix upstream questions: out-of-memory, dropped non-detects, and rivers in the wrong basin
Adding an industry filter to block A turned a working question into a failing
one. "What facilities are within 5 km upstream from PFOS samples in York
County" answers in 3s with block A empty; add NAICS 221320 to it and both
projections time out in query planning at 30s.

`sideIsConstrained` was a boolean: region, filter or IRI pin, any of them means
constrained. An industry filter set it to true, the reorder stopped firing, and
the query went back to seeding from block A, which now meant every sewage
treatment facility in the country rather than the 118 sample points in one
county. The comment two lines above it already said "narrowed is not small"; the
predicate did not act on its own warning.

It now returns a rank, and the body leads with the higher one:

  3  an IRI pin. The executor chose this slice.
  2  a region. Bounded by geography, and the graph is indexed for it.
  1  an entity filter only. Selective in kind, unbounded in extent.
  0  nothing.

Ties keep the anchor-first order, so this only changes cases where both sides
are constrained at different levels. Everything where one side was constrained
and the other was not keeps the plan it had, which the snapshot confirms: of 375
shapes, 2 move, both variants with a region on one side and a bare filter on the
other.

Measured on the live endpoint, all with block A carrying no region:

  1 code,  unbounded   429 planning timeout  ->  35 facilities, 16s
  1 code,  5 km        429 planning timeout  ->  22 facilities,  8s
  8 codes, 5 km        429 planning timeout  ->  88 facilities,  3s
  8 codes, unbounded   429 planning timeout  -> 192 facilities, 15s

21 codes still fails. That is a different cost, the `|| EXISTS { ?ic
fio:subcodeOf ?sel }` clause evaluated per selected code and duplicated inside
the bounded aggregate, and it is not addressed here.

One shipped dashboard question changes plan. `samples-downstream-waste-indiana`
has samples scoped to Indiana and NAICS 5622 with no region, so its anchor is
nationwide and its target is one state. Both orders were run: identical answers,
136 samples and 415 facilities either way, reordered no slower at 13s and 2s
against 16s and 3s. `check-query-joins` asserted that no prebuilt may reorder,
which described the old predicate rather than a rule; it now expects this one to
and says why.

30 new assertions pin the ranking in both directions, and a mutation reverting
it to the boolean fails them.
QUERY-MATRIX.md Part 11: what the boolean predicate cost, the rank that replaced
it, the measured before and after with block A unscoped, and the two things it
does not fix.

DEBUGGING.md gains a second 2026-09-20 entry, with the one-variable isolation
table showing that block A had to be small rather than merely filtered, and the
note that a long industry selection fails for a different reason.

W38 entry 17. Written to be read next to entry 15, since this is the same class
of bug one level up and was found by doing the obvious next thing with the map
entry 15 fixed.
…r it

Yesterday's entries said industry selections beyond about ten codes fail, and
blamed the `FILTER(?ic = ?sel || EXISTS { ?ic fio:subcodeOf ?sel })` clause.
Both halves were wrong, and both were written from a measurement taken before
the narrowness ranking landed.

21 codes answers fine at 5 km: 101 facilities in 13s, 84 samples in 23s. It
fails only unbounded, at 430 MB. Code count is a cost on the closure, and a
distance bound already caps it, so a long selection is a reason to set a
distance rather than to shorten the list.

    codes   5 km                    unbounded
    1       22 facilities,  8s      35 facilities, 16s
    8       88 facilities,  3s     192 facilities, 15s
    21     101 facilities, 13s     500, out of memory at 430 MB

The clause is cleared as well, by trying the fix rather than reasoning about it.
`fio:subcodeOf` is materialised transitively (325110 is a direct subcode of 32,
325 and 3251), so `?ic fio:subcodeOf? ?sel` is an exact rewrite of the FILTER.
Measured on the 21-code bounded query: 13s with the FILTER, query-planning
timeout with the path. The FILTER is the fast form. Recorded as a rejected
design so nobody optimises it again.
Narrowing a question by industry made it fail: rank how narrow a side is, instead of asking whether it is
@railway-app

railway-app Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the explorer-app-pr-60 environment in sawgraph-explorer

Service Status Web Updated
sawgraph-api 🕒 Building (View Logs) Web Sep 25, 2026 at 1:40 am UTC
sawgraph-web 🕒 Building (View Logs) Web Sep 25, 2026 at 1:40 am UTC
1 service not affected by this PR
  • sawgraph-db-dev

@railway-app
railway-app Bot temporarily deployed to sawgraph-explorer / explorer-app-pr-60 September 25, 2026 01:40 Destroyed
@prayaslashkari
prayaslashkari merged commit ca2be08 into main Sep 25, 2026
2 of 4 checks passed

This branch was successfully deployed

1 inactive deployment
sawgraph-explorer / explorer-app-pr-60 — 01a68995 Deployed Sep 25, 2026 by railway-app[bot]
sawgraph-explorer / development — 01a68995 Deployed Sep 20, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants