fix(indexing): index hand-written /api/ pages as prose, not endpoints - #816
Merged
Merged
Conversation
The indexer called any /api/<api>/<slug> route an endpoint unless its slug contained "overview", and rebuilt it from OpenAPI schema files. Hand-written pages such as /api/platform-api/getting-started and /authentication have no schema files, so they were indexed as an empty endpoint template: a duplicated title, an empty "## Endpoint" section, and none of the page text (234 characters for a 2.5 KB page). Classify a route as an endpoint when the OpenAPI generator wrote its <slug>.api.mdx, the same source the schema files come from. Against the current build this gives 167 endpoints and 217 prose pages; the two platform pages go from endpoint to prose and nothing else changes. The docs MCP provider had its own copy of the slug rule and it had already drifted: it treated every platform-api page as prose, so each of the 36 platform endpoint fetches missed on the id lookup and made a second Glean call by URL. It now requests both possible ids in one call, so it does not need the rule at all.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The devdocs indexer treats any
/api/<api>/<slug>route as an endpoint unless the slug containsoverview. It then rebuilds the page from OpenAPI schema files. Hand-written pages under/api/have no schema files, so they get indexed as an empty endpoint template: a duplicated title, an empty## Endpointsection, and a one-line description..mddocs_fetch/api/platform-api/getting-started/api/platform-api/authenticationThe docs MCP provider (
api/glean-search-provider.mjs) had its own copy of that slug rule, and the two copies had already diverged. The provider only applied the rule toclient-apiandindexing-api, so it treated every Platform API page as prose. As a result, each of the 36 Platform API endpoint fetches asked for the wrong document id, missed, and made a second Glean call by URL.Change
scripts/indexing/data_client.py): a route counts as an endpoint when the OpenAPI generator wrote itsdocs/api/**/<slug>.api.mdx. That's the same source the schema files come from.ApiRoute.is_overviewis removed.getdocumentsfor both possible ids (infoPageandapiReference) in a single call and keeps whichever one exists. Every page type now takes one call. The URL fallback stays for pages that neither id finds./api/. The README describes the rule.Verification
getdocuments. They cover a guide, prose under/api/, and an endpoint page, each found in one call, plus the URL fallback, a missing page, an empty document, and a 429. Onmain, the endpoint and fallback cases fail because main makes two lookups.pnpm test(46 files, 389 tests),typecheck,format:check,snippets:check, the build, indexerpytest(80 tests), andglean-idx test --phase mock(384 documents) all pass.Rollout
Document ids include the object type, so the next index run moves the two prose pages to
infoPageids. The full upload's stale-document deletion removes the oldapiReferenceentries. The provider requests both ids, sodocs_fetchworks before and after that run.