Skip to content

Refuse GoogleOther with a bare 403 - #6

Merged
ralyodio merged 1 commit into
mainfrom
fix/refuse-googleother-ua
Sep 25, 2026
Merged

ralyodio merged 1 commit into
mainfrom
fix/refuse-googleother-ua

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Follow-up to #5. GoogleOther is still crawling at ~120/min (1,768 × 200 and 1,230 × 402 in its last 3,000 requests on dev2), because Google caches robots.txt for up to a day.

The gate now answers any GoogleOther user agent with a 45-byte 403 before the gateway or throttle, except /robots.txt, which it needs to read to learn the refusal. Googlebot (Search) is untouched.

Tests: 72/72 against a real Postgres, including a new test (GoogleOther: 403 on /api and /servers, 200 on /robots.txt; Googlebot: 200 on /servers). Lint is clean.

🤖 Generated with Claude Code

robots.txt refuses it since #5, but Google caches robots.txt for up to a
day, and meanwhile it walks every filter combination of /servers and
/api/v1/search at ~120/min. Answer it in the gate with a few bytes and
no data. robots.txt stays readable so it can learn the refusal;
Googlebot is untouched.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@ralyodio
ralyodio merged commit 28190e7 into main Sep 25, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant