Found while fixing #492 (PR #503) and deliberately left out of it. Both halves verified against the live page just now, on main at d2d9bb80c — not inferred from the code.
Two independent breakages in one crawler
pcapkit/vendor/ipx/socket.py scrapes https://en.wikipedia.org/wiki/Internetwork_Packet_Exchange#Socket_number.
1. Wikipedia rejects requests' default User-Agent.
default-UA GET -> HTTP 403
browser-UA GET -> HTTP 200, 113743 bytes
So the crawler cannot fetch the page at all as written. Wikipedia has been progressively blocking generic library User-Agents; this is not specific to this repo.
2. Even with a working fetch, the table it wants is gone. The page now carries exactly one wikitable, and it is the wrong one:
wikitable count on the live page: 1
table 0: headers ['Octets', 'Field']
That is the IPX header format table. The well-known-socket-number table this crawler parsed has been removed from the article. So there is nothing left to scrape even after fixing the User-Agent — the data source itself is gone.
What #503 did instead, and why that is not a fix
#503 needed Socket(0x0000) to exist, because it is IPX's own class default and a legitimate wire value (see #492). Rather than depend on a scrape that currently yields nothing, process() now prepends the member unconditionally, with a comment saying so.
That is deliberately a workaround. It gets the one needed member in regardless of the scrape, but it does not make the crawler work, and the next regeneration will still lose every socket the scrape used to supply. Nobody should read #503 as having fixed this.
Options, none of them free
- Fix the User-Agent and find a new source. The IANA-style authority for IPX socket numbers no longer really exists — Novell's registry is long defunct — so the realistic sources are an archived revision of the Wikipedia article, or a hand-maintained table checked into the vendor script.
- Pin to an archived revision.
oldid= on the Wikipedia URL, or a Wayback snapshot. Stable and reproducible, but frozen; acceptable given the registry is effectively closed.
- Retire the scrape and hand-maintain the enum, as a deliberate decision recorded in the vendor script. Given IPX is a dead protocol with a closed socket registry, this may honestly be the right answer — a crawler for data that will never change again is cost without benefit.
My own leaning is 3 or 2, precisely because IPX is not going to gain new socket numbers. But that is a maintenance-policy call rather than a technical one, so it belongs to whoever owns the vendor pipeline.
Worth checking at the same time
Other vendor crawlers may hit the same 403. This one was noticed only because #492 sent someone into it. A sweep of pcapkit/vendor/** for crawlers that fetch with a default User-Agent would size that, and is not something I have done — this issue reports one crawler, verified, and nothing wider.
Found while fixing #492 (PR #503) and deliberately left out of it. Both halves verified against the live page just now, on
mainatd2d9bb80c— not inferred from the code.Two independent breakages in one crawler
pcapkit/vendor/ipx/socket.pyscrapeshttps://en.wikipedia.org/wiki/Internetwork_Packet_Exchange#Socket_number.1. Wikipedia rejects
requests' default User-Agent.So the crawler cannot fetch the page at all as written. Wikipedia has been progressively blocking generic library User-Agents; this is not specific to this repo.
2. Even with a working fetch, the table it wants is gone. The page now carries exactly one
wikitable, and it is the wrong one:That is the IPX header format table. The well-known-socket-number table this crawler parsed has been removed from the article. So there is nothing left to scrape even after fixing the User-Agent — the data source itself is gone.
What #503 did instead, and why that is not a fix
#503 needed
Socket(0x0000)to exist, because it is IPX's own class default and a legitimate wire value (see #492). Rather than depend on a scrape that currently yields nothing,process()now prepends the member unconditionally, with a comment saying so.That is deliberately a workaround. It gets the one needed member in regardless of the scrape, but it does not make the crawler work, and the next regeneration will still lose every socket the scrape used to supply. Nobody should read #503 as having fixed this.
Options, none of them free
oldid=on the Wikipedia URL, or a Wayback snapshot. Stable and reproducible, but frozen; acceptable given the registry is effectively closed.My own leaning is 3 or 2, precisely because IPX is not going to gain new socket numbers. But that is a maintenance-policy call rather than a technical one, so it belongs to whoever owns the vendor pipeline.
Worth checking at the same time
Other vendor crawlers may hit the same 403. This one was noticed only because #492 sent someone into it. A sweep of
pcapkit/vendor/**for crawlers that fetch with a default User-Agent would size that, and is not something I have done — this issue reports one crawler, verified, and nothing wider.