Skip to content

Fix percent-encoded base64 data URL reads - #2222

Open
bIackr0se wants to merge 1 commit into
fsspec:masterfrom
bIackr0se:fix/percent-encoded-base64-data
Open

bIackr0se wants to merge 1 commit into
fsspec:masterfrom
bIackr0se:fix/percent-encoded-base64-data

Conversation

@bIackr0se

Copy link
Copy Markdown
Contributor

Reading data:;base64,%5A%6D%39%76 returns six incorrect bytes instead of b"foo". Escaped padding and + or / characters can also cause padding errors.

Percent-decode the URL body before base64 decoding. This follows the decoding order in the WHATWG data URL processor.

with fsspec.open("data:;base64,%5A%6D%39%76", "rb") as f:
    print(f.read().hex())
# Before: e40e83dfdefa
# After:  666f6f (b"foo")

The tests cover escaped padding, alphabet characters, literal +, all 256 byte values, file size and byte ranges. Unescaped base64, empty bodies and decoded bytes containing a literal percent sequence have compatibility checks.

Verification on macOS arm64 with Python 3.14.7:

  • Six regression cases fail against the upstream decoder and pass with the fix.
  • Data, core and API tests: 105 passed, 6 skipped, 1 expected failure.
  • All repository pre-commit hooks passed.
  • A built wheel installed in a fresh environment passed 11 consumer cases, including binary reads, copies to local files, metadata, ranges and a UTF-8 text read.

@bIackr0se
bIackr0se marked this pull request as ready for review October 7, 2026 22:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant