Skip to content

fix(huggingface): preserve batching when pad_token_id is zero - #2605

Open
loyce-cheng wants to merge 1 commit into
SeldonIO:masterfrom
loyce-cheng:bug/huggingface-zero-pad-token
Open

loyce-cheng wants to merge 1 commit into
SeldonIO:masterfrom
loyce-cheng:bug/huggingface-zero-pad-token

Conversation

@loyce-cheng

Copy link
Copy Markdown

Description

Preserve the configured HuggingFace pipeline batch size when the tokenizer's pad_token_id is 0. The truthiness check treated this valid ID as missing, either reducing the batch size to 1 or overwriting the existing padding token with the EOS fallback.

Changes Made

  • Check explicitly for None before applying the missing-padding fallback.
  • Add five model-download-free regression cases covering zero/nonzero padding IDs, preservation with and without EOS, and the existing single-batch fallback when both tokens are missing.
  • Correct the tiny-BERT integration test to expect its requested batch size of 10, since its padding ID is 0.

Related Issues

Fixes #2251.

Validation

  • Before the fix, the two new zero-padding cases fail: one changes batch size 8 to 1, the other overwrites padding ID 0. The other three cases pass.
  • After the fix: pytest runtimes/huggingface/tests/test_common.py -k 'not test_load_pipeline' -q selects 29 passing tests, including the existing tiny-BERT cases. The two existing distilgpt2 loader/export cases were not run.
  • A separate CPU smoke test using a locally created tiny BERT confirms that two differently sized inputs are processed together in one padded batch after the fix; before it, they run in two separate single-item batches.
  • Black, Flake8, git diff --check, and mypy (--follow-imports=silent) pass for the two changed files.

Local testing used Python 3.12 on Windows with the runtime's pinned Transformers 4.41.2. MLServer imports POSIX SIGQUIT on startup, so a local launcher temporarily aliases that constant solely to import the real modules; no server startup or signal handling is tested. This launcher is outside the PR. Linux CI and the full runtime suite remain unverified locally.

Checklist

  • Code follows the project's style guidelines
  • Targeted tests related to the change pass
  • No documentation changes are necessary
  • Code is reviewed by at least one other team member
  • No breaking changes

@CLAassistant

CLAassistant commented Oct 2, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

batch_size is forced to 1 when tokenizer. pad_tokeni_id is 0 in huggingface runtime

2 participants