Skip to content

[Question] Question about the 1,800+ Sequence Batch Size in Section 6 #305

Description

@qaqjx

First of all, thank you for the excellent work and for open-sourcing BatchGen. We really appreciate the detailed system design and results.

Image

We are trying to better understand the memory accounting behind the 8×H20 results in Section 6, especially the statement that baselines are limited to batch sizes of 8–16 sequences, while BatchGen can use 1,800+ sequences by offloading KV cache to host memory.

Could you clarify what exactly the reported 1,800+ sequences batch size refers to?

In particular, does it mean the MoE-side batch size after combining multiple attention sub-batches, or the number of actively decoding sequences whose current KV caches are resident in GPU memory?

During decoding, is the attention computation split into smaller sub-batches, where each sub-batch restores the required KV cache from host memory, computes the hidden states, buffers them for the later MoE COMBINE step, and then evicts/offloads the corresponding KV cache from GPU memory?

Thank you again for your time, and we look forward to your reply.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions