vLLM-Omni DLO mental model

Resident blocks stay.
Tail blocks flow through buffers.

Example: --dlo-resident-layers 20 on a 50-block DiT.

Important: streamed blocks are not copied back to CPU after compute. Their CPU weights already exist. The GPU buffer is simply overwritten by a later block.

Animation

1. Weight placement

Blocks 0–19: resident Current streamed block Other CPU-backed tail blocks
Block 0Block 49

2. What is physically happening

CPU pinned memory

Master/offloaded tail weights remain here.

GPU resident region — Blocks 0–19

Loaded once when denoising begins; reused across all denoising steps.

CPU → H2D prefetch → GPU buffer A
GPU buffer A
Block 20
compute now
GPU buffer B
Block 21
prefetch next

3. Ping-pong buffer timeline

Buffer A holds block 20, then later gets overwritten by block 22. Buffer B holds block 21, then later gets overwritten by block 23. No GPU→CPU weight copy is needed.

4. Across denoising steps

Blocks 0–19

Stay on GPU throughout the denoising stage.

Blocks 20–49

Stream from CPU again on every denoising step.

After denoising

Resident layers can be released before VAE decode so VAE reuses HBM.