Encoder-Prefill Disaggregation#
Vime supports native vLLM Encoder-Prefill Disaggregation (EPD) for
multimodal rollout through encoder server groups.
Data flow#
Encoder engines are started and health-checked before consumers. They are not registered with the regular or PD router.
For each multimodal request:
The rollout client sends one image-only
/v1/chat/completionsrequest per image to an encoder engine.The encoder publishes the image embeddings through ECConnector.
After all images are primed, the rollout continues through the existing
/v1/chat/completions/renderand/inference/v1/generatesequence.With separate prefill and decode groups, the existing PD router and KVConnector continue to coordinate the prefill-to-decode stage.
Text-only samples bypass the encoder and use the existing rollout path.
Supported topologies#
Encoder plus a combined prefill/decode worker:
vllm:
- name: default
server_groups:
- worker_type: encoder
num_gpus: 1
- worker_type: regular
num_gpus: 1
Encoder plus separate prefill and decode workers:
vllm:
- name: default
server_groups:
- worker_type: encoder
num_gpus: 1
- worker_type: prefill
num_gpus: 1
- worker_type: decode
num_gpus: 1
Within one model, Vime rejects incomplete or mixed layouts. In particular,
encoder+decode, encoder+prefill, and mixtures of regular with
prefill/decode are invalid.
EC connector#
The current EPD implementation is single-node only. Vime defaults the connector
to vLLM’s built-in ECExampleConnector and assigns each model an isolated
/dev/shm/vime-ec-* path. The connector creates the directory on its first
write, and the rollout manager removes it on shutdown. /dev/shm is a
memory-backed filesystem, so this minimal implementation avoids ordinary disk
I/O while still exchanging encoder embeddings through safetensors files.
Vime assigns the native EC role from the worker type:
Worker type |
Native EC role |
|---|---|
|
|
|
|
|
|
|
no EC configuration |
As with PD’s KV connector, group overrides retain the highest priority. The
default ECExampleConnector cannot span nodes;
multi-node EPD will be enabled after vLLM lands a production EC connector.
Encoder groups default enable_prefix_caching to false for compatibility
with the current vLLM EC producer path. A group-level override can explicitly
enable it when supported by the selected model and vLLM version.
Encoder-only mode#
Vime enables mm_encoder_only by default for encoder groups, so vLLM skips the
language component and only keeps the multimodal encoder. Weight updates may
still send the full model: vLLM’s weight loader skips the language component
replaced by StageMissingLayer. A group-level override can explicitly set
mm_encoder_only: false when the full model is required on an encoder engine.
Custom rollout contract#
The default rollout, streaming rollout, and Geo3K multi-turn rollout perform priming automatically. A custom multimodal generate function must prime all images before render/generate:
from vime.rollout.vllm_rollout import prime_encoder
await prime_encoder(args, messages, model_name="actor")
The helper is a no-op for text-only and non-EPD deployments. It sends one concurrent request per image, distributes requests round-robin across encoder URLs, and uses the standard HTTP retry policy. Exhausted retries propagate to the rollout; Vime does not silently encode on the consumer.
EPD requires a vLLM build that provides ECExampleConnector and
producer/consumer EC roles.