[WIP] Upstream encoder/decoder support based on multiple blocktables #161

afeldman-nm · 2024-04-02T19:41:28Z

vLLM currently supports decoder-only models. This PR

Adds support for T5 (encoder/decoder model) using the latest vLLM Attention wrapper
Augments SequenceGroups with optional encoder sequence representations
Modifies block_manager to support both self- and cross-attention blocktables
model_runner prepare_prompt and prepare_decode can optionally prepare input_metadata for encoder-self-attention and decoder-cross-attention, in addition to decoder-self-attention (current supported behavior)

This PR when finished

Addresses vLLM issue 187 Adding support for encoder-decoder models, like T5 or BART vllm-project/vllm#187
Represents the completion of work started by js8544 to integrate encoder/decoder support Add Encoder-decoder model support and T5 Model support vllm-project/vllm#3117
Supports work on Whisper integration [WiP] Whisper Implementation #147 Whisper support vllm-project/vllm#180

BEFORE SUBMITTING, PLEASE READ THE CHECKLIST BELOW AND FILL IN THE DESCRIPTION ABOVE

PR Checklist (Click to Expand)

Thank you for your contribution to vLLM! Before submitting the pull request, please ensure the PR meets the following criteria. This helps vLLM maintain the code quality and improve the efficiency of the review process.

PR Title and Classification

Only specific types of PRs will be reviewed. The PR title is prefixed appropriately to indicate the type of change. Please use one of the following:

[Bugfix] for bug fixes.
[CI/Build] for build or continuous integration improvements.
[Doc] for documentation fixes and improvements.
[Model] for adding a new model or improving an existing model. Model name should appear in the title.
[Frontend] For changes on the vLLM frontend (e.g., OpenAI API server, LLM class, etc.)
[Kernel] for changes affecting CUDA kernels or other compute kernels.
[Core] for changes in the core vLLM logic (e.g., LLMEngine, AsyncLLMEngine, Scheduler, etc.)
[Hardware][Vendor] for hardware-specific changes. Vendor name should appear in the prefix (e.g., [Hardware][AMD]).
[Misc] for PRs that do not fit the above categories. Please use this sparingly.

Note: If the PR spans more than one category, please include all relevant prefixes.

Code Quality

The PR need to meet the following code quality standards:

We adhere to Google Python style guide and Google C++ style guide.
Pass all linter checks. Please use format.sh to format your code.
The code need to be well-documented to ensure future contributors can easily understand the code.
Include sufficient tests to ensure the project to stay correct and robust. This includes both unit tests and integration tests.
Please add documentation to docs/source/ if the PR modifies the user-facing behaviors of vLLM. It helps vLLM user understand and utilize the new features or changes.

Notes for Large Changes

Please keep the changes as concise as possible. For major architectural changes (>500 LOC excluding kernel/data/config/test), we would expect a GitHub issue (RFC) discussing the technical design and justification. Otherwise, we will tag it with rfc-required and might not go through the PR.

What to Expect for the Reviews

The goal of the vLLM team is to be a transparent reviewing machine. We would like to make the review process transparent and efficient and make sure no contributor feel confused or frustrated. However, the vLLM team is small, so we need to prioritize some PRs over others. Here is what you can expect from the review process:

After the PR is submitted, the PR will be assigned to a reviewer. Every reviewer will pick up the PRs based on their expertise and availability.
After the PR is assigned, the reviewer will provide status update every 2-3 days. If the PR is not reviewed within 7 days, please feel free to ping the reviewer or the vLLM team.
After the review, the reviewer will put an action-required label on the PR if there are changes required. The contributor should address the comments and ping the reviewer to re-review the PR.
Please respond to all comments within a reasonable time frame. If a comment isn't clear or you disagree with a suggestion, feel free to ask for clarification or discuss the suggestion.

Thank You

Finally, thank you for taking the time to read these guidelines and for your interest in contributing to vLLM. Your contributions make vLLM a great tool for everyone!

…-project#2802)

…t#2730)

Co-authored-by: Cade Daniel <edacih@gmail.com>

Co-authored-by: zhangdacheng <zhangdacheng@ainirobot.com> Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>

…ject#3046)

Signed-off-by: Tao He <sighingnow@gmail.com>

… pkgs (vllm-project#3070)

…ompt_lens is treated as a list in T5

…oder mode; removed encoder/decoder argument of Sequence

…on of relative position encoding based on packed-variable-length-sequences

…ct T5 inference result. Nothing is broken by this commit, unless there is a subsequent commit with changes in order to pass regression tests.

…ks wrong though. Added not_causal option for attn_bias to kernel interface contracts; also switched to batch size 1 to avoid incorrectness likely caused by packed-variable-sequence-length mask having zeroes rather than -inf's

…adata has correct blocktable, slot_mapping=None, and correct (max) context length(s) (derived from prompt); decode-phase decoder self-attention relative position encoding mask has 1 x K geometry where 1 is the number of new tokens generated in a step and K is context length padded to the nearest multiple of block size, and also mask is reshuffled with contiguous (); ensured general correctness of cross-attention input_metadata; modified T5 example script to prevent HF/vLLM T5 instances from being length limited; net effect: batch-size 1 seems to work but batch-size >1 not supported

andy-neuma · 2024-09-03T17:38:34Z

REPO is getting archived ...

ronensc and others added 30 commits February 21, 2024 18:18

Update comment (vllm-project#2934)

d7f3964

Added early stopping to completion APIs (vllm-project#2939)

5574081

Migrate MistralForCausalLM to LlamaForCausalLM (vllm-project#2868)

344020c

Use Llama RMSNorm custom op for Gemma (vllm-project#2974)

95529e3

chore(vllm): codespell for spell checking (vllm-project#2820)

93dc5a2

Optimize GeGLU layer in Gemma (vllm-project#2975)

fd5dcc5

[FIX] Fix a bug in initializing Yarn RoPE (vllm-project#2983)

c530e2c

Remove Flash Attention in test env (vllm-project#2982)

6f32cdd

Include tokens from prompt phase in counter_generation_tokens (vllm…

4caf704

…-project#2802)

Fix nvcc not found in vlm-openai image (vllm-project#2781)

57f0449

[Fix] Fissertion on YaRN model len (vllm-project#2984)

f7c1234

Port metrics from aioprometheus to prometheus_client (vllm-projec…

ef978fe

…t#2730)

Add LogProbs for Chat Completions in OpenAI (vllm-project#2918)

70f3e8e

Optimize Triton MoE Kernel (vllm-project#2979)

cfc15a1

Co-authored-by: Cade Daniel <edacih@gmail.com>

[Minor] Remove gather_cached_kv kernel (vllm-project#3043)

d6e4a13

[Minor] Remove unused config files (vllm-project#3039)

d9f726c

Don't use cupy when enforce_eager=True (vllm-project#3037)

c1c0d00

Fix stablelm (vllm-project#3038)

4dd6416

Support Orion model (vllm-project#2539)

48a8f4a

Co-authored-by: zhangdacheng <zhangdacheng@ainirobot.com> Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>

fix get_ip error in pure ipv6 environment (vllm-project#2931)

2410e32

[Minor] Fix type annotation in fused moe (vllm-project#3045)

4bd18ec

Support logit bias for OpenAI API (vllm-project#3027)

e0ade06

[Minor] Fix StableLMEpochForCausalLM -> StableLmForCausalLM (vllm-pro…

8b430d7

…ject#3046)

Enable GQA support in the prefix prefill kernels (vllm-project#3007)

71bcaf9

Signed-off-by: Tao He <sighingnow@gmail.com>

multi-lora documentation fix (vllm-project#3064)

a868310

Restrict prometheus_client >= 0.18.0 to prevent errors when importing…

e46fa5d

… pkgs (vllm-project#3070)

[Neuron] Support inference with transformers-neuronx (vllm-project#2569)

3b7178c

Add LoRA support for Gemma (vllm-project#3050)

929b4f2

t5-small

dd82ba3

fix

f2fd579

afeldman-nm added 16 commits March 22, 2024 15:21

t5 Sampler does not pass vocab size to constructor; input_metadata.pr…

cbfba8e

…ompt_lens is treated as a list in T5

add_request now correctly swaps decoder_prompt, prompt in encoder/dec…

501551c

…oder mode; removed encoder/decoder argument of Sequence

Added cross_input_metadata field to InputMetadata

08435e4

wip multi blocktable

6e459a2

wip

8e1ca33

plumbing dummy input metadata structures into model

e097732

plumbed encoder/decoder input metadata all the way into t5

2a44585

first pass at T5 encoder support

91a4608

inefficient but effective & Attention-wrapper-compatible implementati…

d0c5e36

…on of relative position encoding based on packed-variable-length-sequences

wip cross-attention

3737d5b

first pass at enc/dec support that runs e2e but doesn't produce corre…

38946ed

…ct T5 inference result. Nothing is broken by this commit, unless there is a subsequent commit with changes in order to pass regression tests.

to pass regression tests: removed debug prints

3c39f55

wip vllm, examples => fp32

4ec2fde

works on bsz = 1

38f55ed

wip

c1258b4

afeldman-nm marked this pull request as draft April 2, 2024 19:45

afeldman-nm added 11 commits April 3, 2024 14:19

passing with t5-small

0af1022

refactoring out print statements

f5242a0

fix to pass regression tests

de0fd31

WIP google/flan-t5-xxxx

5a67647

removed print statement

ed05d47

batched enc/dec example

d5a8b92

wip, trying prompt padding

f555f5d

bs >1 prefill works

2c12b44

small change to examples

dba02b2

fix to support case where num prompts != 2

db201b6

andy-neuma closed this Sep 3, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[WIP] Upstream encoder/decoder support based on multiple blocktables #161

[WIP] Upstream encoder/decoder support based on multiple blocktables #161

afeldman-nm commented Apr 2, 2024

andy-neuma commented Sep 3, 2024

[WIP] Upstream encoder/decoder support based on multiple blocktables #161

[WIP] Upstream encoder/decoder support based on multiple blocktables #161

Conversation

afeldman-nm commented Apr 2, 2024

PR Title and Classification

Code Quality

Notes for Large Changes

What to Expect for the Reviews

Thank You

andy-neuma commented Sep 3, 2024