deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m by mosheisland · Pull Request #772 · deepspeedai/DeepSpeedExamples

mosheisland · 2023-10-17T06:41:18Z

Bloom-560m model has high variance in its last LN layer weight. This causes accuracy issues in bf16 stage2 training. Therefore, reset the parameters of the last LN layer before training. This is a good practice in any case where we replace the classifier that follows the LN.

In addition, in case we are using only optimize lora, we need to force the training of the LN parameters that were reset.

Note that current fix uses plain initialization of final LN. A separate commit will provide support for zero3 initialization.

Change-Id: I323d8947907eb4a1cc0fa6354bdaf0cbbf33a68d

Bloom-560m model has high variance in its last LN layer weight. This causes accuracy issues in bf16 stage2 training. Therefore, reset the parameters of the last LN layer before training. This is a good practice in any case where we replace the classifier that follows the LN. In addition, in case we are using only optimize lora, we need to force the training of the LN parameters that were reset. Note that current fix uses plain initialization of final LN. A separate commit will provide support for zero3 initialization. Change-Id: I323d8947907eb4a1cc0fa6354bdaf0cbbf33a68d Signed-off-by: Moshe Island <misland@habana.ai>

lekurile

LGTM!

Bloom-560m model has high variance in its last LN layer weight. This causes accuracy issues in bf16 stage2 training. Therefore, reset the parameters of the last LN layer before training. This is a good practice in any case where we replace the classifier that follows the LN. In addition, in case we are using only optimize lora, we need to force the training of the LN parameters that were reset. Note that current fix uses plain initialization of final LN. A separate commit will provide support for zero3 initialization. Change-Id: I323d8947907eb4a1cc0fa6354bdaf0cbbf33a68d Signed-off-by: Moshe Island <misland@habana.ai> Co-authored-by: Moshe Island <misland@habana.ai>

mosheisland requested review from RezaYazdaniAminabadi, ShadenSmith, arashb, awan-10, conglongli, duli2012, eltonzheng, jeffra, minjiaz, mrwyattii, samyam, tjruwase, xiaoxiawu-microsoft and yaozhewei as code owners October 17, 2023 06:41

tjruwase requested review from lekurile and removed request for RezaYazdaniAminabadi, ShadenSmith, arashb, conglongli, duli2012, eltonzheng, jeffra, minjiaz, mrwyattii, samyam, tjruwase, xiaoxiawu-microsoft and yaozhewei October 17, 2023 13:30

lekurile approved these changes Oct 17, 2023

View reviewed changes

tjruwase merged commit 185e25c into deepspeedai:master Oct 17, 2023

mosheisland deleted the 9_fix_bloom_stage2_bf16_acc branch November 22, 2023 07:52

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m#772

deepspeed-chat: fix bf16 stage2 accuracy for bloom-560m#772
tjruwase merged 1 commit intodeepspeedai:masterfrom
mosheisland:9_fix_bloom_stage2_bf16_acc

mosheisland commented Oct 17, 2023

Uh oh!

lekurile left a comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Conversation

mosheisland commented Oct 17, 2023

Uh oh!

lekurile left a comment

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants