Conversation
werwrewe
marked this pull request as draft
September 28, 2026 03:00
werwrewe
force-pushed
the
dev-swift-v5-fa3
branch
2 times, most recently
from
September 28, 2026 03:12
6286976 to
e1565d5
Compare
werwrewe
marked this pull request as ready for review
September 28, 2026 03:42
FA2/FA3/FA4 share transformers' single flash_attention_forward entry and differ only in the leaf kernel, but AttentionInterface.get_interface is an exact-key lookup. Registering the SP wrapper only under flash_attention_2 means an FA3 config silently trains without sequence parallelism. Register the wrapper under every FlashAttention name and add CPU-only regression tests.
dadb3f0 moved _dispatch_generation from vllm_sampler_tq.py to generation_submissions.py but left the test importing it from the old location, breaking pytest collection on dev-swift-v5.
werwrewe
force-pushed
the
dev-swift-v5-fa3
branch
from
September 28, 2026 07:00
3db53b1 to
a192898
Compare
werwrewe
marked this pull request as draft
September 28, 2026 07:03
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR type
PR information
What this PR does
Add FlashAttention-3 support for CUDA via the kernelize config entry, as a pure
leaf replacement:
flash_attention_3(kernel/ops/flash_attention3/): registered withbackends=('cuda',), available only on CUDA withflash-attn-3installed.Tests
Qwen3 dense SFT on 2×H800 (SP=2), batch_size=12, sequence length ~4096,
peak memory measured with
torch.cuda.max_memory_allocated():FA3 gives ~9% per-step speedup over the FA2 baseline with identical peak memory
and matching loss/grad-norm curves.