Conversation
Due to float flooring math, the computed gradient_accumulation_steps value can be zero. Set min to 1.
|
Example: Without fix, the following training config will generate zero gradient_accumulation_steps torchrun nodes = 8 |
|
Yep. For instance, in your case, you'll end up with a batch size of 32x8 = 256 and not 128. This PR still can be useful as a safety in the mean time, but this more of hiding a misconfig. |
|
What do you think of an alternative fix where:
|
Yep, what I had in mind was a consistency check of the entangled params, warning and auto fix if possible. A few sanity checks could avoid quite some issues there. |
Due to float flooring math, the computed gradient_accumulation_steps value can be zero. Set min to 1.