Will Held @williamheld.com · Nov 7

Yes! Token packing has been the standard since RoBERTa. Excerpt below! The intuition is that the model quickly learns to not attend across [SEP] boundaries and packing avoids "wasting" compute on padding tokens required to make the variable batch size consistent.

18 likes 2 replies

?

Replies

Pasquale Minervini · Nov 8

Hey! 🙂 we analysed what happens during pre-training, and for causal LMs, intra-document causal masking helps quite a bit both in terms of pre-training dynamics and downstream task performance: arxiv.org/abs/2402.13991

Will Held · Nov 7

More recent works take the "don't attend across documents" even further and manually mask out the attention across documents. This gives you the compute-efficiency without the "it's weird to attend across documents at all" aspect.