Benjamin Warner @benjaminwarner.dev · Feb 10

When we finetune ModernBERT-Large-Instruct on task specific datasets, the generative MLM head is better or nearly equal to standard classification heads.

0 likes 1 replies

?

Replies

Benjamin Warner · Feb 10

Can all encoders be instruction-tuned? Can we replicate ModernBERT's results with an older model like RoBERTa or peer model like GTE-en-MLM? No. And it's not close.