Jeremy Howard @howard.fm · Dec 19

6 years after BERT, we have a replacement: ModernBERT! @answerdotai, @LightOnIO (et al) took dozens of advances from recent years of work on LLMs, and applied them to a BERT-style model, including updates to the architecture and the training process, eg alternating attention.

26 likes 1 replies

?

Replies

Jeremy Howard · Dec 19

Fancy GenAI stuff like GPT 4 is too big, slow, private, and expensive for many jobs. Consider that the original GPT-1 was 117m params. Llama 3.1, by contrast, has up to 405 billion params! 😲 These models are slow, expensive, and *not yours to control*.