walking mirage @wolf.observer · May 23

because the finetuning tweets i used were massively outweighed by the original 40GB training corpus, all they did was influence the *flavor* of the text emitted by the model. without them, it emits random semi-coherent text; with them, it emits random semi-coherent text that resembles tweets.

13 likes 1 replies

?

Replies

walking mirage · May 23

at 355 million parameters, the model was a tiny fraction of today's LLMs, except the very smallest of the category designed to run on mobile devices. you could do something similar with modern foundation models, but I think it'd be boring; they're far too coherent. they know too much.