Ryan Heuser @ryanheuser.com · May 16

Shannon measured the information rate of English at ~1 bit per character. According to a byte-level LLM measuring next-character predictability in LLM & human text (diaries, abstracts, dreams, fiction), aligned models produce sub-English information rates & LLM text is more predictable than humans'.

12 likes 1 replies

?

Replies

Ben Schmidt · May 16

What kind of human text? First thing that occurs to me with this is that copy-edited human text should have significantly lower entropy than un-copy-edited.