fry69 @fry69.dev · May 18

How fast? Well, I get ~30 token/s on my M5 Max 128GB reliably. Tested with context sizes up to 256k The speed goes down with longer contexts of course, see here for a benchmark ->

4 likes 1 replies

?

Replies

jake · May 19

"the speed goes down with longer contexts" is carrying a lot of water. do the math on the area under the curve for prefill on 256k tokens, and see that your laptop is probably maxing out its cores for ~45min just to get to first token. from there, sure, speed is good, for same KV-cache/prefix