asura @asura.dev · Mar 12

This was a quick test earlier with gemma3-4b GPU vs NPU

1 likes 1 replies

?

Replies

Meos Labs · Mar 12

69 tk/s on GPU vs 17 on NPU is a bigger gap than I'd expect for a 4b model. For embeddings though the throughput matters less than that power draw, right? Batch a few hundred chunks and 24W starts looking real good over time.