Chrissy LeMaire @funbucket.dev · Mar 29

Interesting difference. The table below shows idle vs. maxed out doing work/inference, not just a loaded model. Before inference, only memory spiked as the model was loaded into memory. Not GPU or CPU. As an infrastructure person and toolmaker newish to AI, I find this all fascinating!

6 likes 0 replies

?