Interesting difference. The table below shows idle vs. maxed out doing work/inference, not just a loaded model. Before inference, only memory spiked as the model was loaded into memory. Not GPU or CPU. As an infrastructure person and toolmaker newish to AI, I find this all fascinating!
6 likes 0 replies
?