moogle.dev @moogle.dev · Jun 8

Testing Parallelism with Gemma 12B #llm #llama.cpp #llamacpp on RTX 5080 mmm (laptop) but really, anime was a mistake. gotta say the performance for a 12B DENSITY is absolute crazy. Qwen 9B would struggle a bit, google be cooking a lot.

0 likes 0 replies

?