#LLM performance on #AMD Radeon AI PRO R9700 (32GB VRAM) using #ROCm 7.2: input/output tokens per second and VRAM usage at a 128k context length with Q8 KV cache. Ollama + Q4_K_M models.
0 likes 0 replies
?
#LLM performance on #AMD Radeon AI PRO R9700 (32GB VRAM) using #ROCm 7.2: input/output tokens per second and VRAM usage at a 128k context length with Q8 KV cache. Ollama + Q4_K_M models.
0 likes 0 replies