Kubernetes isn’t built for LLMs—until now. Antonio Cardace introduces llm-d: smarter scheduling, KV-cache-aware routing & max GPU usage. Learn how to run distributed inference at scale – fast, flexible, and lock-in free. stackconf.eu/talks/c... #stackconf
1 likes 0 replies
?