Flashes Home Search Notifications Sign in Post
stackconf @stackconf.eu · Apr 22

Kubernetes isn’t built for LLMs—until now. Antonio Cardace introduces llm-d: smarter scheduling, KV-cache-aware routing & max GPU usage. Learn how to run distributed inference at scale – fast, flexible, and lock-in free. stackconf.eu/talks/c... #stackconf

1 likes 0 replies

?

Legal

Privacy Policy Terms and Conditions

Contact

Support FAQs

Follow

Bluesky Instagram Threads
Flashes for Bluesky Get it on the App Store