Smaine Kahlouch (Smana) · May 14
The stack is built for evolution:
✅ vLLM on NVIDIA L4 (fp8)
✅ Crossplane InferenceService (1 YAML = the stack)
✅ KEDA scaling on GPU batch/KV signals
✅ Envoy AI Gateway + Iris for smart semantic routing
✅ Amazon S3 Files for shared POSIX storage