Lab 11: KServe Inference with HAMi DRA GPU Sharing
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Install HAMi DRA 0.2.3 and the Ascend DRA driver on an Ascend 310P3 node, watch the webhook turn huawei.com/Ascend310P requests into ResourceClaims, and verify two-Pod NPU sharing, memory quota enforcement, and scheduler capacity accounting.
The same outcome through Kubernetes-native Dynamic Resource Allocation (experimental).
Enforce vGPU count, memory, and compute quotas for HAMi workloads before Pods reach the scheduler.