Lab 11: KServe Inference with HAMi DRA GPU Sharing
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Install HAMi on a GPU cluster and schedule SGLang inference services with GPU partitioning.
Install HAMi v2.10.0 and verify per-Pod MIG placement, mixed profiles, selective reclamation, restart recovery, and multi-GPU spillover.
Install HAMi 2.10.0 and the Ascend device plugin, then verify template matching, multi-Pod sharing, capacity exhaustion, whole-card exclusion, and monitoring on an Ascend 310P3.
Install HAMi DRA 0.2.3 and the Ascend DRA driver on an Ascend 310P3 node, watch the webhook turn huawei.com/Ascend310P requests into ResourceClaims, and verify two-Pod NPU sharing, memory quota enforcement, and scheduler capacity accounting.
Simulate 8 A100 GPUs with HAMi scheduling features, no real GPU needed.