$npx -y skills add microsoft/azure-skills --skill airunway-aks-setupSet up AI Runway on AKS — from bare cluster to running model. Covers cluster verification, controller install, GPU assessment, provider setup, and first deployment. WHEN: \"setup AI Runway\", \"onboard AKS cluster\", \"install AI Runway\", \"airunway setup\", \"deploy model to AK
| 1 | # AI Runway AKS Setup |
| 2 | |
| 3 | This skill walks users from a bare Kubernetes cluster to a running AI model deployment. Follow each step in sequence unless the user provides `skip-to-step N` to resume from a specific phase. |
| 4 | |
| 5 | > **Cost awareness:** GPU node pools incur significant compute charges (A100-80GB can cost $3–5+/hr). Confirm the user understands cost implications before provisioning GPU resources. |
| 6 | |
| 7 | ## Prerequisites |
| 8 | |
| 9 | This skill assumes an AKS cluster already exists. If the user does not have a cluster, hand off to the `azure-kubernetes` skill first to provision one (with a GPU node pool unless CPU-only inference is acceptable), then return here. |
| 10 | |
| 11 | ## Quick Reference |
| 12 | |
| 13 | | Property | Value | |
| 14 | |----------|-------| |
| 15 | | Best for | End-to-end AI Runway onboarding on AKS | |
| 16 | | CLI tools | `kubectl`, `make`, `curl` | |
| 17 | | MCP tools | None | |
| 18 | | Related skills | `azure-kubernetes` (cluster setup), `azure-diagnostics` (troubleshooting) | |
| 19 | |
| 20 | ## When to Use This Skill |
| 21 | |
| 22 | Use this skill when the user wants to: |
| 23 | - Set up AI Runway on an existing AKS cluster from scratch |
| 24 | - Install the AI Runway controller and CRDs |
| 25 | - Assess GPU hardware compatibility for model deployment |
| 26 | - Choose and install an inference provider (KAITO, Dynamo, KubeRay) |
| 27 | - Deploy their first AI model to AKS via AI Runway |
| 28 | - Resume a partially-complete AI Runway setup from a specific step |
| 29 | |
| 30 | ## MCP Tools |
| 31 | |
| 32 | This skill uses no MCP tools. All cluster operations are performed directly via `kubectl` and `make`. |
| 33 | |
| 34 | ## Rules |
| 35 | |
| 36 | 1. Execute steps in sequence — load the reference for each step as you reach it |
| 37 | 2. Report cluster state at each step: ✓ healthy, ✗ missing/failed |
| 38 | 3. Ask for user confirmation before any install or deployment action |
| 39 | 4. If a step is already complete, report status and skip to the next step |
| 40 | 5. If the user provides `skip-to-step N`, start at step N; assume prior steps are complete |
| 41 | |
| 42 | ## Steps |
| 43 | |
| 44 | | # | Step | Reference | |
| 45 | |---|------|-----------| |
| 46 | | 1 | **Cluster Verification** — context check, node inventory, GPU detection | [step-1-verify.md](references/steps/step-1-verify.md) | |
| 47 | | 2 | **Controller Installation** — CRD + controller deployment | [step-2-controller.md](references/steps/step-2-controller.md) | |
| 48 | | 3 | **GPU Assessment** — detect GPU models, flag dtype/attention constraints | [step-3-gpu.md](references/steps/step-3-gpu.md) | |
| 49 | | 4 | **Provider Setup** — recommend and install inference provider | [step-4-provider.md](references/steps/step-4-provider.md) | |
| 50 | | 5 | **First Deployment** — pick a model, deploy, verify Ready | [step-5-deploy.md](references/steps/step-5-deploy.md) | |
| 51 | | 6 | **Summary** — recap, smoke test, next steps | [step-6-summary.md](references/steps/step-6-summary.md) | |
| 52 | |
| 53 | ## Error Handling |
| 54 | |
| 55 | | Error / Symptom | Likely Cause | Remediation | |
| 56 | |-----------------|--------------|-------------| |
| 57 | | No kubeconfig context | Not connected to a cluster | Run `az aks get-credentials` or equivalent | |
| 58 | | Controller in CrashLoopBackOff | Config or RBAC issue | `kubectl logs -n airunway-system -l control-plane=controller-manager --previous` | |
| 59 | | Provider not ready | Image pull or RBAC issue | `kubectl logs <pod-name> -n <namespace>` for the provider pod | |
| 60 | | ModelDeployment stuck in Pending | GPU scheduling failure or provider not ready | `kubectl describe modeldeployment <name> -n <namespace>` events | |
| 61 | | `bfloat16` errors at inference | T4 or V100 lacks bfloat16 support | Add `--dtype float16` to serving args | |
| 62 | |
| 63 | For full error handling and rollback procedures, see [troubleshooting.md](references/troubleshooting.md). |