bygoogle· 116 skills
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for generic batch jobs or HPC task queues (use gke-batch-hpc instead).
$npx -y skills add google/skills --skill gke-inferenceInstalls into the current project.
Run `npx skills use "https://github.com/google/skills" --skill "google/skills/gke-inference"` and follow the generated skill instructions now. Read its complete output, redirecting it to a temporary file first if necessary. Resolve relative paths from the supporting-files directory it provides.
Use the skills in "https://github.com/google/skills" that are relevant to the current task. Run `npx skills add "https://github.com/google/skills"` and select the relevant skills, then follow their instructions.