intel/gpu-ai-skills

Agent Skills for running, benchmarking, and profiling Hugging Face models on Intel GPUs — PyTorch, vLLM-XPU, and SGLang-XPU

16Stars on the repository
26Mods indexed here, across every type
5d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

intel/gpu-ai-skills

Skill Claude CodeCodex

Create a CUDA-to-XPU migration assessment for an existing AI repo. Identify CUDA-specific assumptions, route to the right XPU skills, produce a migration report. Use when the user has a CUDA repo, notebook, Dockerfile, launch script, HF / vLLM / SGLang workload, or Triton kernel and asks to migrate it to Intel Arc /…

16 5d ago A 163 tokens original Apache-2.0

llamacpp-xpu-run

02

intel/gpu-ai-skills

Skill Claude CodeCodex

Run a GGUF model on an Intel GPU using llama.cpp's SYCL backend (Level Zero) with the official intel.Dockerfile. Covers building the Docker image from source at a pinned tag, launching llama-server with an OpenAI-compatible API, device selection, multi-GPU layer splitting, all recommended runtime env vars…

16 5d ago A 163 tokens original Apache-2.0

model-can-it-fit

03

intel/gpu-ai-skills

Skill Claude CodeCodex

Estimate whether a Hugging Face decoder-only LLM, MoE, or VLM fits in Intel GPU VRAM for a quantization, context length, concurrency, runtime, and tensor-parallel setting. Use for memory-fit or max-model-len planning before launch. Reports weights, KV cache, activations, framework overhead, and first mitigation. Not…

16 5d ago A 106 tokens original Apache-2.0

intel/gpu-ai-skills

Skill Claude CodeCodex

Recommend how to configure vLLM-XPU for a Hugging Face decoder-only LLM on Intel Arc B-series GPUs: choose quantization, KV dtype, DP/TP layout, max concurrency, and max context using roofline math against published hardware specs. Use when the user asks "How should I configure vLLM?", requests the best vLLM…

16 5d ago A 121 tokens original Apache-2.0

sglang-xpu-bench

05

intel/gpu-ai-skills

Skill Claude CodeCodex

Benchmark a running SGLang-XPU server on an Intel GPU using sglang.benchserving. Measures TTFT, TPOT, ITL, end-to-end latency, and throughput against the OpenAI-compatible endpoint. Use after sglang-xpu-run. Not for vLLM servers (use vllm-xpu-bench) or no-server PyTorch (use torch-xpu-bench).

16 5d ago C 93 tokens original Apache-2.0

sglang-xpu-run

06

intel/gpu-ai-skills

Skill Claude CodeCodex

Serve a Hugging Face safetensors model on an Intel GPU using SGLang's XPU backend with the OpenAI-compatible API. Covers pulling the pre-built intel/sglang-dev:latest image, fixing the render-group and UMD/kernel compatibility issues that affect non-root sglang images, the SYCLUR / Level Zero env vars needed on…

16 5d ago D 170 tokens original Apache-2.0

torch-xpu-bench

07

intel/gpu-ai-skills

Skill Claude CodeCodex

Benchmark a Hugging Face model on an Intel GPU through pure PyTorch + Transformers, single-process, no HTTP server. Measures generate() throughput in tokens/sec, time-to-first-token, decode-step latency, and peak XPU memory. Also covers diffusion and encoder-only models via references/non-llm-snippets.md. Use after…

16 5d ago A 96 tokens original Apache-2.0

torch-xpu-profile

08

intel/gpu-ai-skills

Skill Claude CodeCodex

Profile a Hugging Face model on Intel GPU at the PyTorch level with torch.profiler and Kineto. Captures CPU + XPU timeline, exports Chrome trace, identifies hottest kernels and async-overlap gaps. Use when the user asks why a model is slow, which op is the bottleneck, or where the GPU is idle. Not for profiling inside…

16 5d ago A 118 tokens original Apache-2.0

torch-xpu-run

09

intel/gpu-ai-skills

Skill Claude CodeCodex

Run an arbitrary Hugging Face safetensors model on an Intel GPU using upstream PyTorch (>= 2.8) with the built-in torch.xpu device. Covers loading from the Hub, picking the right dtype, autocast, multi-GPU with accelerate's devicemap, and the CUDA -> XPU code translation a user has to do once. Use for the Transformers…

16 5d ago A 143 tokens original Apache-2.0

vllm-xpu-bench

10

intel/gpu-ai-skills

Skill Claude CodeCodex

Benchmark a running vLLM-XPU OpenAI-compatible server on an Intel GPU using vllm bench. Measures TTFT (time-to-first-token), TPOT (time-per-output-token), ITL (inter-token latency), end-to-end latency, and throughput under concurrency. Covers online (vllm bench serve) and offline (vllm bench throughput) modes…

16 5d ago C 124 tokens original Apache-2.0

vllm-xpu-profile

11

intel/gpu-ai-skills

Skill Claude CodeCodex

Profile a running vLLM-XPU server with torch.profiler around a window of real requests, either via /startprofile and /stopprofile HTTP endpoints or via vllm bench --profile for offline runs. Use to find the dominant op under real concurrent traffic. Not for pure PyTorch (use torch-xpu-profile), SYCL kernel-level…

16 5d ago A 105 tokens original Apache-2.0

vllm-xpu-run

12

intel/gpu-ai-skills

Skill Claude CodeCodex

Serve a Hugging Face safetensors model on an Intel GPU with upstream vLLM-XPU's OpenAI-compatible API, or check whether a model or architecture is currently documented on XPU. Covers live support lookup, image choice, container launch, known serve-flag requirements, model-impl fallback, and attention/quant…

16 5d ago A 129 tokens original Apache-2.0

xpu-container-run

13

intel/gpu-ai-skills

Skill Claude CodeCodex

Launch a Docker container with Intel GPU access on Linux. Encodes the correct combination of --device /dev/dri, render-group access, --ipc=host, ZEAFFINITYMASK pinning, Hugging Face cache mount, and --entrypoint /bin/bash for interactive use. Use when running any Intel-XPU container (vLLM-XPU, sglang-xpu, torch-XPU…

16 5d ago A 146 tokens original Apache-2.0

xpu-deploy-plan

14

intel/gpu-ai-skills

Skill Claude CodeCodex

Plan an end-to-end Intel XPU model deployment by chaining existing skills. Calls xpu-runtime-preflight (readiness), model-can-it-fit (sizing), model-config-recommend (flags), and the selected runtime skill (vllm-xpu-run / sglang-xpu-run / torch-xpu-run), then writes a single PLAN.md with one exact launch command…

16 5d ago A 149 tokens original Apache-2.0

xpu-discover

15

intel/gpu-ai-skills

Skill Claude CodeCodex

Inventory Intel GPUs (Arc, Arc Pro, Data Center GPU Max) on a Linux host. Detect devices, check driver health, list processes using each XPU, run a quick diagnostic, and read live utilisation.

16 5d ago A 48 tokens original Apache-2.0

intel/gpu-ai-skills

Skill Claude CodeCodex

Before loading a Hugging Face model on Intel XPU, detect its actual type (text generation, text encoder, seq2seq, vision classification, vision-language, audio encoder, audio seq2seq, multimodal VL, diffusion, time-series, reward model, masked LM) so the agent picks the right AutoModel class and input kwargs. Prevents…

16 5d ago A 141 tokens original Apache-2.0

xpu-port

17

intel/gpu-ai-skills

Skill Claude CodeCodex

Execute a single-target CUDA-to-XPU port of a PyTorch repo with libcst-based scan, mechanical rewrite, and CPU FP64 vs target-dtype correctness verify on one forward pass. Use when the request says "port" — "port my repo to XPU", "port my repo at to XPU", "rewrite the CUDA calls to XPU", "apply the mechanical…

16 5d ago A 193 tokens original Apache-2.0

intel/gpu-ai-skills

Skill Claude CodeCodex

Profile Intel-XPU workloads at the SYCL / Level Zero kernel level via Intel pti-gpu's unitrace. Captures per-API-call and per-kernel timing, memory transfers, oneCCL / MPI events, and hardware counters PyTorch-level profilers cannot see. Use when a hot op is already known at the torch.profiler layer and the user needs…

16 5d ago A 127 tokens original Apache-2.0

intel/gpu-ai-skills

Skill Claude CodeCodex

Run a read-only go/no-go preflight before any Intel GPU/XPU skillpack work. Checks driver health, /dev/dri permissions, render/video groups, Docker, /dev/shm, disk, proxy, and optional container-level XPU visibility. Use when the user asks whether a machine is ready for XPU model work or needs a reusable lab readiness…

16 5d ago A 98 tokens original Apache-2.0

xpu-system-setup

20

intel/gpu-ai-skills

Skill Claude CodeCodex

First-time setup for Intel XPU/GPU hosts. Detects what's missing and installs xpu-smi, configures user groups (render), sets up Intel GPU PPA repository, installs Level Zero runtime, installs Docker, and runs a post-setup verification gate. Prompts before each installation by default (use --auto for unattended). Also…

16 5d ago B 160 tokens original Apache-2.0

template-skill

21

intel/gpu-ai-skills

Skill Claude CodeCodex

Replace this with a one-paragraph description of what the skill does and when an agent should use it. Include keywords likely to appear in user prompts.

16 5d ago A 34 tokens original Apache-2.0