Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/Amal-David/mlx-porting-skillnpx agentmods add skills/amal-david/mlx-porting-skill/mlx-model-portingWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/amal-david/mlx-porting-skill/mlx-model-porting)<a href="https://agentmods.dev/skills/amal-david/mlx-porting-skill/mlx-model-porting"><img src="https://agentmods.dev/badge/skills/amal-david/mlx-porting-skill/mlx-model-porting.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00244 | $0.03124 |
| Opus 5 | $0.00122 | $0.01562 |
| Sonnet 5 | $0.00049 | $0.00625 |
| Haiku 4.5 | $0.00024 | $0.00312 |
Grade A, and why
mlx-model-porting scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 147 lines — stays where its author put it; the contents beside it link to each section on GitHub.
MLX model porting and optimization
Mission
Produce or inspect a correct, reproducible, architecture-aware MLX implementation. Correctness before speed. Every speed or memory claim must name hardware, software versions, workload, baseline, and quality gate.
Six families have scaffolds (MoE/SSM synthetic, others runbook-guided); four have worked packets under examples/: Qwen2.5, BGE, t5-small, HuBERT. Exact output is the built-in metric.
When to use this skill
Port, convert, run, inspect, quantize, package, or publish a PyTorch/Hugging Face model or MLX project on Apple Silicon, or fix any parity, NaN/Inf, shape, tokenizer, preprocessing, output, performance, memory, cache, serving, benchmark, or provenance issue in a port. Not for CUDA/non-Apple targets, ML theory without an MLX target, or training from scratch.
Trigger map
| Signal | Load |
|---|---|
Port/convert/run request, config.json, safetensors index, model directory, or Hub id. |
intake, then Workflow 1-6; worked chain for dense decoders, routed runbooks otherwise. |
| User points at an existing local MLX project, running MLX app, or completed MLX port. | inspector mode plus inspect_mlx_project.py |
| User asks "what can I do with this model?", asks for capability fit, or wants model-specific advice. | model advisor playbook |
| A known architecture family needs the right runbook. | model support map, then the architecture table in Workflow step 4 below |
| Dense decoder, BERT encoder, T5 encoder-decoder, HuBERT/Wav2Vec2 acoustic encoder, sparse MoE, or selective SSM. | Use scaffold_port.py; select capture/parity mode dense-decoder (default), encoder, encoder-decoder, asr, or ssm. |
| NaN, Inf, cosine-similarity drift, parity failure, or garbage output appears versus the source. | failure atlas |
| Weight conversion, key mapping, tensor rename, transpose, reshape, split, merge, or shape transform is in scope. | core porting method |
| The user says "make it faster" but no profile, workload, or baseline exists yet. | benchmarking |
| Speedup plan, how techniques combine, or expected compound gains. | compound stacks |
| KV cache, long context, recurrent state, attention memory, or prefill/decode memory is the bottleneck. | attention and KV cache |
| Quantization or "4-bit". | guide, quality gate |
| Structured local optimization sweep. | loop |
| Decoding, serving, speculative decoding, batching, streaming, or API runtime behavior is requested. | decoding and serving |
Compile behavior, mx.compile, custom kernel, graph capture, Metal, or operation fusion comes up. |
compile and kernels |
| Publish, release, checkpoint conversion, model card, provenance, or license packaging is requested. | packaging and publication |
| The user asks for "50-100 optimization ideas", a deep model-specific hunt, or research-backed candidates. | hypothesis-led learning |
| Vision-language, multimodal, or image+text (VLM) input appears. | multimodal/omni runbook |
| Diffusion, flow-matching, or image/video generation appears. | diffusion/flow runbook |
| Text-to-speech, vocoder, or audio generation appears. | flow-TTS, autoregressive audio |
| ASR, transcription, or streaming speech appears. | ASR, streaming speech |
| Sparse mixture-of-experts or top-k expert routing appears. | MoE runbook |
| Selective state-space, Mamba, or linear-attention hybrid appears. | SSM/hybrid runbook |
| Graph, GNN, message passing, node/edge features, or sparse graph workload appears. | graph message passing runbook |
| Classic CV detection, segmentation, keypoints, depth, OCR, or non-generative vision appears. | non-generative CV runbook |
| Time-series, forecasting, tabular sequence, anomaly detection, or temporal model appears. | time-series forecasting runbook |
What ships with it
60 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- assets/architectures.yaml 27 KB
- assets/BENCHMARK_REPORT.md 5.5 KB
- assets/benchmarks/attestations/dependencies/18e18afcaccafade98daf13a54092927904649e1dd4eba8299ab717d5d94ff45.bin 659 B
- assets/benchmarks/attestations/dependencies/29ef203c13d9bebb6b8cad6aadb44d1ad495e2bbc19184ca5415b6a505eb36f6.bin 618 B
- assets/benchmarks/attestations/dependencies/311bb75620c546b7ad092f6864e45452a3d5a4fc8a698b3e691769484e2f0469.bin 1010 KB
- assets/benchmarks/attestations/dependencies/3a2364d8096a54c302c820e24beabb976d2b8b9502ca865337711307ed8a4d61.bin 1.8 KB
- assets/benchmarks/attestations/dependencies/3bbfbecc997379696313844c6056ffaf7c88d59300dcb92adb47709eb2410e9e.bin 1.4 KB
- assets/benchmarks/attestations/dependencies/4b0b2f3a82a33037efafd221d83c4add53ef7915b7fedaf04ed546d8c3f7ba75.bin 16 KB
- assets/benchmarks/attestations/dependencies/623e1dd10d518a2461ef30229cf63ea36cb9b61b93ebb7f5371ec5b05f479aae.bin 9.7 KB
- assets/benchmarks/attestations/dependencies/700eff903c2312233b464bb0f49f0d9d22241d93801a8c137b2e0119c150e822.bin 12 KB
- assets/benchmarks/attestations/dependencies/7bdd4fd41047a617ec84f7eb214cde3aaa3c3e4c932550468bae1204d1df6550.bin 303 B
- assets/benchmarks/attestations/dependencies/8f61a6438d46b55ce5579edf247f9988bd98bd801a37e4ccdaf295f1c3aa5bc0.bin 14 KB
- assets/benchmarks/attestations/dependencies/9050a138850c3d4e1f10057ea46425ae41ee80900a98524468742f3a44db91e5.bin 9.4 KB
- assets/benchmarks/attestations/dependencies/90e8e529c1cc2e48d81caf697247b5f46912ab83446014087c7116075542b235.bin 4.4 KB
- assets/benchmarks/attestations/dependencies/92cfbc9ba53d4b2d7681a9096881d8cce4c7f5aff06ce054efbd6d454cf8c14f.bin 12 KB
- assets/benchmarks/attestations/dependencies/933c0376d99e23e3348569d9029809caeb66e96e14f89d80c120cb280e871f14.bin 8.5 KB
- assets/benchmarks/attestations/dependencies/99bbc9e8f50ab6aa1942d8afde9ae4f5f36fd3af31c916f769b66108299676af.bin 3.5 KB
- assets/benchmarks/attestations/dependencies/a129a2c913db5bb7004a76e189c1121d8957cac174934a4b4b671200199dd659.bin 8.1 KB
- assets/benchmarks/attestations/dependencies/b325de6ce79d3cb39a0541aecd42a60296ad37f8dde5720c9a3775811151ce1b.bin 20 KB
- assets/benchmarks/attestations/dependencies/ba01c5485dfe46d9819ecf016f844ef2ca2e721418e805fe0e4ddec852a5a8c3.bin 6.8 KB
- assets/benchmarks/attestations/dependencies/c02ca21d6abf70d741db7a322e8cdde793b97604a926d0014223b22b011b5adf.bin 5.4 KB
- assets/benchmarks/attestations/dependencies/c055c2e5658faa050c0814e2eaac1ad5d31212a199e659f729abe6e2c6c86fa9.bin 5.7 KB
- assets/benchmarks/attestations/dependencies/d1a2088211cfe5b4cb9723da93243efdc0a2b94e5cf74f7d1af31cc6a42aaf1e.bin 4.5 KB
- assets/benchmarks/attestations/dependencies/d1e740f70dc8a2dcce8f24f39c1c43db1824724312314919cb40109a81ec6fe8.bin 14 KB
- assets/benchmarks/attestations/dependencies/e34b902aa24fe1971420b7ef33935f322c7cf24d6f16911d46171f770ee60bd0.bin 7.6 KB
- assets/benchmarks/attestations/dependencies/e538cddf7479f1aad592a0058e9a1eb29120fbbb4c6e6507b54ed1fbad02fb45.bin 151 B
- assets/benchmarks/attestations/dependencies/e62a3b11aff60a72009bd7690502043cdfdcd493ff71a27f557a7f299ca74ec8.bin 22 KB
- assets/benchmarks/attestations/dependencies/f16d40bace757e068a3afac593b5df900743ea69f71d42bd34efd352567981dd.bin 9.6 KB
- assets/benchmarks/attestations/dependencies/f748ea4f10995bf30ed6ba76ed3539f23c18cdf541031e1c0968dd60dc98f723.bin 321 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-001/challenge.json 1.2 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-001/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-001/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-001/quality-contract.json 511 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-002/challenge.json 1.2 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-002/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-002/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-002/quality-contract.json 511 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-003/challenge.json 1.2 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-003/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-003/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-003/quality-contract.json 511 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-004/challenge.json 1.2 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-004/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-004/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-004/quality-contract.json 511 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-005/challenge.json 1.2 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-005/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-005/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/measure-005/quality-contract.json 511 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/warmup-001/challenge.json 1.2 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/warmup-001/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/warmup-001/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-bf16/runs/warmup-001/quality-contract.json 511 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-001/challenge.json 1.1 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-001/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-001/output.json 54 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-001/quality-contract.json 510 B
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-002/challenge.json 1.1 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-002/evidence.json 13 KB
- assets/benchmarks/attestations/qwen2.5-0.5b-port-f32/runs/measure-002/output.json 54 B
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 147 lines · 244 tokens per session scan A eeeca7c2efc3
mlx-model-porting is a skill published in the GitHub repository Amal-David/mlx-porting-skill (5 stars, last pushed today), licensed Apache-2.0. It adds 244 tokens to every session and 3,124 once invoked, about $0.0012 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
add-new-model
Use this skill when the user wants to add or port a new model architecture to MLX-VLM — mapping a Hugging Face modeltype to a new file under mlxvlm/models, writing the ModelConfig, matching layer/weight names, reusing a similar existing model, adding a test class, and validating the port. Covers vision-language…
convert-quantize
Use this skill when the user wants to convert a Hugging Face model to MLX or quantize/dequantize one with mlxvlm.convert, including bits and group size, quant modes (affine, mxfp4, nvfp4, mxfp8), RTN vs AWQ, mixed-bit recipes, dtype casts, calibration (text or multimodal), local vs Hub paths, revisions, uploading to…
cli-inference
Use this skill when the user wants to run or debug MLX-VLM inference from the command line, including uv run mlxvlm.generate, image/audio/video inputs, local model paths, Hugging Face model IDs, deterministic repro commands, and CLI errors around processors, prompts, model loading, or missing weights.
server-inference
Use this skill when the user wants to run or debug MLX-VLM server inference, including uv run mlxvlm.server, /v1/models, /v1/chat/completions, /v1/responses, streaming, OpenAI-compatible clients, health checks, metrics, model unload/reload, adapters, trust-remote-code, and server request/response failures.
hf-cache-models
Use this skill when the user wants to list, inspect, or report MLX-VLM model candidates available in the local Hugging Face cache directory, including the server's opt-in hf-cache discovery mode, cache-dir overrides, JSON output, or issue-ready cached model lists.
exporting-to-fhir
Convert OpenMed NER output (entities from openmed.analyzetext) into FHIR R4 resources — Condition, MedicationStatement, Observation — using OpenMed's built-in FHIR R4 export helpers in openmed.clinical.exporters. Covers the verified CodeableConcept builder (coding, codeableconcept, systemuri), deterministic fullUrl…