Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mindspore-ai/akg/cpu-optimization-armnpx skills add mindspore-ai/akg --skill cpu-optimization-armgit clone --depth 1 https://github.com/mindspore-ai/akgWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mindspore-ai/akg/cpu-optimization-arm)<a href="https://agentmods.dev/skills/mindspore-ai/akg/cpu-optimization-arm"><img src="https://agentmods.dev/badge/skills/mindspore-ai/akg/cpu-optimization-arm.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00030 | $0.04896 |
| Opus 5 | $0.00015 | $0.02448 |
| Sonnet 5 | $0.00006 | $0.00979 |
| Haiku 4.5 | $0.00003 | $0.00490 |
Grade A, and why
cpu-optimization-arm scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 503 lines — stays where its author put it; the contents beside it link to each section on GitHub.
ARM CPU 性能优化指南
1. ARM 架构特性与优化策略
1.1 架构标识
- 架构: aarch64 (ARM 64-bit, ARMv8-A)
- 主要厂商: ARM, Apple Silicon (M1/M2/M3), AWS Graviton, 华为鲲鹏
- SIMD 扩展: NEON (Advanced SIMD)
1.2 核心优化原则
- 利用 NEON 并行性: 使用 NEON 指令同时处理多个数据
- 消除数据依赖: 避免连续指令间的寄存器依赖
- 优化缓存使用: 按行优先访问,提高缓存命中率
- 减少分支预测失败: 循环展开,减少条件判断
2. NEON SIMD 向量化优化
2.1 基本概念
NEON (Advanced SIMD) 是 ARM 的 SIMD 指令集:
- 寄存器宽度: 128 位
- 并行处理能力:
- 4 个 float32(单精度浮点)
- 2 个 float64(双精度浮点)
- 16 个 int8, 8 个 int16, 4 个int32, 2 个 int64
2.2 编译器自动向量化
推荐方式: 让编译器自动向量化,通过编译选项启用:
# 在 load_inline 中添加 ARM 向量化选项
op_module = load_inline(
name="custom_op",
cpp_sources=cpp_source,
extra_cflags=[
"-O3", # 最高优化级别
"-mcpu=native", # 针对当前 ARM CPU 优化
"-ftree-vectorize", # 启用自动向量化
"-ffast-math", # 快速数学优化(可选)
],
verbose=True
)
注意: ARM 使用 -mcpu=native 而不是 -march=native。
2.3 循环优化示例
简单方式(未优化):
torch::Tensor elementwise_add(torch::Tensor a, torch::Tensor b) {
if (!a.is_contiguous()) a = a.contiguous();
if (!b.is_contiguous()) b = b.contiguous();
torch::Tensor output = torch::zeros_like(a);
auto a_ptr = a.data_ptr<float>();
auto b_ptr = b.data_ptr<float>();
auto out_ptr = output.data_ptr<float>();
int64_t numel = a.numel();
// 简单循环
for (int64_t i = 0; i < numel; ++i) {
out_ptr[i] = a_ptr[i] + b_ptr[i];
}
return output;
}
优化方式(循环展开,便于 NEON 向量化):
torch::Tensor elementwise_add_optimized(torch::Tensor a, torch::Tensor b) {
if (!a.is_contiguous()) a = a.contiguous();
if (!b.is_contiguous()) b = b.contiguous();
torch::Tensor output = torch::zeros_like(a);
auto a_ptr = a.data_ptr<float>();
auto b_ptr = b.data_ptr<float>();
auto out_ptr = output.data_ptr<float>();
int64_t numel = a.numel();
// 循环展开 4 倍(匹配 NEON 对 float32 的处理能力)
int64_t i = 0;
int64_t step = 4;
for (; i + step <= numel; i += step) {
out_ptr[i] = a_ptr[i] + b_ptr[i];
out_ptr[i + 1] = a_ptr[i + 1] + b_ptr[i + 1];
out_ptr[i + 2] = a_ptr[i + 2] + b_ptr[i + 2];
out_ptr[i + 3] = a_ptr[i + 3] + b_ptr[i + 3];
}
// 处理剩余元素
for (; i < numel; ++i) {
out_ptr[i] = a_ptr[i] + b_ptr[i];
}
return output;
}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 503 lines · 30 tokens per session scan A 8268e3e6a4ec
cpu-optimization-arm is a skill published in the GitHub repository mindspore-ai/akg (259 stars, last pushed 25d ago), licensed Apache-2.0. It adds 30 tokens to every session and 4,896 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
integrated-browser
Use this when working on the VS Code integrated browser ("browserView") to understand its architecture and mental model. Covers the embedded Chromium browser, its editor tab, navigation, overlay/layout, sessions, and agent browser tools under src/vs/platform/browserView and src/vs/workbench/contrib/browserView.
gke-compute-classes
Configures, optimizes, and troubleshoots GKE ComputeClasses. Use when configuring Spot VMs with on-demand fallback, targeting specific accelerators (GPUs/TPUs) or machine families, restricting ComputeClass access, or debugging pending pods related to node pool auto-creation. Do not use for cluster-level Node Auto…
jetson-diagnostic
Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.
doca-socket-relay
Use this skill when the operator is driving the DOCA Socket Relay to bridge a socket-oriented host application onto a BlueField DPU peer without rewriting it — picking the deployment shape (in-process, sidecar, or BlueField service container), configuring the host-side socket and the DPU-side forwarding endpoint…
offensive-z-wave
Z-Wave attack methodology — sniffing with Z-Force / EZ-Wave / RTL-SDR + ZniffMobile, S0 (legacy) network-key derivation flaw and key reuse, S2 (modern) ECDH commissioning analysis, replay/injection on unauthenticated nodes, default-key brute-force on test deployments, and home-automation hub pivots. Use when targeting…
hsb-flash
Flash the FPGA on an HSB board connected to an NVIDIA devkit. Supports HSB Lattice boards (FPGA versions 2407, 2412, 2507, 2510) and Leopard Imaging VB1940 "all-in-one" cameras (FPGA versions 2507, 2510). Uses release-specific YAML manifests and board-type-specific program commands. Lattice and VB1940 commands must…