Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mindspore-ai/akg/cpu-optimization-x64npx skills add mindspore-ai/akg --skill cpu-optimization-x64git clone --depth 1 https://github.com/mindspore-ai/akgWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mindspore-ai/akg/cpu-optimization-x64)<a href="https://agentmods.dev/skills/mindspore-ai/akg/cpu-optimization-x64"><img src="https://agentmods.dev/badge/skills/mindspore-ai/akg/cpu-optimization-x64.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00034 | $0.03735 |
| Opus 5 | $0.00017 | $0.01868 |
| Sonnet 5 | $0.00007 | $0.00747 |
| Haiku 4.5 | $0.00003 | $0.00374 |
Grade A, and why
cpu-optimization-x64 scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 381 lines — stays where its author put it; the contents beside it link to each section on GitHub.
x64 CPU 性能优化指南
1. x64 架构特性与优化策略
1.1 架构标识
- 架构: x86_64 (也称为 x64, AMD64)
- 主要厂商: Intel, AMD
- SIMD 扩展: AVX, AVX2, AVX-512
1.2 核心优化原则
- 利用 SIMD 并行性: 使用 AVX/AVX2/AVX-512 指令同时处理多个数据
- 优化缓存使用: 按行优先访问,提高缓存命中率
- 减少分支预测失败: 循环展开,减少条件判断
- 内存对齐: 确保数据对齐到 32/64 字节边界
2. SIMD/AVX 向量化优化
2.1 基本概念
AVX (Advanced Vector Extensions) 是 x86-64 的 SIMD 指令集扩展:
- AVX: 256 位寄存器,可同时处理 8 个 float32 或 4 个 float64
- AVX2: 增强的 AVX,支持整数运算
- AVX-512: 512 位寄存器,可同时处理 16 个 float32 或 8 个 float64
2.2 编译器自动向量化
推荐方式: 让编译器自动向量化,通过编译选项启用:
# 在 load_inline 中添加向量化选项
op_module = load_inline(
name="custom_op",
cpp_sources=cpp_source,
extra_cflags=[
"-O3", # 最高优化级别
"-march=native", # 针对当前 CPU 架构优化
"-ftree-vectorize", # 启用自动向量化
],
verbose=True
)
2.3 循环优化示例
简单方式(未优化):
torch::Tensor elementwise_add(torch::Tensor a, torch::Tensor b) {
if (!a.is_contiguous()) a = a.contiguous();
if (!b.is_contiguous()) b = b.contiguous();
torch::Tensor output = torch::zeros_like(a);
auto a_ptr = a.data_ptr<float>();
auto b_ptr = b.data_ptr<float>();
auto out_ptr = output.data_ptr<float>();
int64_t numel = a.numel();
// 简单循环
for (int64_t i = 0; i < numel; ++i) {
out_ptr[i] = a_ptr[i] + b_ptr[i];
}
return output;
}
优化方式(循环展开,便于向量化):
torch::Tensor elementwise_add_optimized(torch::Tensor a, torch::Tensor b) {
if (!a.is_contiguous()) a = a.contiguous();
if (!b.is_contiguous()) b = b.contiguous();
torch::Tensor output = torch::zeros_like(a);
auto a_ptr = a.data_ptr<float>();
auto b_ptr = b.data_ptr<float>();
auto out_ptr = output.data_ptr<float>();
int64_t numel = a.numel();
// 循环展开 8 倍(匹配 AVX 寄存器宽度)
int64_t i = 0;
int64_t step = 8;
for (; i + step <= numel; i += step) {
out_ptr[i] = a_ptr[i] + b_ptr[i];
out_ptr[i + 1] = a_ptr[i + 1] + b_ptr[i + 1];
out_ptr[i + 2] = a_ptr[i + 2] + b_ptr[i + 2];
out_ptr[i + 3] = a_ptr[i + 3] + b_ptr[i + 3];
out_ptr[i + 4] = a_ptr[i + 4] + b_ptr[i + 4];
out_ptr[i + 5] = a_ptr[i + 5] + b_ptr[i + 5];
out_ptr[i + 6] = a_ptr[i + 6] + b_ptr[i + 6];
out_ptr[i + 7] = a_ptr[i + 7] + b_ptr[i + 7];
}
// 处理剩余元素
for (; i < numel; ++i) {
out_ptr[i] = a_ptr[i] + b_ptr[i];
}
return output;
}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 381 lines · 34 tokens per session scan A d0921abb5c78
cpu-optimization-x64 is a skill published in the GitHub repository mindspore-ai/akg (259 stars, last pushed 26d ago), licensed Apache-2.0. It adds 34 tokens to every session and 3,735 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
attribute-driven-design
Drive architecture from quality-attribute scenarios using the iterative ADD 3.0 method, producing a reviewable design workbook.
cpu-optimization-arm
ARM CPU 架构性能优化技巧、NEON SIMD 向量化、数值稳定性和调试策略.
cpu-optimization-x64
AVX (Advanced Vector Extensions) 是 x86-64 的 SIMD 指令集扩展:.
pypto-optimization
PyPTO 性能优化规则与调参顺序。适用于需要优化 tile/loop/归约性能、比较不同 tile 方案、解释同一算子不同 tile 性能差异(尤其 softmax/logsoftmax/reduction/norm/loss)的场景.
tilelang-cuda-optimization
TileLang CUDA 性能优化通用策略、最佳实践和调试技巧汇总。适用于需要提升 TileLang 内核性能、遇到编译/运行错误需要排查、或需要了解 TileLang 平台限制的内核代码生成和优化场景.
tilelang-cuda-patterns
TileLang CUDA 核心编程模式(逐元素、归约、矩阵乘法、GEMV)的标准实现范式和代码模板。适用于需要快速确定算子属于哪种编程模式、或需要了解 TileLang 各模式基本代码结构的内核代码生成场景.