Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add mindspore-ai/akg --skill triton-ascend-case-elemwise-castgit clone --depth 1 https://github.com/mindspore-ai/akgWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mindspore-ai/akg/triton-ascend-case-elemwise-cast)<a href="https://agentmods.dev/skills/mindspore-ai/akg/triton-ascend-case-elemwise-cast"><img src="https://agentmods.dev/badge/skills/mindspore-ai/akg/triton-ascend-case-elemwise-cast/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/mindspore-ai/akg/triton-ascend-case-elemwise-cast"><img src="https://agentmods.dev/badge/skills/mindspore-ai/akg/triton-ascend-case-elemwise-cast.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00068 | $0.00675 |
| Opus 5 | $0.00034 | $0.00338 |
| Sonnet 5 | $0.00014 | $0.00135 |
| Haiku 4.5 | $0.00007 | $0.00068 |
Grade A, and why
triton-ascend-case-elemwise-cast scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Int8 到 FP16 类型转换优化案例
任务特征
- 操作类型:Elementwise,类型转换操作
- 数据尺寸:(128, 1024, 1024),shape较大
- 数据类型:输入int8,输出fp16
- 任务特点:可以按照轴的顺序(可flatten为一根轴),外层并行,内层向量化,若UB存不下,可考虑多次切分
优化:二次切分 + 用满UB
# Triton 内核实现:将BLOCK_SIZE分块,每次搬运TILE_SIZE大小的数据
configs = [
triton.Config({"BLOCK_SIZE": 65536, "TILE_SIZE": 65536}), # 核数2048, 性能最优!用满UB且无二次切分
triton.Config({"BLOCK_SIZE": 65536, "TILE_SIZE": 32768}), # 核数2048,但UB未用满
triton.Config({"BLOCK_SIZE": 2097152, "TILE_SIZE": 65536}), # 核数64, 并行度低
triton.Config({"BLOCK_SIZE": 4194304, "TILE_SIZE": 65536}), # 核数32, 并行度更低
]
# 内核操作:
block_start = pid * BLOCK_SIZE
for i in range(0, BLOCK_SIZE, TILE_SIZE):
offsets = block_start + tl.arange(0, TILE_SIZE)
mask = offsets < n_elements
input_data = tl.load(input_ptr + offsets, mask=mask)
output_data = tl.cast(input_data, tl.float16)
tl.store(output_ptr + offsets, output_data, mask=mask)
优化内容
- triton 内核部分使用for循环,尝试进行二次切分,每次搬运TILE_SIZE大小的数据,提高UB的利用率
- 在一定范围内提高核数,并尝试用满UB
- 核内没有二次切分时性能最优(BLOCK_SIZE = TILE_SIZE = 65536)
总结
- 当数据的shape较大时,为了获得更佳的性能,切分值设置尽量能被shape的大小整除
- 对于单纯的Elementwise操作,将多根轴的元素展开为一根轴,然后在这根轴上进行切分
- 将block分配给每个线程块,若UB存不下,可考虑多次切分(二次切分)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 50 lines · 68 tokens per session scan A 48b16dc72d35
triton-ascend-case-elemwise-cast is a skill published in the GitHub repository mindspore-ai/akg (259 stars, last pushed 1mo ago), licensed Apache-2.0. It adds 68 tokens to every session and 675 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
triton-ascend-case-matmul-swizzle2d
A case study for optimizing matrix multiplication in Triton on Ascend hardware using 2D block swizzling and a fixed number of cores.
triton-ascend-case-reduction-weighted-swiglu
A performance-optimization method for a specialized 3D tensor operation using Weighted SwiGLU during backward computation. It reshapes data and divides work to improve parallel processing while keeping within on-chip memory limits.
triton-ascend-case-elemwise-cast
A performance-tuning skill for converting large arrays from one data type to another, such as int8 to fp16. It uses two levels of splitting to improve use of shared on-chip memory on Ascend hardware.
triton-ascend-case-elemwise-zeros
An optimization for creating small tensors with zeros, arange, or full in Triton Ascend. Triton Ascend is a programming environment for running tensor operations on Ascend hardware.
gke-compute-classes
Configures, optimizes, and troubleshoots GKE ComputeClasses. Use when configuring Spot VMs with on-demand fallback, targeting specific accelerators (GPUs/TPUs) or machine families, restricting ComputeClass access, or debugging pending pods related to node pool auto-creation. Do not use for cluster-level Node Auto…
jetson-diagnostic
Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.