Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Tasachii/Tasachii-Tools --skill experiment-loggit clone --depth 1 https://github.com/Tasachii/Tasachii-ToolsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/tasachii/tasachii-tools/experiment-log)<a href="https://agentmods.dev/skills/tasachii/tasachii-tools/experiment-log"><img src="https://agentmods.dev/badge/skills/tasachii/tasachii-tools/experiment-log/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/tasachii/tasachii-tools/experiment-log"><img src="https://agentmods.dev/badge/skills/tasachii/tasachii-tools/experiment-log.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00113 | $0.00783 |
| Opus 5 | $0.00056 | $0.00392 |
| Sonnet 5 | $0.00023 | $0.00157 |
| Haiku 4.5 | $0.00011 | $0.00078 |
Grade A, and why
experiment-log scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
experiment-log — การทดลองที่ไม่ได้จด เท่ากับไม่ได้ทำ
Ledger อยู่ที่ EXPERIMENTS.md ที่ root ของ repo โปรเจกต์นั้น (อยู่กับโค้ด version ไปด้วยกัน — ไม่ใช่ ~/Documents) ถ้ายังไม่มีให้สร้างพร้อม entry แรกจากสิ่งที่รู้
Entry template
## EXP-007 — 2026-07-08
**สมมติฐาน:** เพิ่ม feature X น่าจะช่วยเพราะ <เหตุผลที่คิดก่อนรัน>
**เปลี่ยนจาก:** EXP-005 (best ปัจจุบัน) — เปลี่ยนอย่างเดียว: <diff>
**Config:** <hyperparams สำคัญ> · commit `abc123` · seed 42
**ผล:** <metric หลัก> = 0.847 (best เดิม 0.851, baseline 0.71)
**คำตัดสิน:** KILL — แพง feature engineering แต่ metric ลด · เรียนรู้: <อะไร>
กติกาที่บังคับตอนจด
- สมมติฐานมาก่อนรัน — ถ้า user จะรันโดยไม่มีเหตุผลว่าทำไมน่าจะดีขึ้น ให้ถามหนึ่งครั้ง ("คาดว่าจะดีขึ้นเพราะอะไร") ถ้าตอบว่าลองมั่วๆ ก็จดว่า
exploratoryตรงๆ — ห้ามแต่งสมมติฐานย้อนหลัง - เปลี่ยนทีละอย่าง — ถ้า run ใหม่เปลี่ยนหลายอย่างพร้อมกัน (features + lr + โมเดล) เตือนว่าผลจะบอกไม่ได้ว่าอะไรช่วย ถ้ายืนยันก็จดว่าเปลี่ยนอะไรบ้างครบทุกตัว
- อ้าง run ก่อนหน้า — ทุก entry ระบุว่า diff จาก run ไหน ไม่ใช่ลอยๆ
- metric จาก harness เดิมเท่านั้น — เลขที่คำนวณคนละวิธีห้ามลง ledger ปนกัน (ถ้ายังไม่มี harness ชี้ไปสกิล ml-baseline ก่อน)
- คำตัดสินชัด — KEEP (best ใหม่) / KILL (ทางตัน + เรียนรู้อะไร) / INVESTIGATE (ผลแปลก ต้องขุด) — entry ที่ไม่มีคำตัดสินคือภาระคนอ่าน
พฤติกรรมอัตโนมัติ
- ก่อนรันของใหม่: scan ledger ว่าไอเดียนี้เคยลอง (และตายไป) แล้วหรือยัง — เจอให้ทัก พร้อมเหตุผลตอนนั้น
- เปิด session ใหม่: อ่าน ledger แล้ว brief ได้ใน 3 บรรทัด — best ปัจจุบันคืออะไร, ล่าสุดลองอะไร, ค้างอะไรอยู่
- ถูกถาม "สรุปหน่อย": ตาราง top-5 runs + ทิศทางที่ยังไม่ได้ลองที่มีเหตุผลรองรับ
- จด failure ให้ละเอียดเท่า success — KILL ที่จดดีคือสิ่งที่กันทีม (และตัวเองเดือนหน้า) ไม่ให้เสียเวลาซ้ำ
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 35 lines · 113 tokens per session scan A 78d427dc29e1
experiment-log is a skill published in the GitHub repository Tasachii/Tasachii-Tools (2 stars, last pushed 1mo ago), licensed MIT. It adds 113 tokens to every session and 783 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
arboreto
Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for…
pyhealth
Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality, readmission, length-of-stay, drug recommendation, sleep staging, ICD coding, EEG events), instantiating models (Transformer…
torchdrug
Build and troubleshoot TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and knowledge graph reasoning. Use when code imports torchdrug or needs its datasets, models, tasks, or Engine.
deepspot-m
Generate transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Use when you need spatial gene expression in log1p-CPM for 224x224 tiles at about 20x, want to query protein-coding genes by symbol instead of a fixed panel, or want to run prediction across a whole slide after tiling with…
nemo-mbridge-perf-expert-parallel-overlap
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.
pick-a-pii-model
Select an on-device OpenMed PII model from the committed registry by language, runtime format, and size budget, then require recall validation before deployment. Use when an agent must choose a local PII detector for CPU, Apple Silicon, or a mobile export without relying on live model discovery.