Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ericrisco/rsc-harness --skill replicategit clone --depth 1 https://github.com/ericrisco/rsc-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ericrisco/rsc-harness/replicate)<a href="https://agentmods.dev/skills/ericrisco/rsc-harness/replicate"><img src="https://agentmods.dev/badge/skills/ericrisco/rsc-harness/replicate/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/ericrisco/rsc-harness/replicate"><img src="https://agentmods.dev/badge/skills/ericrisco/rsc-harness/replicate.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00073 | $0.02574 |
| Opus 5 | $0.00036 | $0.01287 |
| Sonnet 5 | $0.00015 | $0.00515 |
| Haiku 4.5 | $0.00007 | $0.00257 |
Grade A, and why
replicate scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 211 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Replicate platform operations
This skill is about how a model runs in production on Replicate — clients, async, deployments,
Cog packaging, webhooks, scaling, and spend. It is the platform-engineering counterpart to image
prompt craft. If the question is what prompt, aspect ratio, or model family produces a good image,
that is replicate-images, not this skill. Here the mental model is: a prediction is a job. You
either wait for it, poll it, or get pinged about it — and where it runs (shared cold pool vs a private
warm deployment) is a cost-and-latency dial you set deliberately.
Pinned facts (verified 2026-06-02): Python client replicate 1.x (latest 1.0.7), Python 3.8+;
JS client replicate on npm; auth via REPLICATE_API_TOKEN. A 2.0.0aN alpha exists on PyPI but
is NOT the default — pin replicate>=1,<2 so a fresh install never silently pulls it.
Decision: how should this model run?
Pick the row by latency tolerance and whether your process can block. Do not default to run() for
everything — a 10-minute job inside a web request will time out and burn a worker.
| Pattern | Latency | Blocks your process? | Cost shape | Use when |
|---|---|---|---|---|
replicate.run(...) |
seconds | Yes — waits to completion | per-prediction, shared pool | interactive/quick calls, scripts, CLIs |
predictions.create() + poll |
minutes | Yes, but you control the loop | per-prediction, shared pool | long job, a worker can babysit it |
predictions.create(webhook=...) |
minutes+ | No — fire and forget | per-prediction, shared pool | long job, the request must return now |
| Deployment (private endpoint) | low + steady | depends on call style above | warm floor + per-prediction | sustained traffic, need warm/private/autoscale cap |
Rule: if a human or HTTP request is waiting longer than a few seconds, do not block on run() —
switch to predictions + webhook. Why: synchronous timeouts kill the request but the GPU job keeps
running and billing.
Auth & install
What ships with it
6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 211 lines · 73 tokens per session scan A a0f8c4c44250
replicate is a skill published in the GitHub repository ericrisco/rsc-harness (82 stars, last pushed yesterday), licensed MIT. It adds 73 tokens to every session and 2,574 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
assessing-vector-and-embedding-weaknesses
Test vector stores for embedding inversion, cross-tenant leakage, and poisoning.
detecting-data-and-model-poisoning
Identify poisoned training data and backdoored models across the ML pipeline.
detecting-model-extraction-attacks
Detect model stealing, model inversion, and membership inference performed through inference-API abuse by monitoring query patterns, applying output perturbation, and red-teaming your own model's extractability.
testing-for-system-prompt-leakage
Extract and defend system prompts plus embedded secrets and routing logic.
defending-llms-with-guardrails
Deploy Llama Guard, NeMo Guardrails, and LLM Guard input/output scanners as runtime defenses.
red-teaming-llms-with-garak
Run NVIDIA garak probe suites against an LLM endpoint to test for jailbreaks, prompt injection, data leakage, and toxic generation, then interpret the hit-rate report for triage and reporting.