Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Jaganpro/sf-skills --skill sf-ai-agentforce-testinggit clone --depth 1 https://github.com/Jaganpro/sf-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jaganpro/sf-skills/sf-ai-agentforce-testing)<a href="https://agentmods.dev/skills/jaganpro/sf-skills/sf-ai-agentforce-testing"><img src="https://agentmods.dev/badge/skills/jaganpro/sf-skills/sf-ai-agentforce-testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/jaganpro/sf-skills/sf-ai-agentforce-testing"><img src="https://agentmods.dev/badge/skills/jaganpro/sf-skills/sf-ai-agentforce-testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- Socket pass
- Snyk pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00085 | $0.02071 |
| Opus 5 | $0.00043 | $0.01035 |
| Sonnet 5 | $0.00017 | $0.00414 |
| Haiku 4.5 | $0.00009 | $0.00207 |
Grade A, and why
sf-ai-agentforce-testing scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
- Do **not** use raw `curl` for OAuth token validation in the ECA flow; use the provided credential tooling. How it starts
The opening of the file, as written. The whole thing — 248 lines — stays where its author put it; the contents beside it link to each section on GitHub.
sf-ai-agentforce-testing: Agentforce Test Execution & Coverage Analysis
Use this skill when the user needs formal Agentforce testing: multi-turn conversation validation, CLI Testing Center specs, topic/action coverage analysis, preview checks, or a structured test-fix loop after publish.
When This Skill Owns the Task
Use sf-ai-agentforce-testing when the work involves:
sf agent testworkflows- multi-turn Agent Runtime API testing
- topic routing, action invocation, context preservation, guardrail, or escalation validation
- test-spec generation and coverage analysis
- post-publish / post-activate test-fix loops
Delegate elsewhere when the user is:
- building or editing the agent itself → sf-ai-agentforce or sf-ai-agentscript
- running Apex unit tests → sf-testing
- creating seed data for actions → sf-data
- analyzing session telemetry / STDM traces → sf-ai-agentforce-observability
Core Operating Rules
- Testing comes after deploy / publish / activate.
- Use multi-turn API testing as the primary path when conversation continuity matters.
- Use CLI Testing Center as the secondary path for single-utterance and org-supported test-center workflows.
- Interactive and programmatic CLI preview use standard
sf org login webauthentication; ECA is only required for Agent Runtime API testing, not for live preview. - Fixes to the agent should be delegated to sf-ai-agentscript when Agent Script changes are needed.
- Do not use raw
curlfor OAuth token validation in the ECA flow; use the provided credential tooling.
Script path rule
Use the existing scripts under:
~/.claude/skills/sf-ai-agentforce-testing/hooks/scripts/
These scripts are pre-approved. Do not recreate them.
What ships with it
57 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- assets/agentscript-test-spec.yaml 10 KB
- assets/basic-test-spec.yaml 3.3 KB
- assets/cli-auth-guardrail-tests.yaml 9.0 KB
- assets/cli-deep-history-tests.yaml 10 KB
- assets/comprehensive-test-spec.yaml 12 KB
- assets/context-vars-test-spec.yaml 7.8 KB
- assets/custom-eval-test-spec.yaml 10 KB
- assets/escalation-tests.yaml 11 KB
- assets/guardrail-tests.yaml 10 KB
- assets/multi-turn-agentscript-comprehensive.yaml 7.2 KB
- assets/multi-turn-comprehensive.yaml 7.7 KB
- assets/multi-turn-context-preservation.yaml 4.1 KB
- assets/multi-turn-escalation-flows.yaml 4.1 KB
- assets/multi-turn-topic-routing.yaml 4.0 KB
- assets/standard-test-spec.yaml 5.8 KB
- assets/test-plan-template.yaml 3.2 KB
- CREDITS.md 2.7 KB
- hooks/scripts/agent_api_client.py 27 KB runs code
- hooks/scripts/agent_discovery.py 38 KB runs code
- hooks/scripts/credential_manager.py 17 KB runs code
- hooks/scripts/generate_multi_turn_scenarios.py 30 KB runs code
- hooks/scripts/generate-test-spec.py 23 KB runs code
- hooks/scripts/multi_turn_fix_loop.py 15 KB runs code
- hooks/scripts/multi_turn_test_runner.py 77 KB runs code
- hooks/scripts/parse-agent-test-results.py 17 KB runs code
- hooks/scripts/rich_test_report.py 9.4 KB runs code
- hooks/scripts/run-automated-tests.py 16 KB runs code
- hooks/scripts/test-fix-loop.sh 10 KB runs code
- hooks/scripts/trace_analyzer.py 20 KB runs code
- LICENSE 1.1 KB
- README.md 4.5 KB
- references/agent-api-reference.md 15 KB
- references/agentic-fix-loops.md 30 KB
- references/agentscript-agents.md 3.9 KB
- references/agentscript-testing-patterns.md 12 KB
- references/automated-testing.md 3.8 KB
- references/cli-commands.md 34 KB
- references/cli-testing-details.md 6.8 KB
- references/connected-app-setup.md 5.6 KB
- references/coverage-analysis.md 16 KB
- references/credential-convention.md 2.0 KB
- references/deep-conversation-history-patterns.md 12 KB
- references/eca-setup-guide.md 7.6 KB
- references/execution-protocol.md 2.7 KB
- references/interview-wizard.md 6.9 KB
- references/key-insights.md 1.5 KB
- references/known-issues.md 7.0 KB
- references/multi-turn-execution.md 4.4 KB
- references/multi-turn-testing.md 17 KB
- references/results-scoring.md 4.1 KB
- references/scoring-rubric.md 1001 B
- references/swarm-execution.md 5.8 KB
- references/test-plan-format.md 1.3 KB
- references/test-spec-reference.md 31 KB
- references/test-templates.md 1.5 KB
- references/topic-name-resolution.md 7.6 KB
- references/trace-analysis.md 9.3 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 248 lines · 85 tokens per session scan A 62b72df123f6
sf-ai-agentforce-testing is a skill published in the GitHub repository Jaganpro/sf-skills (423 stars, last pushed 4mo ago), licensed MIT. It adds 85 tokens to every session and 2,071 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
research-engineer
An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.
tika-eval-compare
Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".
neuron-evaluation-engineer
Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
atmos-validation
Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.