cli-test

A test plan for the blz command-line tool, a program operated from a terminal. It covers setup, search, source and index management, configuration, output formats, integrations, and error cases.

In plain words
What is it for?
Use it to check installation and help text, exercise search and source operations, verify JSON and text output, test shell and piping behavior, and confirm useful exit codes.
Why use it?
It helps reveal whether the command works correctly in normal use, with invalid input, broken networks, damaged caches, permission problems, and larger datasets.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/outfitter-dev/blz/cli-test
Clone the repo
git clone --depth 1 https://github.com/outfitter-dev/blz
Per session 9 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 483 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00009 $0.00483
Opus 5 $0.00005 $0.00242
Sonnet 5 $0.00002 $0.00097
Haiku 4.5 $0.00001 $0.00048

Measured yesterday against content hash 904c7dd09e90, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

cli-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.factory/commands/cli-test.md · 65 lines

How it starts

The opening of the file, as written. The whole thing — 65 lines — stays where its author put it; the contents beside it link to each section on GitHub.

blz CLI Comprehensive Test

Please perform a comprehensive test of the blz CLI tool with the following focus: $ARGUMENTS

Test Plan

1. Environment & Setup

  • Verify blz is properly installed and accessible
  • Check version information with blz --version
  • Validate configuration directory structure
  • Test help output for main command and subcommands

2. Core Functionality Tests

  • Search Operations: Test various search patterns, filters, and output formats
  • Source Management: Add, list, update, and remove sources
  • Index Operations: Test indexing, updates, and cache management
  • Configuration: Verify config file handling and environment variables

3. Edge Cases & Error Handling

  • Invalid input validation
  • Network connectivity issues
  • Corrupted cache handling
  • Permission errors
  • Large dataset performance

4. Output Format Testing

  • Text output formatting and readability
  • JSON output structure and validity
  • Error message clarity and helpfulness
  • Search result relevance and ranking

5. Integration Testing

  • Shell completion functionality
  • Pipe/redirection compatibility
  • Exit codes for scripting
  • Configuration precedence (env vars vs config files)

Expected Deliverables

  1. Test Summary: Overall health status of the CLI
  2. Performance Metrics: Search latency, indexing speed, memory usage
  3. Issue Report: Any bugs, inconsistencies, or UX problems found
  4. Improvement Suggestions: Specific recommendations with priority levels

Special Instructions

  • Use verbose output (-v or --verbose) when available to capture detailed logs
  • Test with both small and large document sets if possible
  • Verify that all documented features work as described in help text
  • Pay attention to consistency across similar commands (flags, output format, etc.)
  • Test error recovery and graceful degradation scenarios

If you identify any issues, please:

  1. Document exact reproduction steps
  2. Include relevant error messages or unexpected outputs
  3. Suggest specific fixes or improvements
  4. Prioritize issues by severity (critical, high, medium, low)

Read the full file on GitHub · 65 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 65 lines · 9 tokens per session scan A 904c7dd09e90

Subscribe to this mod's changes

cli-test is a command published in the GitHub repository outfitter-dev/blz (27 stars, last pushed 20d ago), licensed MIT. It adds 9 tokens to every session and 483 once invoked, about $0.0000 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.