Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/opendcai/dataflow-webui/text2qa-sample-evaluatornpx skills add OpenDCAI/DataFlow-WebUI --skill text2qa-sample-evaluatorgit clone --depth 1 https://github.com/OpenDCAI/DataFlow-WebUIWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/opendcai/dataflow-webui/text2qa-sample-evaluator)<a href="https://agentmods.dev/skills/opendcai/dataflow-webui/text2qa-sample-evaluator"><img src="https://agentmods.dev/badge/skills/opendcai/dataflow-webui/text2qa-sample-evaluator.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00031 | $0.00921 |
| Opus 5 | $0.00015 | $0.00461 |
| Sonnet 5 | $0.00006 | $0.00184 |
| Haiku 4.5 | $0.00003 | $0.00092 |
Grade A, and why
text2qa-sample-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 128 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Text2QASampleEvaluator Operator Reference
Text2QASampleEvaluator evaluates QA pairs across 4 dimensions, generating 8 output columns (grades + feedbacks for each dimension).
1. Import
from dataflow.operators.core_text import Text2QASampleEvaluator
2. Constructor
Text2QASampleEvaluator(
llm_serving=llm_serving,
)
| Parameter | Required | Default | Description |
|---|---|---|---|
llm_serving |
Yes | None | LLM service object |
3. run() Signature
op.run(
storage=self.storage.step(),
input_question_key="question",
input_answer_key="answer",
)
# returns: list of 8 output column names
| Parameter | Required | Default | Description |
|---|---|---|---|
storage |
Yes | None | Storage step object |
input_question_key |
No | "generated_question" |
Question column name |
input_answer_key |
No | "generated_answer" |
Answer column name |
4. Output Columns (8 columns)
| Column Name (Default) | Description |
|---|---|
question_quality_grades |
Question quality scores |
question_quality_feedbacks |
Question quality feedback |
answer_alignment_grades |
Answer alignment scores |
answer_alignment_feedbacks |
Answer alignment feedback |
answer_verifiability_grades |
Answer verifiability scores |
answer_verifiability_feedbacks |
Answer verifiability feedback |
downstream_value_grades |
Downstream value scores |
downstream_value_feedbacks |
Downstream value feedback |
Note: Column names use plural suffix (grades/feedbacks), not singular.
5. Usage Example
from dataflow.operators.core_text import Text2QASampleEvaluator
from dataflow.serving import APILLMServing_request
from dataflow.utils.storage import FileStorage
class MyPipeline:
def __init__(self):
self.storage = FileStorage(
first_entry_file_name="./data/qa_pairs.jsonl",
cache_path="./cache",
file_name_prefix="step",
cache_type="jsonl"
)
self.llm_serving = APILLMServing_request(
api_url="https://api.openai.com/v1/chat/completions",
key_name_of_api_key="DF_API_KEY",
model_name="gpt-4o",
max_workers=10
)
self.evaluator = Text2QASampleEvaluator(
llm_serving=self.llm_serving
)
def forward(self):
self.evaluator.run(
storage=self.storage.step(),
input_question_key="question",
input_answer_key="answer"
)
if __name__ == "__main__":
pipeline = MyPipeline()
pipeline.forward()
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 128 lines · 31 tokens per session scan A 40fc4dec039c
text2qa-sample-evaluator is a skill published in the GitHub repository OpenDCAI/DataFlow-WebUI (224 stars, last pushed 8d ago), licensed Apache-2.0. It adds 31 tokens to every session and 921 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
article-writing
Write articles, guides, blog posts, tutorials, newsletter issues, and other long-form content in a distinctive voice derived from supplied examples or brand guidance. Use when the user wants polished written content longer than a paragraph, especially when voice consistency, structure, and credibility matter.
ljg-learn
Deep concept anatomist that deconstructs any concept through 8 exploration dimensions (history, dialectics, phenomenology, linguistics, formalization, existentialism, aesthetics, meta-philosophy) and compresses insights into an epiphany. Use when user asks to explain, dissect, or deeply understand a concept, term, or…
eli5
Explain research, papers, or technical ideas in plain English with minimal jargon, concrete analogies, and clear takeaways. Use when the user says "ELI5 this", asks for a simple explanation of a paper or research result, wants jargon removed, or asks what something technically dense actually means.
code-documenter
Use when adding docstrings, creating API documentation, or building documentation sites. Invoke for OpenAPI/Swagger specs, JSDoc, doc portals, tutorials, user guides.
deck-course-module
暖纸背景 + Playfair, 左侧学习目标常驻, 含 MCQ 自测页.
pedagogy-review
Holistic pedagogical review of a lecture deck (.qmd or .tex). Checks narrative arc, prerequisite assumptions, worked examples, notation clarity, and deck-level pacing. Use when user says "pedagogy review", "does this teach well?", "is the flow right?", "will students follow?", "review the narrative", or before…