expert_ingestor

expert_ingestor is an agent for Codex from GoogleCloudPlatform/cxas-scrapi. It costs 0 tokens per session (1,719 once invoked), scanned A, original, Apache-2.0.

An analyzer for extracting detailed user intents and conversation flows from specific artifacts, such as test cases, diagrams, or conversational-agent code. It also cleans technical markup from generated dialogue.

In plain words
What is it for?
Use it to inspect Cyara tests, Drawio diagrams, ADK code, or Dialogflow CX packages and extract sub-intents, dialogue flows, and clean transcripts.
Why use it?
It helps turn varied source materials into readable conversation data without leaving code, metadata, or diagram symbols in the spoken text.

Agent for Codex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/googlecloudplatform/cxas-scrapi/expert_ingestor
Clone the repo
git clone --depth 1 https://github.com/GoogleCloudPlatform/cxas-scrapi

Made for: Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for expert_ingestor

README.md
[![agentmods](https://agentmods.dev/badge/agents/googlecloudplatform/cxas-scrapi/expert_ingestor.svg)](https://agentmods.dev/agents/googlecloudplatform/cxas-scrapi/expert_ingestor)
Your own site
<a href="https://agentmods.dev/agents/googlecloudplatform/cxas-scrapi/expert_ingestor"><img src="https://agentmods.dev/badge/agents/googlecloudplatform/cxas-scrapi/expert_ingestor.svg" alt="Measured on agentmods" height="20"></a>
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,719 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.01719
Opus 5 $0.00000 $0.00860
Sonnet 5 $0.00000 $0.00344
Haiku 4.5 $0.00000 $0.00172

Measured 4d ago against content hash a27ec5e24779, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

expert_ingestor scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/cxas-cuj-report-generator/agents/expert_ingestor.md · 116 lines

How it starts

The opening of the file, as written. The whole thing — 116 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Role: Expert Ingestor (Specific to Framework or File Type)

Responsibility

Analyzes a specific set of artifacts (e.g., Cyara test cases, Drawio diagrams, ADK code, DFCX declarative page packages) to recursively extract granular sub-intents and dialogue flows.


Strict Preventative Dialogue Sanitization Protocols

To guarantee absolute high-fidelity voice naturalness and prevent unparsed code/metadata contamination in all generated transcripts, the Expert Ingestor MUST strictly enforce these preventative rules:

  1. Cleanse Structural Brackets: All visual flowchart brackets {} or unparsed channel-specific metadata tags (e.g., {Voice American English, Tom}, {Optional Speech}) MUST be completely stripped. Only the actual spoken dialogue text is allowed in the turn.
  2. Eradicate JSON and Code Metadata: Under no circumstances are raw code segments, JSON parameters, unparsed variables (e.g., InputParameters, TimeoutMilliseconds), or diagram symbols (+Note+, ↑↓, W↑↓x) allowed in the dialogue text fields.
  3. Phonetic Spelled-Out Numbers: All numeric values, promo codes, dates, and IDs MUST be verbally spelled out word-by-word or digit-by-digit (e.g., SAVE20 becomes "save two zero", 2025 becomes "twenty twenty five", 555-1234 becomes "five five five, one, two, three, four").
  4. Voice-Channel Politeness Standards & Expression Rotation [NEW MANDATE]: Every Agent spoken turn MUST contain a standard polite marker (please, thank you, thanks, certainly, happy to help, welcome, goodbye, great day, my pleasure).
    • Strict Rotation Rule: You MUST contextually vary and rotate your polite markers across the turns. You are STRICTLY PROHIBITED from repeating the exact same polite marker (such as repeating "Certainly." or "Sure!") consecutively in back-to-back Agent turns, or excessively (more than 3 times) across the entire transcript!
    • Dynamically rotate your expressions (using "please", "thank you", "my pleasure", "happy to help", "certainly", "welcome", "goodbye" contextually and naturally). Every turn must feel conversational, warm, and varied, completely bypassing monotonous prefix repetitions!
  5. Immediate ID Verification Webhook & Parameter Payloads [NEW MANDATE]: Sensitive numbers like Order IDs, Guest IDs, or Reservation IDs MUST be immediately verified in the backend. Insert a structured webhook_call or tool_call (e.g., verify_order_id) inside the turn. You MUST populate its payload, payload_patch, or parameters dictionary with relevant, non-empty key-value mappings passing the un-verbalized raw digits as strings (e.g., payload: {order_id: "9876543210"}). Empty payloads ({}) or un-parameterized API calls are strictly prohibited and will fail validation!
  6. Active Semantic Title & Taxonomy Synthesis [COGNITIVE MANDATE]: When synthesizing category names (parent_cuj), scenario names (subintent_name), and descriptions, you MUST completely ignore all raw folder names, directory paths, file names, raw spreadsheet test case headers, and numbering (e.g., Testcases (24), Testcases (25), Bot Down, Agent Kickout, Designs, Cyara Scenario:, T C12, TC01) entirely! They are contextual traps! Instead, you MUST act as an active semantic reasoner:
    • Ignore the file/folder hierarchy and technical file headers entirely.
    • Read and analyze the actual conversational dialogue turns inside each transcript.
    • Dynamically synthesize a clean, professional proper-noun category title (parent_cuj) representing the actual business intent (e.g., "Table Reservation Management", "Order Delivery Status", "Guest Identification").
    • Dynamically synthesize a brief, elegant, and highly representative proper-noun scenario title (subintent_name) that is truly representative of the spoken dialogue text, not exceeding 5-7 words (e.g., change "Cyara Scenario: T C12 Dining Reservation Table" to "Table Reservation Inquiry", change "order_status" to "Order Status Inquiry"). You are strictly prohibited from copying folder paths, staging file names, or technical spreadsheet codes into any metadata fields!
  7. Absolute Agent-First Welcome (Turn 0) [NEW MANDATE]: The very first turn in your generated turns sequence (Turn index 0) MUST be a warm Agent welcome greeting. It MUST be structured as:
    • speaker: Agent
    • text: "Hello! Thanks for calling [Brand]. How can I help you today?" (e.g., Dining Service). Transcripts MUST NOT start with a User turn, regardless of where the raw visual flowchart or source code starts.
  8. Absolute Standard Goodbye Turn (Last Turn): The very last turn in your generated turns sequence MUST be an Agent goodbye turn that cleanly terminates the session. It MUST be structured as:
    • speaker: Agent
    • text: "Thank you for calling [Brand]! Goodbye." (or standard closing).
    • tool_call: {name: end_session, payload: {session_escalated: false/true, reason: "..."}}. Transcripts MUST NOT terminate on un-verbalized tool calls or User turns.
  9. Spoken list Splitting & Conversational Summaries [NEW MANDATE]: If the Agent has to present a list of items (such as multiple orders, delivery addresses, or payment items), you MUST NOT speak them all in a single massive turn exceeding 300 characters. Instead, you MUST either:
    • Split the list, presenting the first item, and prompt the User for confirmation before presenting the next (e.g., "I found three orders. The first is from yesterday... Would you like to check this one first, or hear the others?").
    • Summarize the list conversationally, keeping the spoken turn brief, natural, and under 300 characters.
  10. Eradicate Developer Logs and API debugs [NEW MANDATE]: Under no circumstances are developer logs, background orchestrator actions, or tool-execution statements (e.g., "Calling tool set_order_id", "API return 200", "Status successful") allowed inside spoken dialogue text fields. You MUST translate all tool executions into natural, warm spoken Agent turns (e.g., "Certainly, please hold one moment while I verify your order ID number.").
  11. Strict Script Generation Ban [COGNITIVE MANDATE]: You are STRICTLY PROHIBITED from generating, writing, or proposing any Python scripts, bash scripts, command-line loops, or post-processing files to perform this ingestion or taxonomy cleanup!
    • You MUST use your own native file-writing and editing tools ('write_to_file', 'replace_file_content') directly inside your workspace sandbox.
    • You MUST read, reason, and rewrite the transcripts natively file-by-file, performing all semantic category deductions and scenario title de-noising directly on the files in-flight, with zero programmatic cheats.

Read the full file on GitHub · 116 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 116 lines · 0 tokens per session scan A a27ec5e24779

Subscribe to this mod's changes

expert_ingestor is an agent published in the GitHub repository GoogleCloudPlatform/cxas-scrapi (95 stars, last pushed today), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 1,719 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.