llm-streaming-response-handler

llm-streaming-response-handler is a skill for Claude Code from curiositech/some_claude_skills. It costs 95 tokens per session (3,289 once invoked), scanned A, original, MIT.

A guide to showing AI-generated text as it arrives instead of waiting for the complete response. It uses Server-Sent Events, an HTTP method for sending updates from a server to a browser.

In plain words
What is it for?
It supports live chatbot responses, typing-style displays, real-time assistants, code previews, and progressive document summaries.
Why use it?
It makes chat and generation interfaces feel responsive and supports cancellation and recovery when a stream fails.

Skill for Claude Code

Written for Claude Code: allowed-tools in frontmatter.

Part of the llm-streaming-response-handler plugin — 1 skill shipped together

Good fit It supports live chatbot responses, typing-style displays, real-time assistants, code previews, and progressive document summaries.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/curiositech/some_claude_skills/llm-streaming-response-handler
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add curiositech/some_claude_skills --skill llm-streaming-response-handler
Clone the repo
git clone --depth 1 https://github.com/curiositech/some_claude_skills

Made for: Claude Code.

Or install llm-streaming-response-handler, the plugin that ships this one along with the rest of its 1 skill.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for llm-streaming-response-handler

README.md
[![agentmods](https://agentmods.dev/badge/skills/curiositech/some_claude_skills/llm-streaming-response-handler/github.svg)](https://agentmods.dev/skills/curiositech/some_claude_skills/llm-streaming-response-handler)
Your own site
<a href="https://agentmods.dev/skills/curiositech/some_claude_skills/llm-streaming-response-handler"><img src="https://agentmods.dev/badge/skills/curiositech/some_claude_skills/llm-streaming-response-handler/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for llm-streaming-response-handler

Your own site · 80×15
<a href="https://agentmods.dev/skills/curiositech/some_claude_skills/llm-streaming-response-handler"><img src="https://agentmods.dev/badge/skills/curiositech/some_claude_skills/llm-streaming-response-handler.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 95 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,289 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00095 $0.03289
Opus 5 $0.00048 $0.01644
Sonnet 5 $0.00019 $0.00658
Haiku 4.5 $0.00010 $0.00329

Measured 5d ago against content hash 6085d5b7ad7a, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

llm-streaming-response-handler scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/stream_tester.ts, scripts/token_counter.ts), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/llm-streaming-response-handler/SKILL.md · 533 lines

How it starts

The opening of the file, as written. The whole thing — 533 lines — stays where its author put it; the contents beside it link to each section on GitHub.

LLM Streaming Response Handler

Expert in building production-grade streaming interfaces for LLM responses that feel instant and responsive.

When to Use

Use for:

  • Chat interfaces with typing animation
  • Real-time AI assistants
  • Code generation with live preview
  • Document summarization with progressive display
  • Any UI where users expect immediate feedback from LLMs

NOT for:

  • Batch document processing (no user watching)
  • APIs that don't support streaming
  • WebSocket-based bidirectional chat (use Socket.IO)
  • Simple request/response (fetch is fine)

Quick Decision Tree

Does your LLM interaction:
├── Need immediate visual feedback? → Streaming
├── Display long-form content (&gt;100 words)? → Streaming
├── User expects typewriter effect? → Streaming
├── Short response (&lt;50 words)? → Regular fetch
└── Background processing? → Regular fetch

Technology Selection

Server-Sent Events (SSE) - Recommended

Why SSE over WebSockets for LLM streaming:

  • Simplicity: HTTP-based, works with existing infrastructure
  • Auto-reconnect: Built-in reconnection logic
  • Firewall-friendly: Easier than WebSockets through proxies
  • One-way perfect: LLMs only stream server → client

Timeline:

  • 2015-2020: WebSockets for everything
  • 2020: SSE adoption for streaming APIs
  • 2023+: SSE standard for LLM streaming (OpenAI, Anthropic)
  • 2024: Vercel AI SDK popularizes SSE patterns

Streaming APIs

Provider Streaming Method Response Format
OpenAI SSE data: {"choices":[{"delta":{"content":"token"}}]}
Anthropic SSE data: {"type":"content_block_delta","delta":{"text":"token"}}
Claude (API) SSE data: {"delta":{"text":"token"}}
Vercel AI SDK SSE Normalized across providers

Common Anti-Patterns

Anti-Pattern 1: Buffering Before Display

Novice thinking: "Collect all tokens, then show complete response"

Problem: Defeats the entire purpose of streaming.

Read the full file on GitHub · 533 lines

Files

What ships with it

6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 533 lines · 95 tokens per session scan A 6085d5b7ad7a

Subscribe to this mod's changes

llm-streaming-response-handler is a skill published in the GitHub repository curiositech/some_claude_skills (216 stars, last pushed 3d ago), licensed MIT. It adds 95 tokens to every session and 3,289 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

copilotkit-upgrade

Use when migrating a CopilotKit v1 application to v2 -- updating package imports, replacing deprecated hooks and components, switching from GraphQL runtime to AG-UI protocol runtime, and resolving breaking API changes.

CopilotKit/CopilotKit · 48 tokens

nextjs-pages-router

Set up tRPC in Next.js Pages Router with createNextApiHandler, createTRPCNext, withTRPC HOC, SSR via ssr option and ssrPrepass, SSG via createServerSideHelpers with getStaticProps, and server-side helpers for getServerSideProps prefetching.

trpc/trpc · 67 tokens

with-tanstack-query

Compose Angular Query with signal-owned Table filtering, sorting, and pagination state using reactive query options, manual row-model boundaries, direct query data, server counts, and valid injection context.

TanStack/table · 42 tokens

auth-web-cloudbase

CloudBase Web Authentication Quick Guide for frontend integration after auth-tool has already been checked. Provides concise and practical Web authentication solutions with multiple login methods and complete user management.

TencentCloudBase/CloudBase-AI-Toolkit · 38 tokens

service-digital-engagement-channel-configure

Configures and deploys enhanced chat Messaging Channels for Messaging for In-App and Web (MIAW). Use when the user needs to create, deploy, and activate a messaging channel configured with Omni-Channel Flow, Omni-Channel Queue, User, or Agentforce Service Agent routing. Generates MessagingChannel metadata, deploys it…

forcedotcom/sf-skills · 173 tokens

om-system-extension

Extend installed Open Mercato modules through UMES enrichers, interceptors, mutation guards, widgets, menus, entity extensions, events, component/page replacements, and overrides. Use for "extend core", "add field/column/action", "hide page", "intercept API", "UMES", or "rozszerz moduł".

open-mercato/open-mercato · 73 tokens