Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add LeoLin990405/r-analytics-skill --skill r-nlpgit clone --depth 1 https://github.com/LeoLin990405/r-analytics-skillWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/leolin990405/r-analytics-skill/r-nlp)<a href="https://agentmods.dev/skills/leolin990405/r-analytics-skill/r-nlp"><img src="https://agentmods.dev/badge/skills/leolin990405/r-analytics-skill/r-nlp/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/leolin990405/r-analytics-skill/r-nlp"><img src="https://agentmods.dev/badge/skills/leolin990405/r-analytics-skill/r-nlp.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00029 | $0.01071 |
| Opus 5 | $0.00015 | $0.00535 |
| Sonnet 5 | $0.00006 | $0.00214 |
| Haiku 4.5 | $0.00003 | $0.00107 |
Grade A, and why
r-nlp scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 152 lines — stays where its author put it; the contents beside it link to each section on GitHub.
R NLP Skill
Sub-skills
| Sub-skill | Description |
|---|---|
| r-nlp-text | tidytext, quanteda, text2vec |
| r-nlp-topic | LDA, STM, topic modeling |
| r-nlp-sentiment | sentimentr, syuzhet, lexicons |
Natural Language Processing and text mining in R.
Core NLP Packages
| Package | Description |
|---|---|
| tidytext ★ | Tidy text mining (tidyverse style) |
| quanteda ★ | Quantitative text analysis |
| tm | Comprehensive text mining framework |
| text2vec ★ | Fast vectorization & word embeddings |
| NLP | Basic NLP functions |
| openNLP | Apache OpenNLP interface |
Text Processing
| Package | Description |
|---|---|
| stringr | Consistent string manipulation |
| stringi | ICU-based string processing |
| SnowballC | Snowball stemmers |
| koRpus | Text analysis package |
| utf8 | UTF-8 text handling |
Sentiment Analysis
| Package | Description |
|---|---|
| syuzhet | Sentiment extraction (3 dictionaries) |
| sentimentr | Sentence-level sentiment |
| tidytext | Sentiment lexicons (AFINN, Bing, NRC) |
Topic Modeling
| Package | Description |
|---|---|
| topicmodels | LDA and CTM topic models |
| LDAvis | Interactive topic model visualization |
| stm | Structural topic models |
Other
| Package | Description |
|---|---|
| zipfR | Word frequency distributions |
| MonkeyLearn | MonkeyLearn API interface |
| corporaexplorer | Dynamic text collection exploration |
Quick Examples
# tidytext workflow
library(tidytext)
library(dplyr)
# Tokenize
df %>%
unnest_tokens(word, text) %>%
anti_join(stop_words) %>%
count(word, sort = TRUE)
# Sentiment analysis
df %>%
unnest_tokens(word, text) %>%
inner_join(get_sentiments("bing")) %>%
count(sentiment)
# TF-IDF
df %>%
unnest_tokens(word, text) %>%
count(document, word) %>%
bind_tf_idf(word, document, n)
# quanteda
library(quanteda)
corpus <- corpus(texts)
tokens <- tokens(corpus, remove_punct = TRUE)
dfm <- dfm(tokens) %>%
dfm_remove(stopwords("en")) %>%
dfm_trim(min_termfreq = 5)
# Topic modeling
library(topicmodels)
dtm <- cast_dtm(df, document, word, n)
lda <- LDA(dtm, k = 5, control = list(seed = 1234))
topics <- tidy(lda, matrix = "beta")
# text2vec word embeddings
library(text2vec)
it <- itoken(texts, tokenizer = word_tokenizer)
vocab <- create_vocabulary(it)
vectorizer <- vocab_vectorizer(vocab)
tcm <- create_tcm(it, vectorizer, skip_grams_window = 5)
glove <- GloVe$new(rank = 50)
word_vectors <- glove$fit_transform(tcm, n_iter = 10)
What ships with it
18 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- r-nlp-sentiment/sentimentr/SKILL.md 3.5 KB
- r-nlp-sentiment/SKILL.md 3.1 KB
- r-nlp-sentiment/syuzhet/SKILL.md 3.1 KB
- r-nlp-text/hunspell/SKILL.md 1.7 KB
- r-nlp-text/quanteda/SKILL.md 1.6 KB
- r-nlp-text/SKILL.md 2.3 KB
- r-nlp-text/stringdist/SKILL.md 2.1 KB
- r-nlp-text/text2vec/SKILL.md 1.5 KB
- r-nlp-text/tidytext/SKILL.md 3.3 KB
- r-nlp-text/tm/SKILL.md 1.5 KB
- r-nlp-text/tokenizers/SKILL.md 2.0 KB
- r-nlp-topic/LDAvis/SKILL.md 3.8 KB
- r-nlp-topic/SKILL.md 2.1 KB
- r-nlp-topic/stm/SKILL.md 3.6 KB
- r-nlp-topic/topicmodels/SKILL.md 3.7 KB
- sub-skills/r-nlp-text/sub-skills/quanteda/SKILL.md 1.6 KB
- sub-skills/r-nlp-text/sub-skills/text2vec/SKILL.md 1.5 KB
- sub-skills/r-nlp-text/sub-skills/tm/SKILL.md 1.5 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 152 lines · 29 tokens per session scan A caec69b10106
r-nlp is a skill published in the GitHub repository LeoLin990405/r-analytics-skill (5 stars, last pushed 6mo ago), licensed MIT. It adds 29 tokens to every session and 1,071 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
bio-applied-molecular-evolution
Test Hardy-Weinberg equilibrium, simulate Wright-Fisher drift/selection, and compute dN/dS, Tajima's D, and Fst with NumPy/SciPy. Use for neutral theory, molecular clock divergence time, selection scans, or effective population size (Ne) questions.
advanced-string-structures
Build tries, Aho-Corasick, and suffix arrays with Kasai LCP to index DNA/text and match many patterns in one pass. Use for genome motif scanning, k-mer indexing, longest-repeat search, or BWA/FM-index groundwork.
ai-science-esm2-embeddings
Generate ESM2 protein embeddings (fair-esm/transformers) and predict structure with ESMFold. Use when embedding sequences, scoring mutations zero-shot, annotating protein function, or doing fast MSA-free structure prediction.
ai-science-geneformer-scgpt
Tokenize scRNA-seq via Geneformer gene-rank or scGPT expression-bin encoding; annotate cell types, simulate in-silico knockouts. Use for foundation-model cell annotation, Geneformer/scGPT tokenization, or perturbation prediction.
ai-science-zero-shot-mutation
Score protein point mutations zero-shot with ESM-1v/ESM-2 masked-LM log-odds, ensembled, benchmarked on ProteinGym DMS. Use when predicting mutation effects, ranking missense variants, scoring VUS fitness with no labels.
bio-applied-advanced-ngs
Assemble genomes de novo: greedy OLC, de Bruijn graph/Eulerian path, N50/L50/NG50 stats, SPAdes/Flye/hifiasm CLI usage. Use when choosing k-mer size, picking an assembler for Illumina/ONT/HiFi reads, or scoring contiguity.