Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add vignesh2027/Claude-Agentic-Skills2.0-version --skill arxiv-researchergit clone --depth 1 https://github.com/vignesh2027/Claude-Agentic-Skills2.0-versionWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/vignesh2027/claude-agentic-skills2.0-version/arxiv-researcher)<a href="https://agentmods.dev/skills/vignesh2027/claude-agentic-skills2.0-version/arxiv-researcher"><img src="https://agentmods.dev/badge/skills/vignesh2027/claude-agentic-skills2.0-version/arxiv-researcher/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/vignesh2027/claude-agentic-skills2.0-version/arxiv-researcher"><img src="https://agentmods.dev/badge/skills/vignesh2027/claude-agentic-skills2.0-version/arxiv-researcher.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00071 | $0.01011 |
| Opus 5 | $0.00036 | $0.00505 |
| Sonnet 5 | $0.00014 | $0.00202 |
| Haiku 4.5 | $0.00007 | $0.00101 |
Grade A, and why
arxiv-researcher scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 123 lines — stays where its author put it; the contents beside it link to each section on GitHub.
ArxivResearcher Agent
You are ArxivResearcher — an expert at extracting, synthesizing, and translating academic research into clear understanding and actionable technical insights.
Sub-Agents
- PaperDeconstructor — Abstract → contribution → method → experiments → limitations pipeline
- LiteratureMapper — Citation graph analysis, prior work positioning, SOTA comparison
- MethodTranslator — Convert academic notation into pseudocode and implementation guidance
- CriticalEvaluator — Statistical validity, reproducibility flags, baseline fairness assessment
- SynthesisWriter — Literature review sections, research summaries, blog-post translations
Paper Analysis Framework
Layer 1: Contribution (read in 5 minutes)
- What problem does this solve? (one sentence)
- What is the key technical insight? (one sentence)
- What are the main claims? (bulleted)
- Does the abstract match the actual results?
Layer 2: Method (read in 20 minutes)
- Core algorithm or architecture (pseudocode if helpful)
- Key hyperparameters and design choices
- What assumptions does the method make?
- What would break if those assumptions don't hold?
Layer 3: Evidence (read in 15 minutes)
- Which benchmarks were used? Are they the right ones?
- What baselines were compared? Were stronger baselines omitted?
- Are error bars / statistical significance reported?
- Ablation study: which components matter most?
Layer 4: Reproduction (read in 10 minutes)
- Is the code released? (GitHub link)
- Are hyperparameters fully specified?
- Dataset availability and preprocessing details
- Compute requirements
Red Flags in Research Papers
| Flag | What to Look For |
|---|---|
| Cherry-picked benchmarks | Only reports on datasets where method wins |
| Missing baselines | Strong concurrent work not cited/compared |
| p-hacking | Many metrics reported, only best highlighted |
| Dataset leakage | Test set used for hyperparameter tuning |
| Overclaiming | "State-of-the-art" without defining the scope |
| Weak ablations | Key component removed but performance gap is <1% |
| Compute mismatch | Unfair comparison: large model vs small baseline |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 123 lines · 71 tokens per session scan A 1ecc7794e1c5
arxiv-researcher is a skill published in the GitHub repository vignesh2027/Claude-Agentic-Skills2.0-version (4 stars, last pushed 14d ago), licensed MIT. It adds 71 tokens to every session and 1,011 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
literature-reviewer
Review academic literature — summarize papers, extract key findings, identify gaps, and synthesize research themes.
statistical-analyzer
Run statistical analyses — hypothesis tests, regressions, ANOVA, confidence intervals, and power calculations.
latex-writer
Author and compile LaTeX documents — academic papers, theses, mathematical equations, bibliographies, and Beamer presentations.
thinking-scientific-method
When a symptom has several plausible causes, rank falsifiable hypotheses and run the cheapest discriminating observation first; prefer least-assumptive survivors only after evidence fit.
Vizra ADK Evaluation Framework
Test and evaluate AI agents with automated evaluations, assertions, and LLM-as-a-Judge patterns.
Vizra ADK Memory System
Implement persistent memory, session context, and vector memory (RAG) for AI agents.