machine learning agents

84 tagged machine learning, measured the same way as everything else here.

Browse within: data-science 21forecasting 15mlforecast 15neuralforecast 15deep-learning 14prompt-engineering 9Data Processing 7cpp 7data-pipeline 7jupyterlab 7Apple Silicon 6MLOps 6agent-skill 6benchmarking 6

AnastasiyaW/codex-claude-code-config

Agent

Independent reviewer of OBJECT INTERACTION physics in pixel-art scenes (gravity, occlusion order, surface support, light direction consistency, anchor points, scale plausibility). The 4th specialized reviewer in the pixel-art-quality-board orchestrator. Use when the user asks "do objects interact correctly", "is…

147 7d ago A 156 tokens original MIT

AnastasiyaW/codex-claude-code-config

Agent

Orchestrator agent that runs 4 specialized pixel-art reviewers in parallel (style + animation + composition + interaction), then synthesizes their verdicts into a single PASS/NEEDSWORK/REJECT decision with prioritized fixes. Use when the user asks to "review my pixel art quality", "score this animation"…

147 7d ago A 163 tokens original MIT

agentic-workflows

27

CliMA/ClimaAtmos.jl

Agent

GitHub Agentic Workflows (gh-aw) - Create, debug, and upgrade AI-powered workflows with intelligent prompt routing.

125 4d ago A 24 tokens original Apache-2.0

gke-cluster-runner

28

vlasenkoalexey/tpu_performance_autoresearch_wiki

Agent Claude Code

Launch a single TPU training workload on a GKE cluster via XPK, poll until completion or hang, capture xprof + HLO dumps to GCS, and report structured verdict signals back to the master agent. Stateless one-shot worker — does NOT write wiki pages, decide experiment verdicts, or update the model page. Use for every…

54 6d ago A 100 tokens original MIT

kernel-verifier

29

vlasenkoalexey/tpu_performance_autoresearch_wiki

Agent Claude Code

Independent verifier for kernel-family experiments (the Roles section's verifier for the pallas lane). Given a final candidate kernel + the naive baseline, it independently re-benchmarks both in a fresh process, re-runs numerical parity, captures traces/LLO dumps with the canonical flag set, runs the hypothesis-firing…

54 6d ago A 168 tokens original MIT

profile-analyzer

30

vlasenkoalexey/tpu_performance_autoresearch_wiki

Agent Claude Code

Analyze a single completed experiment's xprof trace + HLO dump and return structured ## Profile + ## HLO Dump markdown sections that slot directly into the experiment page. Phases: Phase 1 walks xprof (bucket attribution, dominant ops, memory profile); Phase 2 walks HLO (module sizes, fusion verification, regression…

54 6d ago A 250 tokens original MIT

google-colab-expert

31

andisab/swe-marketplace

Agent

Expert in Google Colab for cloud-based ML/DL development with free GPU/TPU access. Specializes in Colab 2025 features (Gemini AI integration, google.colab.ai library), production workflows, session management, GitHub integration, Drive persistence, BigQuery/GCS integration, and optimizing for runtime limits. Use for…

21 14d ago A 282 tokens original MIT

data-jupyter-expert

32

andisab/swe-marketplace

Agent

Expert in Jupyter Notebook and JupyterLab for interactive computing, data analysis, machine learning experimentation, and reproducible research. Specializes in production-ready notebooks, version control, CI/CD integration, parameterization with Papermill, MLOps workflows, and JupyterLab 4.4+ modern features including…

21 14d ago A 260 tokens original MIT

agent-brain

35

cpuguy96/StepCOVNet

Agent

Human-maintained inventory of .cursor/rules and skills. Updated by the agent during agent-brain-refresh — not generated by a script.

21 8d ago A 0 tokens original Apache-2.0

cpuguy96/StepCOVNet

Agent

Status (2026-06-30): Tide gate PASS — scratch iter175 / champion v8. Do not start new tide iter runs unless reproducing. Next AR work: gate-10song-smoke (EXPERIMENTLOG.md § Current phase).

21 8d ago A 0 tokens original Apache-2.0

self-journal

37

cpuguy96/StepCOVNet

Agent

Iterative improvements to agent behavior, process, and conventions on this project. Not a substitute for EXPERIMENTLOG.md or DISCUSSIONNOTES.md — those hold research findings; this holds how we work better.

21 8d ago A 0 tokens original Apache-2.0

phase_4

39

jeremylongshore/plugins-nixtla

Agent

Contract: This agent runs automated verification scripts to validate Phase 2-3 conclusions with empirical data.

9 7d ago A 0 tokens

phase_5

40

jeremylongshore/plugins-nixtla

Agent

Contract: This agent synthesizes all phase outputs into actionable recommendations with implementation plans.

9 7d ago A 0 tokens

coverage-skeptic

41

Amal-David/mlx-porting-skill

Agent

Look for blind spots, unsupported architecture families, missing validation gates, and overclaimed optimizations.

5 3d ago A 0 tokens original Apache-2.0

data-profiler

44

StamKavid/last-ds-mile

Agent

Fast structural profiling sweep for a dataset — shape, dtypes, missingness, cardinality, duplicate keys. Use during /ds-data or /ds-explore for a quick first-pass profile. Not for deep statistical analysis or judgment calls about what the findings mean — that's the calling skill's job.

3 24d ago A 64 tokens original MIT

ds-reviewer

45

StamKavid/last-ds-mile

Agent

Runs the ds-method discipline checklist against a notebook or pipeline before /ds-report — baseline present, validation strategy sound, slice performance checked, metric matches the problem. Use before final reporting/handoff, or when asked to sanity-check a DS pipeline end to end. Not for hunting leakage specifically…

3 24d ago A 70 tokens original MIT

leakage-auditor

46

StamKavid/last-ds-mile

Agent

Adversarially hunts for target leakage across a feature pipeline — features that encode the target directly, temporal leakage where future information reaches training data, and validation-split leakage. Use before /ds-model or /ds-report when a metric looks implausibly good, or as a final check before a pipeline…

3 24d ago A 83 tokens original MIT

OpenMarkets

47

danchev/openmarkets

Agent

This chat mode is designed for analyzing market trends, providing insights on financial markets, and assisting with investment strategies. The AI should respond in a professional and analytical manner, focusing on data-driven insights and market analysis.

3 24d ago A 45 tokens AGPL-3.0