Fix CI failures and performance regressions for a Julia package PR by iterating - triage the latest CI results, fix one root cause, verify locally, push. Use when a PR's GitHub Actions or Buildkite CI is failing, or when CI jobs run slower than they do on the main branch. Verifies CPU and GPU compilation locally…
Use when the user wants to search cloud GPUs across providers, provision a GPU pod, install AI frameworks (ComfyUI, SwarmUI, vLLM, Ollama, Open WebUI, Axolotl, H2O LLM Studio), pull models from HuggingFace / Civitai / URLs to a unified store, sync workspaces to S3, manage idle-pod lifecycle, or track GPU spend — all…
Use when (re)indexing a Sweet Search project. Runs the full-profile indexer with GPU model prewarming (CoreML cascade on M3+, candle Metal on M1/M2, ORT CPU elsewhere), kills ORT CPU models during indexing to avoid memory contention, and rewarms them for query readiness on completion. Incremental runs under 20 files…
Fix CI failures and performance regressions for a Julia package PR by iterating - triage the latest CI results, fix one root cause, verify locally, push. Use when a PR's GitHub Actions or Buildkite CI is failing, or when CI jobs run slower than they do on the main branch. Verifies CPU and GPU compilation locally…
Use when the user wants to run ComfyUI image/video generation on a rented remote GPU instead of paid API subscriptions: rent a GPU box, bootstrap the stack, connect the agent over SSH, drive workflows via the ComfyUI API, and deliver results to Telegram.
Search the web, arxiv, Crossref, PubMed, OpenAlex, or Wikipedia, and fetch page text, using common/tools/search.sh. Use whenever the user asks to search online, look something up, or find papers while served by a local model, and whenever the built-in WebSearch tool fails or is silently dropped by the endpoint.
Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables 20x realtime transcription. SRT/VTT/TTML/CSV subtitles, speaker diarization, URL/YouTube input, batch processing with ETA, transcript search, chapter detection, per-file language map.
Use when optimizing, tuning, diagnosing, or profiling CUDA, CUTLASS, Triton, PyTorch, vLLM, TensorRT-LLM, or another GPU workload; when assessing an NCU, Nsys, or PyTorch Profiler report; or when the test workload, correctness checks, measurement path, or target environment is incomplete.
Rent a GPU on vast.ai, run experiments on it over SSH/rsync, pull artifacts back, and destroy it. Use when a task needs a GPU (training, CUDA, large-model inference).
Provides guidance for writing custom Metal compute kernels using MLX's mx.fast.metalkernel() API for Apple Silicon GPUs (M1, M2, M3, M4). Covers kernel patterns for RMSNorm, LayerNorm, softmax, attention variants (causal, sliding window, GQA), and reductions. Includes profiling, debugging, simdgroupmatrix MMA (M3+)…
Use AutoDL safely through its official developer APIs and documentation. Trigger for AutoDL balance, Container Instance Pro, elastic deployment, images, GPU stock, NFS, duration packages, WeChat notifications, container metrics, billing, storage, data retention, SSH/JupyterLab/VSCode, CUDA environments…
Use Project Bourne to plan, execute, reproduce, inspect, or trace scientific and engineering workloads when durable provenance matters. Trigger for simulations, numerical solvers, ML or training, GPU and HPC runs on local compute, Slurm, PBS, or LSF, reproducible experiments, failed attempts, experiment comparisons…
★not rated 0 5d agoA91 tokens
originalApache-2.0
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: