All-atom protein design using BoltzGen diffusion model. Use this skill when: (1) Need side-chain aware design from the start, (2) Designing around small molecules or ligands, (3) Want all-atom diffusion (not just backbone), (4) Require precise binding geometries, (5) Using YAML-based configuration. For structure…
Predict protein subcellular localization from amino acid sequence using BioT5. Use this skill when: (1) You have a protein sequence and want to know where it localizes in the cell, (2) You need to identify cellular compartment (nucleus, cytoplasm, membrane, etc.), (3) You want quick localization prediction without…
Query PubChem database for chemical structures, similar compounds, and bioactivity data. Use this skill when: (1) Converting drug name to molecular structure (SMILES, SDF), (2) Finding similar compounds for lead optimization, (3) Querying bioactivity data against protein targets, (4) Getting compounds active in…
Expert-in-the-loop retrosynthetic planning workflow. Use when you need to break down complex target molecules into available starting materials, design synthetic routes, or collaborate with a human chemist to refine a proposed synthesis pathway.
Autonomous Retrosynthetic Tree Search Loop. Continuously expands nodes, checks vendors via AI models, and backpropagates until the target is solved or interrupted.
Retrieve proteins with similar structures, sequences, or from the same family. Use this skill when: (1) Finding similar proteins or homologs, (2) Searching for proteins with similar 3D structure, (3) Performing sequence similarity search, (4) Discovering proteins in the same family.
Call accessible chromatin peaks from ATAC-seq BAM files, annotate peaks to genomic features and genes, and identify differentially accessible regions between experimental conditions.
Use this skill when a task involves Geneformer workflows, especially TranscriptomeTokenizer input preparation, tokenized .dataset generation, cell or gene classification with Classifier, embedding extraction with EmbExtractor, and in silico perturbation analysis with InSilicoPerturber.
Use this skill when a task involves the LangCell project for single-cell language-cell modeling, especially zero-shot cell type annotation, few-shot annotation, LangCell-CE finetuning, Geneformer-style tokenization, or preparing text descriptions for candidate cell identities and multimodal cell-text matching…
Use this skill when a task involves the local scGPT project in /DATA/disk0/zhaosy/home/scGPT, especially scGPT preprocessing and binning, checkpoint vocabulary matching, cell embedding extraction, reference mapping, fine-tuning scGPT for integration or annotation, or using scGPT tutorials for GRN, perturbation…
Probabilistic deep learning framework for single-cell multi-omics data analysis. Use this skill when: (1) Analyzing single-cell RNA-seq data with batch correction, (2) Integrating multi-modal data (CITE-seq, ATAC-seq, multi-omics), (3) Performing cell type annotation with scANVI, (4) Spatial transcriptomics…
Prepare your RNA-seq, proteomics, methylation, and other omics datasets for joint integration by applying per-assay normalization, cross-assay batch correction, feature ID alignment, and missing value handling.
Load, inspect, centroid, and extract features from raw LC-MS/MS data files. This is Step 1 of the proteomics pipeline — all downstream peptide identification and quantification steps require centroided, quality-checked spectra as input.
Search MS2 spectra against a protein sequence database to identify peptides and proteins in your sample. Apply target-decoy FDR filtering to control false discovery rate at both PSM and protein levels.
Complete single-cell RNA-seq analysis workflow built on Scanpy and AnnData. Use this skill when: (1) Loading diverse single-cell data formats (10X, h5ad, CSV), (2) Performing quality control and filtering, (3) Normalization, dimensionality reduction, and clustering, (4) Marker gene identification and cell type…
Use this skill when a task involves the local SToFM project in /DATA/disk0/zhaosy/home/SToFM, especially preprocessing spatial transcriptomics data for SToFM, generating cell embeddings with the cell encoder plus SE(2) Transformer pipeline, handling spatial coordinates, or preparing SToFM embeddings for downstream…
Load spatial transcriptomics data from Visium, Xenium, MERFISH, Slide-seq, and other platforms using Squidpy and SpatialData. Use this skill when: (1) Loading Visium spatial transcriptomics data from Space Ranger output, (2) Loading Xenium single-cell resolution spatial data, (3) Loading MERFISH, CosMx, or other…
Structure prediction using Boltz-2, an open biomolecular structure predictor. Use this skill when: (1) Predicting protein complex structures, (2) Validating designed binders, (3) Predicting protein-ligand complexes, (4) Using local GPU resources. For protein complex binding affinity evaluation, use prodigy.
Generate diverse lead compounds for a specific protein target using structure-based drug design with MolCraft. Use this skill when: (1) Designing drug candidates for a known protein target (PDB ID or disease name), (2) Generating structurally diverse molecules with optimized binding affinity, (3) Filtering candidates…
Generate comprehensive drug development progress reports for disease therapeutic targets. Use when user asks about target drug pipeline, clinical trials, or research progress. Triggers on phrases like "target report", "drug development progress", "clinical trial summary", "靶点报告", "药物研发进展", "竞品分析", "专利分析".
Modify molecules based on natural language descriptions using MolT5/BioT5 models. Use this skill when: (1) User wants to modify a molecule to improve specific properties (solubility, potency, etc.), (2) User provides a molecule and asks to "make it more X" or "improve Y", (3) User wants to generate molecule variants…
Query UniProt database for protein sequences, metadata, and search by criteria. Use this skill when: (1) Looking up protein information by UniProt accession ID, (2) Searching proteins by gene name, organism, function, or disease, (3) Retrieving comprehensive protein metadata including domains, PTMs, and annotations.