Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
gptomics avatar

Bio Read Alignment Star Alignment

  • 3 installs
  • 1.1k repo stars
  • Updated July 25, 2026
  • gptomics/bioskills

Align RNA-seq reads with STAR splice-aware alignment, including two-pass mode for novel splice junctions.

About

Aligns RNA-seq reads with STAR, a splice-aware aligner supporting two-pass mode for novel junction discovery. Developers use it for RNA-seq data requiring accurate spliced alignment to a reference.

  • Splice-aware RNA-seq alignment with genome index
  • Two-pass mode for novel splice-junction discovery

Bio Read Alignment Star Alignment by the numbers

  • 3 all-time installs (skills.sh)
  • Ranked #1,661 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/gptomics/bioskills --skill bio-read-alignment-star-alignment

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs3
repo stars1.1k
Last updatedJuly 25, 2026
Repositorygptomics/bioskills

What it does

Align RNA-seq reads with STAR splice-aware alignment, including two-pass mode for novel splice junctions.

Files

SKILL.mdMarkdownGitHub ↗

STAR RNA-seq Alignment

Generate Genome Index

# Basic index generation
STAR --runMode genomeGenerate \
    --runThreadN 8 \
    --genomeDir star_index/ \
    --genomeFastaFiles reference.fa \
    --sjdbGTFfile annotation.gtf \
    --sjdbOverhang 100    # Read length - 1

Index with Specific Read Length

# For 150bp reads, use sjdbOverhang=149
STAR --runMode genomeGenerate \
    --runThreadN 8 \
    --genomeDir star_index_150/ \
    --genomeFastaFiles reference.fa \
    --sjdbGTFfile annotation.gtf \
    --sjdbOverhang 149

Basic Alignment

# Paired-end alignment
STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn reads_1.fq.gz reads_2.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate

Single-End Alignment

STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn reads.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate

Two-Pass Mode

# Two-pass mode for better novel junction detection
STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn r1.fq.gz r2.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate \
    --twopassMode Basic

Quantification Mode

# Output gene counts (like featureCounts)
STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn r1.fq.gz r2.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate \
    --quantMode GeneCounts

Output: sample_ReadsPerGene.out.tab with columns: 1. Gene ID 2. Unstranded counts 3. Forward strand counts 4. Reverse strand counts

ENCODE Options

# ENCODE recommended settings
STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn r1.fq.gz r2.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate \
    --outSAMunmapped Within \
    --outSAMattributes NH HI AS NM MD \
    --outFilterType BySJout \
    --outFilterMultimapNmax 20 \
    --outFilterMismatchNmax 999 \
    --outFilterMismatchNoverReadLmax 0.04 \
    --alignIntronMin 20 \
    --alignIntronMax 1000000 \
    --alignMatesGapMax 1000000 \
    --alignSJoverhangMin 8 \
    --alignSJDBoverhangMin 1

Fusion Detection

# For chimeric/fusion detection
STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn r1.fq.gz r2.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate \
    --chimSegmentMin 12 \
    --chimJunctionOverhangMin 8 \
    --chimOutType Junctions WithinBAM SoftClip \
    --chimMainSegmentMultNmax 1

Output Files

FileDescription
*Aligned.sortedByCoord.out.bamSorted BAM file
*Log.final.outAlignment summary statistics
*Log.outDetailed log
*SJ.out.tabSplice junctions
*ReadsPerGene.out.tabGene counts (if --quantMode)
*Chimeric.out.junctionFusion candidates (if chimeric)

Memory Requirements

# Reduce memory for limited systems
STAR --genomeLoad NoSharedMemory \
    --limitBAMsortRAM 10000000000 \  # 10GB for sorting
    ...

# For very large genomes, limit during index generation
STAR --runMode genomeGenerate \
    --limitGenomeGenerateRAM 31000000000 \  # 31GB
    ...

Shared Memory Mode

# Load genome into shared memory (for multiple samples)
STAR --genomeLoad LoadAndExit --genomeDir star_index/

# Run alignments (faster startup)
STAR --genomeLoad LoadAndKeep --genomeDir star_index/ ...

# Remove from memory when done
STAR --genomeLoad Remove --genomeDir star_index/

Key Parameters

ParameterDefaultDescription
--runThreadN1Number of threads
--sjdbOverhang100Read length - 1
--outFilterMultimapNmax10Max multi-mapping
--alignIntronMax0Max intron size
--outFilterMismatchNmax10Max mismatches
--outSAMtypeSAMOutput format
--quantMode-GeneCounts for counting
--twopassModeNoneBasic for two-pass

Related Skills

  • rna-quantification/featurecounts-counting - Alternative counting
  • rna-quantification/alignment-free-quant - Salmon/kallisto alternative
  • differential-expression/deseq2-basics - Downstream DE analysis
  • read-qc/fastp-workflow - Preprocess reads

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.