Claude
Skills
Sign in
Back

bio-clip-seq-clip-preprocessing

Included with Lifetime
$97 forever

Preprocess CLIP-seq reads (eCLIP, iCLIP, iCLIP2, iCLIP3, irCLIP, PAR-CLIP, FLASH) with protocol-specific UMI extraction, adapter trimming, length filtering, and post-alignment PCR-duplicate collapse. Use when raw CLIP FASTQ must be turned into deduplicated, crosslink-preserving BAM input for peak calling; choosing between two-pass and single-pass adapter trimming; deciding minimum read length; or mapping UMI patterns to specific eCLIP/iCLIP/iCLIP2/iCLIP3 library preps.

General

What this skill does


## Version Compatibility

Reference examples tested with: umi_tools 1.1.5+, cutadapt 4.6+, fastp 0.23.4+, samtools 1.19+, pysam 0.22+, picard 3.1+, preseq 3.2+.

Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
- CLI: `<tool> --version` then `<tool> --help` to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

# CLIP-seq Preprocessing

**"Preprocess raw CLIP reads into UMI-deduplicated, alignable FASTQ"** -> Extract random barcodes, trim adapters without disturbing the 5' truncation site, length-filter to remove unmappable shorts, and (post-alignment) collapse PCR duplicates by UMI + position. The 5' end of the read carries the iCLIP/eCLIP truncation signature one base downstream of the protein-RNA crosslink; preserving this base is the single most important constraint of CLIP preprocessing.

- CLI (eCLIP, paired-end): `umi_tools extract --bc-pattern=NNNNNNNNNN --stdin R1.fq.gz --read2-in R2.fq.gz --stdout R1_umi.fq.gz --read2-out R2_umi.fq.gz`
- CLI (iCLIP/iCLIP2, single-end): `umi_tools extract --bc-pattern=NNNXXXXNN --extract-method=string --stdin R1.fq.gz --stdout R1_umi.fq.gz` (3+2 random Ns flanking a 4 nt library barcode; demultiplex by the X positions first if multiplexed)
- CLI (PAR-CLIP): `umi_tools extract --bc-pattern=NNNN ...` (most protocols use 4 nt random barcodes; verify the lab's exact prep)
- CLI (3' trim only, eCLIP convention): `cutadapt -a AGATCGGAAGAGCACACGTCT -A AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT --quality-base 33 --quality-cutoff 6 -m 18 -o R1.trim.fq.gz -p R2.trim.fq.gz R1_umi.fq.gz R2_umi.fq.gz`

The eCLIP convention is: trim quality and adapter from the 3' end ONLY. The 5' end of read 2 (in paired-end eCLIP) is the truncation site of the RT enzyme at the protein-RNA adduct, located one nucleotide downstream of the crosslink. Trimming the 5' end discards that exact base. `cutadapt -g`, `fastp --trim_front1`, and aggressive quality trimming of the 5' end are all banned for CLIP unless a documented protocol-specific reason exists.

## Read Structure by Protocol

| Protocol | UMI length and location | Truncation-site read | Adapter set | Notes |
|----------|-------------------------|----------------------|-------------|-------|
| eCLIP (Van Nostrand 2016 / ENCODE) | 10 nt at R1 5' end (random nucleotides preceding insert) | R2 5' end (after R1 UMI is stripped) = crosslink -1 | Illumina TruSeq R1 + R2 + inline X1A/X1B inverted adapter for two-pass | ENCODE pipeline normative |
| seCLIP (single-end eCLIP) | 10 nt at R1 5' end | R1 5' end | TruSeq R1 only | ENCODE accepts both eCLIP and seCLIP |
| iCLIP (Konig 2010) | 5 nt random (NNNXXXXNN: 3 N + 4 X library barcode + 2 N), single-end | R1 5' end after barcode strip | L3 adapter at 3' | Multiplexed - demultiplex by the 4 X bases |
| iCLIP2 (Buchbender 2020) | 5 or 9 nt random (NNNXXXXNN or longer), single-end | R1 5' end | L3 adapter at 3' | Increased complexity vs iCLIP; same UMI pattern |
| iCLIP3 (Buchbender et al, bioRxiv 2026.03.01.708747) | 10 nt random + dual sample index, single-end | R1 5' end | TruSeq | Silica-column RNA isolation; non-radioactive; streamlined low-input protocol. Preprint - verify final published version before pinning a pipeline. |
| irCLIP (Zarnegar 2016) | 5 nt random + barcode (similar to iCLIP) | R1 5' end | IR700 adapter + standard TruSeq sequencing adapter | Infrared replaces 32P; otherwise iCLIP-like |
| PAR-CLIP (Hafner 2010) | 0-4 nt depending on prep | T->C transitions within reads (NOT a truncation method) | TruSeq R1 | UMI optional; rely on T->C signature for CL |
| FLASH (Aktas 2020) | Sample barcode + UMI in custom adapter | R1 5' | Custom L3 design | 1.5 day protocol; adapter design proprietary to MPI |
| miCLIP / miCLIP2 (Linder 2015 / Kortel 2021) | iCLIP-style barcodes | R1 5' = m6A -1 (truncation OR C-to-T) | iCLIP L3 | m6A-specific |
| STAMP (Brannan 2021) | NA (no UV) | NA (C-to-U editing) | 10x or bulk RNA-seq adapters | Antibody-free editing-based; preprocess as RNA-seq |

If the read structure differs from the lab's documentation, run `seqkit head -n 1000 R1.fq.gz | seqkit stats -a` and inspect the first 12 bases of 100 random reads. Random-barcode positions show ~25% base composition per position; library/sample barcodes will be fixed across reads.

## Critical Choice: One Adapter Pass vs Two

**One pass (single-end iCLIP, PAR-CLIP):** Single 3' adapter; cutadapt with `-a <L3>` and `-m 18` is sufficient.

**Two passes (eCLIP, paired-end):** The eCLIP library design ligates an inverted X1A/X1B inline adapter that can appear at either end of a short fragment due to read-through. The ENCODE pipeline does pass 1 with the standard 3' adapter, then pass 2 trims a residual 5' adapter that read-through events leave on read 2. Trimming a 5' adapter from R2 is a special case: cutadapt's `-G <ADAPTER>` (uppercase) anchored at R2's 5' end is the right invocation. Do NOT use `-g` (lowercase) for R1 5' trimming - that destroys the truncation site.

**Cutadapt full eCLIP-style invocation:**

```bash
# Pass 1 - 3' adapter (both reads)
cutadapt \
    -a AGATCGGAAGAGCACACGTCT \
    -A AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT \
    --quality-base 33 -q 6 \
    -m 18 \
    -o R1.p1.fq.gz -p R2.p1.fq.gz \
    R1.umi.fq.gz R2.umi.fq.gz

# Pass 2 - read-through 5' adapter on R2 only (ENCODE eCLIP)
cutadapt \
    -G GATCGTCGGACTGTAGAACTCTGAAC \
    --quality-base 33 -q 6 \
    -m 18 \
    -o R1.p2.fq.gz -p R2.p2.fq.gz \
    R1.p1.fq.gz R2.p1.fq.gz
```

The `-q 6` is intentionally permissive. Aggressive `-q 20` quality trimming chews back the 5' end of R2 and breaks truncation-based crosslink-site detection. The `-m 18` minimum-length cutoff is non-negotiable: reads shorter than 18 nt are functionally unmappable (multi-map prevalence > 50%, see CIMS analysis discussions in the CLIP review literature) and must be discarded before alignment.

## Per-Protocol Failure Modes

### eCLIP -- 5' end trimmed by mistake

**Trigger:** Pipeline written for generic RNA-seq applied to eCLIP; `fastp` defaults trim both ends; `cutadapt -g` invoked on R2.

**Mechanism:** The 5' end of R2 in eCLIP is the truncation site = crosslink -1. Quality-trim from 5' or untargeted 5' adapter trim discards that exact base.

**Symptom:** Downstream PureCLIP / iCount / CTK CITS analysis fails to call crosslink sites; peak calling still works but the peaks lose nucleotide precision.

**Fix:** Use `-q 6` (3' only by default in cutadapt) and never `--trim_front2` in fastp. Validate by inspecting BAM read 2 5' positions: 60-90% of unique R2 5' positions should map within 100 nt windows around known RBP binding motifs, not be uniformly distributed.

### iCLIP / iCLIP2 -- Demultiplex confusion with UMI

**Trigger:** Multiplexed iCLIP library; user runs `umi_tools extract` with `NNNNNNNNN` (9 N) when the actual prep is `NNNXXXXNN` (5 random + 4 sample barcode).

**Mechanism:** umi_tools treats the entire 9-base prefix as UMI, losing the sample identity in the middle 4 bases. Reads from different samples are merged.

**Symptom:** Library complexity inflated artificially; per-sample read counts seem high but binding profiles look averaged across samples; sample-specific motifs absent.

**Fix:** Demultiplex BEFORE UMI extraction. Use `umi_tools extract --bc-pattern=NNNXXXXNN --extract-method=string --filter-cell-barcode` with a whitelist of the 4 X-base barcodes; or split the FASTQ with `je demultiplex` against the library barcode table first, then `umi_tools extract --bc-pattern=NNNNN` (the 5 surviving random Ns).

### PAR-CLIP -- T->C mistaken for sequencing error

**Trigger:** Standard variant-calling pipeline applied to PAR-CLIP without recognizing the T->C signature.

**Mechanism:** PAR-CLIP's diagnostic mutation is T->C (4SU adduct pairs with G during RT). Upon
Files: 3
Size: 28.3 KB
Complexity: 43/100
Category: General

Related in General