bio-chipseq-chromatin-state-segmentation
Segments the genome into chromatin states from combinatorial histone modification and chromatin factor ChIP-seq data. Uses ChromHMM (multivariate HMM on binarized signal, v1.27), Segway (Dynamic Bayesian Network on continuous signal), EpiSegMix (flexible-distribution HMM with duration modeling, 2024), EpiLogos (multi-biosample visualization), IDEAS (cell-type-aware joint), and full-stack ChromHMM (Vu Ernst 2022) for cross-cell-type segmentations. Handles state-count selection (15 vs 18 vs 25 states), binarization choice, OverlapEnrichment / NeighborhoodEnrichment downstream analysis, and cross-biosample integration. Use when learning chromatin states from a histone mark panel, characterizing learned states by genomic feature enrichment, or comparing chromatin landscapes across cell types.
What this skill does
## Version Compatibility
Reference examples tested with: ChromHMM 1.27+, Segway 3.0+, EpiSegMix 1.0+, EpiLogos (Meuleman lab), IDEAS 1.20+, samtools 1.19+, bedtools 2.31+. ChromHMM requires Java 8+; runs as `java -mx<MEMORY> -jar ChromHMM.jar <command>`.
# Chromatin State Segmentation
**"Integrate multiple histone modification ChIP-seq tracks into chromatin states"** -> Learn a small set of recurring combinatorial patterns of histone marks (active promoter, active enhancer, poised enhancer, polycomb-repressed, heterochromatic, transcribed, etc.) and segment the genome by which state each region belongs to. Output: per-state genomic intervals, state-by-mark emission matrix, and state-state transition matrix.
- CLI (canonical): ChromHMM `BinarizeBam` -> `LearnModel` -> `OverlapEnrichment` / `NeighborhoodEnrichment`
- CLI (continuous signal): Segway `train` -> `posterior` -> `identify`
- CLI (flexible distributions): EpiSegMix (2024)
- Visualization across biosamples: EpiLogos (Meuleman lab)
- Cell-type-aware joint: IDEAS
Chromatin state segmentation requires a panel of histone marks; minimum 4-5 marks (e.g., H3K4me3, H3K27ac, H3K4me1, H3K36me3, H3K27me3) for meaningful states. With fewer marks, simpler peak-based annotation (chipseq/peak-annotation) is more appropriate.
## Tool Taxonomy
| Tool | Method | Strength | Fails when |
|------|--------|----------|------------|
| **ChromHMM** (Ernst & Kellis 2012; v1.27 current) | Multivariate HMM on binarized 200 bp bins | Canonical; widely used; integrated with Roadmap Epigenomics 15-state model; mature toolchain | Binarization throws away signal quantitation; default 200 bp bins may be too coarse for sharp boundaries |
| **Segway** (Hoffman 2012) | Dynamic Bayesian Network on continuous signal | Higher resolution; uses signal magnitudes not binarized | More complex setup; slower; less standardized output |
| **EpiSegMix** (Schmitz, Aggarwal, Laufer, Walter, Salhab, Rahmann 2024 Bioinformatics 40:btae178) | HMM with flexible read-count distributions + duration modeling | Modern; handles both narrow and broad mark distributions in one model | Newer; smaller user base |
| **EpiLogos** (Meuleman lab) | Multi-biosample visualization tool | Built on top of ChromHMM/Segway segmentations; compare ChromHMM states across 100s of biosamples | Visualization tool, not a segmentation method itself |
| **IDEAS** (Zhang 2016) | Cell-type-aware joint inference | Across-cell-type segmentation respecting cell-type identity | Slower; complex parameter tuning |
| **EpiCSeg** (Mammana 2015) | Negative binomial mixture | Read-count-based; doesn't need binarization | Less standardized output |
| **GenoSTAN** | HMM with various emission distributions | Flexible | Less actively developed |
| **Roadmap 25-state model** (Kundaje 2015) | ChromHMM 25-state precomputed model | Reference for cross-cell-type interpretation | Tied to Roadmap mark panel (5 core marks) |
| **Full-stack ChromHMM** (Vu Ernst 2022) | 100-state segmentation across 1032 datasets / 127 reference epigenomes | Comprehensive cross-tissue annotation | Computationally intensive to retrain |
## ChromHMM Workflow
ChromHMM is the de facto standard. The workflow has 4 stages:
### Step 1: Binarize ChIP-seq signal
```bash
# Build cellMarkFileTable: cell_type<TAB>mark<TAB>file<TAB>(optional control)
cat > cellMarkFileTable.txt << EOF
GM12878 H3K4me3 gm12878_h3k4me3.bam gm12878_input.bam
GM12878 H3K27me3 gm12878_h3k27me3.bam gm12878_input.bam
GM12878 H3K27ac gm12878_h3k27ac.bam gm12878_input.bam
GM12878 H3K4me1 gm12878_h3k4me1.bam gm12878_input.bam
GM12878 H3K36me3 gm12878_h3k36me3.bam gm12878_input.bam
EOF
# Binarize BAMs into 200 bp bins; emission = whether mark exceeds Poisson threshold
java -mx16G -jar ChromHMM.jar BinarizeBam \
-b 200 \
chromsizes_hg38.txt \
bam_dir/ \
cellMarkFileTable.txt \
binarized_output/
```
Output: per-chromosome `_binary.txt` files, one row per 200 bp bin, columns = marks, values 0/1.
### Step 2: Learn model
```bash
# Train HMM with N states; common choices: 15, 18, 25
# 15 states: Ernst & Kellis 2011 model; canonical
# 18 states: extends with additional regulatory states
# 25 states: Roadmap Epigenomics extended model
java -mx16G -jar ChromHMM.jar LearnModel \
-p 8 \
binarized_output/ \
model_15state/ \
15 \
hg38
# Output: model_15state.txt (emission + transition matrices),
# emissions_15.png (visualization), transitions_15.png,
# per-chromosome _segments.bed (state assignments)
# AND automatically runs OverlapEnrichment + NeighborhoodEnrichment
```
### Step 3: Interpret states from emission matrix
| Roadmap 15-state assignments (canonical) |
|-------------------------------------------|
| 1_TssA — Active TSS (high H3K4me3, H3K27ac) |
| 2_TssAFlnk — Flanking TSS (H3K4me3, H3K27ac, lower) |
| 3_TxFlnk — Transcript flanking |
| 4_Tx — Strong transcription (H3K36me3, H3K79me2 if available) |
| 5_TxWk — Weak transcription |
| 6_EnhG — Enhancer in gene body (H3K4me1, H3K27ac) |
| 7_Enh — Generic enhancer (H3K4me1, H3K27ac) |
| 8_ZNF/Rpts — Zinc-finger / repeats |
| 9_Het — Heterochromatin (H3K9me3) |
| 10_TssBiv — Bivalent TSS (H3K4me3 + H3K27me3) |
| 11_BivFlnk — Bivalent flanking |
| 12_EnhBiv — Bivalent enhancer (H3K4me1 + H3K27me3) |
| 13_ReprPC — Polycomb-repressed (H3K27me3) |
| 14_ReprPCWk — Weak Polycomb |
| 15_Quies — Quiescent (no signal) |
### Step 4: Functional enrichment of states
```bash
# OverlapEnrichment: enrichment of each state for external feature sets
java -mx16G -jar ChromHMM.jar OverlapEnrichment \
model_15state/GM12878_15_segments.bed \
/path/to/anchor_files/ \
enrichment_output/GM12878 \
-labels
# NeighborhoodEnrichment: enrichment relative to anchor positions (e.g., TSS)
java -mx16G -jar ChromHMM.jar NeighborhoodEnrichment \
model_15state/GM12878_15_segments.bed \
/path/to/tss_anchors.txt \
enrichment_output/GM12878_TSS \
-labels
```
Anchor files: BED files of features (CGIs, repeats, conserved elements, etc.) for OverlapEnrichment; position files for NeighborhoodEnrichment.
## Choosing State Count
| States | Use case | Mark panel size |
|--------|----------|-----------------|
| 8-10 | Initial exploration; small mark panel (3-4 marks) | 3-5 marks |
| 15 | Roadmap Epigenomics canonical | 5 core (H3K4me3, H3K27ac, H3K4me1, H3K36me3, H3K27me3) |
| 18 | Roadmap extended (adds bivalent states, fine enhancer subtypes) | 5-7 marks |
| 25 | Roadmap Epigenomics extended; cross-cell-type compatibility | 6+ marks |
| 50+ | Full-stack model (Vu Ernst 2022) | Many marks across many cell types |
**Practical workflow:** Train at N=15, 18, 25; compare emission matrices; choose the smallest N where biology is interpretable. Higher N risks over-segmentation (state splitting random variation).
## Segway Workflow
```bash
# Train Segway model (10 states by default; supports more)
segway train \
--num-labels 25 \
--num-instances 3 \
--resolution 100 \
chromsizes_hg38.bed \
h3k4me3.bw h3k27ac.bw h3k4me1.bw h3k36me3.bw h3k27me3.bw \
--traindir traindir/
# Posterior inference + segmentation
segway posterior --traindir traindir/ --identifydir identifydir/ \
chromsizes_hg38.bed \
h3k4me3.bw h3k27ac.bw h3k4me1.bw h3k36me3.bw h3k27me3.bw
# Output: identifydir/segway.bed (state assignments)
```
Segway uses bigWig (continuous signal) vs ChromHMM binarized binary. Trade-off: more information per region (continuous) but more complex training.
## EpiLogos Visualization
EpiLogos doesn't perform segmentation; it visualizes existing ChromHMM/Segway segmentations across many biosamples (epilogos.org).
```bash
# Use precomputed ChromHMM segmentations across multiple cell types
# Web interface: https://epilogos.altius.org/
# Local: github.com/meuleman/epilogos
```
Useful for: cross-cell-type comparison; identifying tissue-specific regulatory states; cohort-level chromatin landscape summaries.
## Full-Stack ChromHMM (Vu Ernst 20Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.