Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
26 Мб
Скачать
224 Bioinformatics of Autoimmune Diseases
FIGURE 7.1 Representative human protein–DNA complexes illustrating four major DNA-binding motifs: (A) HSF1 with a helix–turn–helix domain (PDB: 5D5U); (B) KLF4 zinc nger domain (PDB: 6VTX); (C) Max leucine zipper domain (PDB: 1HLO); and (D) PAX3 homeodomain (PDB: 3CMY). All structures were rendered by the author using PyMOL and publicly available PDB les.
physiological or pathological conditions, including in autoimmune diseases, where their activity is often dysregulated.
In the immune system, TFs play indispensable roles in the development, differentiation, and function of immune cells. They control lineage-specic gene expression programs, govern cytokine production, and maintain the delicate balance between immune activation and immune suppression. Examples of such TFs include FOXP3, which is required for the function of regulatory T cells, and STAT family members, which mediate cytokine signaling and inuence immune cell fate decisions.
TFs operate within complex regulatory networks that integrate multiple signaling pathways, including those activated by cytokines, pattern recognition receptors, and environmental stimuli. These networks are modulated by epigenetic modications such as histone acetylation and DNA methylation, which can enhance or restrict TF binding to DNA. In autoimmune diseases, this regu­latory balance is often disrupted. Overactive TFs may drive the sustained expression of inamma­tory genes, while loss of function or reduced expression can impair the formation or function of regulatory immune cells. Because of their central roles, TFs have emerged as both biomarkers of disease activity and potential targets for therapeutic intervention.
Table 7.1 includes a list of TFs known to play key roles in the development or regulation of
autoimmune diseases, either through genetic mutations, altered expression, or dysregulated path­ways. This table was compiled by the author based on ndings from multiple peer-reviewed stud­ies, including Smith et al. (2020), Zhang et al. (2018), and others. Together, these TFs form an intricate network of gene regulation that maintains immune homeostasis. When their expression is disrupted, whether through genetic mutation, epigenetic modication, or chronic inammatory signaling, the immune system may lose tolerance to self and initiate or perpetuate autoimmune disease. Understanding their expression proles and regulatory pathways is critical for developing targeted therapies to restore immune balance.
7.3 HISTONE MODIFICATIONS AND CHROMATIN DYNAMICS
Histones are core structural proteins that help package DNA into chromatin, a compact and orga­nized form that ts within the cell nucleus. Each nucleosome, the fundamental unit of chromatin consists of approximately 147 base pairs of DNA wrapped around a histone octamer (composed of two copies each of H2A, H2B, H3, and H4) and histone H1 (linker histone) which binds to the DNA
225 DNA-Protein Interactions and Autoimmune Diseases
TA BL E 7.1 Transcription Factors (TFs) Known to Play Roles in Autoimmune Diseases
Transcription Autoimmune Diseases Factor Involved Function/Role
AIRE APS-1, Type 1 diabetes Central tolerance in thymus; promotes expression of tissue-specic
antigens.
FOXP3 IPEX syndrome Master regulator of Tregs (regulatory T cells); maintains peripheral
tolerance.
STAT3 Lupus, RA, MS, Crohn’s Regulates cytokine signaling and T-cell differentiation; hyperactivation
leads to inammation. STAT1 Lupus, MS, RA IFN signaling; contributes to chronic inammation. NF-κB (RELA,
NFKB1) IRF5 SLE, RA, Sjögren’s Promotes expression of type I interferons and inammatory cytokines. IRF7 SLE Major regulator of type I interferon pathway; important in antiviral and
BACH2 Type 1 diabetes, MS, Maintains T-cell balance (effector versus regulatory); suppresses
RORγt (RORC) T-bet (TBX21) IBD, MS, RA GATA3 Allergy, autoimmunity Promotes Th2 differentiation; imbalance with Th1/Th17 contributes to
BLIMP1 SLE, IBD Controls B- and T-cell differentiation; involved in antibody-secreting
(PRDM1) plasma cell development. c-MAF Lupus, RA Regulates IL-10 expression in regulatory T cells and other immune cells. IKZF1 (Ikaros) Lupus, leukemia, RA Controls lymphocyte development; mutations can lead to dysregulated
PAX5 Autoimmunity (via B cell Essential for B cell development and identity.
RUNX1 Autoimmune cytopenias Regulates hematopoiesis and T-cell lineage development. CREM SLE Represses IL-2 expression; altered function promotes SLE-associated
EGR2/3 Lupus, T1D Induced during T-cell anergy and tolerance; dysregulation leads to
RA, SLE, psoriasis, IBD Pro-inammatory signaling; regulates cytokine gene expression.
autoimmune responses.
Crohn’s excessive inammation.
MS, RA, psoriasis Drives Th17 cell differentiation; Th17 cells are pro-inammatory.
Regulates Th1 cell differentiation; involved in IFN-γ production.
disease.
immune responses.
dysfunction)
T-cell defects.
activation.
between nucleosomes and helps compact chromatin into higher-order structures, playing a key role in stabilizing chromatin architecture and regulating gene accessibility (see Section 3.2.6). Each his­tone in the nucleosome octamer possesses a histone tail, which is a exible, unstructured extension emerging primarily from the N-terminal region (and in some cases the C-terminal region) of the protein. These tails protrude from the nucleosome core and are enriched in positively charged amino acids such as lysine and arginine. They serve as key platforms for post-translational modications, including acetylation, methylation, phosphorylation, ubiquitination, and sumoylation. These chemi­cal modications play a vital role in regulating chromatin dynamics and gene expression by altering nucleosome stability, modulating the accessibility of DNA to transcriptional machinery, or acting as binding sites for chromatin remodeling factors and other regulatory proteins.
These post-translational chemical modications serve as molecular signals that recruit or repel transcriptional regulators and chromatin remodeling complexes. Histone acetylation, particularly at lysine residues on histones H3 and H4, neutralizes the positive charge of histones and decreases their afnity for DNA. This results in a more open chromatin conformation that is permissive to transcription. Methylation, depending on the site and number of methyl groups, can either activate
226 Bioinformatics of Autoimmune Diseases
or repress gene transcription. For example, trimethylation of histone H3 at lysine 4 (H3K4me3) is commonly associated with active gene promoters, whereas trimethylation at lysine 27 (H3K27me3) correlates with gene repression.
In the context of autoimmune diseases, normal patterns of histone modication are often altered. These changes can result in the misexpression of genes involved in immune signaling, antigen pre­sentation, and inammatory responses (Allis & Jenuwein, 2016; Hedrich et al., 2014). In systemic lupus erythematosus (SLE), for instance, reduced histone acetylation at the FOXP3 locus has been linked to impaired regulatory T-cell function. In rheumatoid arthritis (RA), increased histone acety­lation in synovial immune cells facilitates the transcription of genes encoding pro-inammatory cytokines such as TNF-alpha and IL-6. Mapping these changes in histone marks offers insight into disease mechanisms and points to new therapeutic approaches, including the use of HDAC inhibi­tors to restore appropriate gene expression patterns.
7.4 EPIGENETIC REGULATION OF IMMUNE IDENTITY
Epigenetic regulation refers to heritable changes in gene function that occur without altering the DNA sequence itself. These changes include histone modications, DNA methylation, chromatin remodeling, and the activity of non-coding RNAs. Together, they create a exible but stable regula­tory framework that guides immune cell development and function.
During hematopoiesis, epigenetic mechanisms are responsible for activating or silencing lineage­specic gene programs, leading to the differentiation of pluripotent progenitor cells into mature immune cells such as T cells, B cells, macrophages, and dendritic cells. Once differentiation is complete, epigenetic marks continue to play a role in determining immune cell responsiveness, plas­ticity, and memory. For instance, the differentiation of helper T-cell subsets, including Th1, Th2, Th17, and regulatory T cells, is orchestrated by unique epigenetic signatures that regulate access to lineage-dening TFs and cytokine genes.
External factors such as infections, diet, stress, and exposure to environmental toxins can inu­ence the epigenetic landscape of immune cells. These exposures can lead to lasting epigenetic changes, a phenomenon known as “epigenetic memory” or “trained immunity.” While this process can enhance immune protection against recurring threats, it can also promote persistent inam­mation in autoimmune diseases. Long-term epigenetic remodeling may lock immune cells into a pathogenic state, even in the absence of continued exposure to the original trigger.
Epigenetic control also plays a key role in silencing retrotransposons and non-coding elements of the genome that could otherwise provoke immune activation. Loss of this silencing can mimic viral infection by generating double-stranded RNA or unmethylated DNA, thereby activating innate immune sensors and exacerbating inammation. These ndings highlight the importance of epi­genetic regulation in preserving immune tolerance and preventing autoimmune disease (Sun et al.,
2015).
7.5 ChIP-Seq AND THE EPIGENOMIC LANDSCAPE OF AUTOIMMUNITY
Chromatin immunoprecipitation followed by sequencing, or ChIP-Seq, is a powerful technique used to study protein–DNA interactions on a genome-wide scale. By using antibodies to isolate DNA fragments bound by specic proteins, such as TFs or histone modications, ChIP-Seq allows researchers to map regulatory elements across the genome and understand how gene expression is controlled.
In the study of autoimmune diseases, ChIP-Seq has been instrumental in identifying changes in the binding patterns of TFs and the distribution of histone modications. For example, in SLE and RA, ChIP-Seq has revealed abnormal binding of TFs such as STAT1, IRF5, and NF-κB at key immune genes. These changes often correspond with altered gene expression proles and reect the rewiring of regulatory circuits that promote chronic inammation and loss of immune tolerance.
227 DNA-Protein Interactions and Autoimmune Diseases
ChIP-Seq also allows for the identication of active and repressed chromatin regions through histone marks such as H3K27ac and H3K4me3. These marks dene enhancers and promoters that are engaged in transcriptional activity. In autoimmune diseases, the locations and intensities of these marks are often shifted, revealing previously unrecognized regulatory elements, sometimes referred to as latent or cryptic enhancers that become active in disease states.
When combined with other genomic technologies, such as RNA sequencing and ATAC-Seq , ChIP-Seq provides a multidimensional view of gene regulation. These integrative approaches help pinpoint causal relationships between epigenomic changes and transcriptional outcomes. Furthermore, many genetic risk variants identied through genome-wide association studies are found in non-coding regions that coincide with ChIP-Seq peaks. This overlap suggests that dis­ease susceptibility is often mediated through changes in gene regulation rather than protein-coding mutations.
Overall, ChIP-Seq offers valuable insights into the molecular mechanisms underlying immune dysregulation in autoimmunity (Barski et al., 2007; Scharer et al., 2019).
By employing ChIP-Seq across a variety of immune cell types, researchers have signicantly advanced our understanding of the epigenetic mechanisms driving autoimmune diseases. These studies typically focus on immune cells that play pivotal roles in immune regulation and disease pathogenesis, using them as the primary sources for DNA extraction in ChIP-Seq experiments. The key immune cell types examined are outlined below.
CD4+ T cells are a critical component of adaptive immunity and differentiate into distinct func­tional subsets, including Th1, Th17, and regulatory T cells (Tregs). Each subset is dened by unique transcriptional and epigenetic signatures. ChIP-Seq has been instrumental in identifying histone modications such as H3K27ac and H3K4me3 at key regulatory loci in these subsets. Furthermore, TFs such as FOXP3, T-bet, and RORγt have been mapped using ChIP-Seq to reveal alterations in binding patterns associated with autoimmune diseases like multiple sclerosis (MS) and RA. These studies have highlighted changes in enhancer activity and silencing of critical regulatory regions, offering insight into cytokine dysregulation and loss of immune tolerance in affected individuals.
B cells, responsible for producing antibodies and presenting antigens, also exhibit epigenomic alterations in autoimmunity. In SLE, ChIP-Seq has uncovered disease-specic super-enhancers that drive expression of genes involved in autoantibody production and hyperactive B cell states. TFs such as IRF4 and PAX5 display aberrant binding in these cells, reecting rewired regulatory net­works that support a pathogenic phenotype.
Dendritic cells serve as key antigen-presenting cells and gatekeepers of peripheral tolerance. When ChIP-Seq is applied to dendritic cells from autoimmune models or patient samples, it reveals transcriptional shifts characterized by increased enhancer activity at pro-inammatory genes. Enhanced binding of NF-κB and STAT proteins has been observed at promoters of cytokine genes, which may explain the overactivation of T cells and the breakdown of immune regulation.
Macrophages contribute to autoimmune inammation through cytokine production, tissue remodeling, and antigen clearance. In chronic autoimmune conditions, macrophages often adopt a persistently activated phenotype. ChIP-Seq has shown that these cells display enriched histone marks such as H3K4me3 and H3K27ac at inammatory loci. These epigenetic signatures support the concept of trained immunity, where previous exposures leave long-lasting effects on the chro­matin landscape, perpetuating inammation even after the initial trigger is removed.
ChIP-Seq analysis across key immune cell types—including CD4+ T cells, B cells, dendritic cells, and macrophages—enables the construction of a cell-type-specic atlas of regulatory altera­tions in autoimmune disease. This comprehensive approach reveals how intrinsic transcriptional programs and extrinsic environmental cues converge on chromatin to shape immune dysregulation and drive disease progression. By mapping the genomic binding patterns of TFs and proling his­tone modications, ChIP-Seq uncovers critical regulatory nodes that govern immune cell function and highlights potential therapeutic targets for restoring immune balance or selectively inhibiting pathogenic responses.
228 Bioinformatics of Autoimmune Diseases
ChIP-Seq can be applied to both bulk cell populations and adapted for single-cell applications, offering methodological exibility based on experimental goals and sample availability. In auto­immune disease research, bulk ChIP-Seq remains widely used to investigate histone landscapes and TF occupancy across immune cell populations such as CD4+ T cells, B cells, and monocytes. Depending on sample preparation, bulk analyses may focus solely on puried immune cells or may include heterogeneous mixtures of immune and non-immune cell types, particularly when tissue biopsies are used without prior enrichment. When sample input is limited or when investigating rare cell types—such as regulatory T cells—single-cell ChIP-Seq provides an indispensable alternative. It allows researchers to resolve epigenetic heterogeneity, pinpoint cell-specic regulatory elements, and detect subtle, disease-relevant chromatin changes that may be obscured in bulk datasets. As such, single-cell ChIP-Seq expands the scope of chromatin research in autoimmune disease by enabling high-resolution insights into cell-type-specic regulatory dynamics.
7.6 STUDY DESIGN
Designing a ChIP-Seq study to investigate autoimmune diseases requires a methodologically rigor­ous approach to ensure that meaningful and reproducible biological insights are obtained. The rst step in designing such a study is the clear denition of the study objective. This includes deter­mining whether the goal is to identify differential TF binding, histone modication patterns, or enhancer usage associated with disease onset, progression, or therapeutic response. The objective should be biologically justied, supported by prior literature or preliminary data, and tailored to specic immune cell types implicated in the autoimmune condition of interest—such as regula­tory T cells in type 1 diabetes or B cells in SLE. The choice of TF or histone mark (e.g., H3K27ac, H3K4me3, FOXP3, STAT1) is determined by the underlying hypothesis, such as aberrant enhancer activity or dysregulated immune tolerance.
Sample acquisition is a critical component of the study design. Samples are typically obtained from peripheral blood mononuclear cells (PBMCs), tissue biopsies (e.g., synovium in RA or brain lesions in MS), or sorted immune cell subsets. It is essential to implement rigorous protocols for sample preservation, cell sorting, and chromatin preparation to maintain consistency across sam­ples. The study population should be grouped based on disease status (e.g., healthy controls versus autoimmune patients), disease severity (e.g., mild versus severe), treatment response (e.g., respond­ers versus non-responders), or other relevant clinical variables. In longitudinal designs, paired sam­ples from the same individuals before and after treatment or at different disease stages can provide additional power to detect dynamic chromatin changes.
Before proceeding to sequencing, several key QC measures must be addressed. These include assessing chromatin fragmentation quality, antibody specicity and efciency (often validated via spike-in or known targets), and the input DNA quantity and quality. Library preparation must be consistent and reproducible, with randomized sample processing to reduce batch effects. Sequencing depth is determined based on the target: TF binding requires higher depth than broad histone modi­cations, and rare cell populations may necessitate single-cell approaches. Experimental replicates, both biological and technical, are essential for downstream statistical robustness.
In ChIP-Seq data analysis, bioinformatic preprocessing includes read alignment (e.g., using Bowtie2), peak calling (e.g., using MACS2), and normalization. For differential analysis, statistical tools such as DiffBind, csaw, or DESeq2 (adapted for ChIP-Seq count data) are commonly employed to compare binding intensities or histone modication levels between groups. Multiple testing cor­rection methods, such as the Benjamini-Hochberg procedure, are applied to control the false dis­covery rate (FDR). In cases where the design includes paired or longitudinal samples, mixed-effects models or paired statistical tests are preferred. Functional enrichment analysis using tools like Genomic Regions Enrichment of Annotations Tool (GREAT), Hypergeometric Optimization of Motif EnRichment (HOMER), or ChIPseeker can identify biological pathways and gene networks associated with differential regulatory regions. Integrating ChIP-Seq data with transcriptomic (e.g.,
229 DNA-Protein Interactions and Autoimmune Diseases
RNA-Seq) or genetic (e.g., GWAS SNPs) datasets further strengthens mechanistic interpretations. Ultimately, a well-structured study design tailored to the biological hypothesis, with appropriate statistical rigor, is fundamental to successfully leveraging ChIP-Seq for uncovering epigenetic mechanisms in autoimmune diseases.
7.7 QUALITY CONTROL
QC is a critical step in ChIP-Seq analysis to ensure that the raw sequencing data are of sufcient quality to support reliable downstream interpretation. The QC process begins with the evaluation of raw FAST Quality (FASTQ) les, which contain the sequence reads generated by the sequencer. Tools such as FastQC are commonly used to assess key quality metrics, including per-base sequence quality, germinal center (GC) content, levels of sequence duplication, and the presence of adapter contamination or overrepresented sequences. Ideally, high-quality ChIP-Seq data should exhibit uniformly high base quality scores and minimal technical artifacts. Poor-quality reads can intro­duce alignment errors and false-positive peak calls, compromising the biological validity of the results.
After initial assessment, adapter trimming and low-quality read ltering are typically performed using tools such as Trimmomatic or Cutadapt. This preprocessing step helps remove sequencing artifacts and retains only high-condence reads for genome alignment, thereby reducing back­ground noise in the dataset. Cleaned reads are then aligned to a reference genome using aligners like Bowtie2 or Burrows–Wheeler Aligner (BWA), and alignment quality is evaluated by examin­ing the proportion of uniquely mapped reads. A low unique mapping rate may indicate technical problems such as degraded DNA, nonspecic antibody binding, or an overabundance of repetitive elements in the sample. High mappability is essential for accurate peak detection and reliable inter­pretation of regulatory regions.
In addition to general mapping statistics, ChIP-Seq-specic quality metrics are assessed to determine the success of the chromatin immunoprecipitation. One key metric is the fraction of reads in peaks (FRiP), which measures the proportion of total reads that fall within condently called peak regions. A high FRiP score suggests effective enrichment of target DNA and a strong signal-to-noise ratio. Other important metrics include the normalized strand cross-correlation coef­cient (NSC) and the relative strand cross-correlation coefcient (RSC), which evaluate enrichment strength and chromatin fragmentation quality. These metrics are typically computed using tools such as phantompeakqualtools or integrated pipelines like the ENCODE ChIP-Seq pipeline.
Duplicate reads are also carefully examined during QC, as excessive duplication may result from PCR amplication bias or indicate poor library complexity. While some duplication is expected, especially in regions of strong enrichment, very high duplication rates may reduce data quality and necessitate protocol renement or deeper sequencing. Together, these quality metrics provide a comprehensive assessment of ChIP-Seq data integrity and are essential for determining whether the dataset meets the thresholds for robust, reproducible analysis. This is particularly important in autoimmune disease research, where small but biologically signicant differences in TF binding or histone modication patterns can have major implications for understanding disease mechanisms.
7.8 ChIP-Seq DATA ANALYSIS PIPELINE
The general ChIP-Seq pipeline (Figure 7.2) begins with data processing, which includes cleaning and aligning the raw sequencing reads to a reference genome. After sequencing, reads are typically stored in FASTQ format. These les are rst subjected to quality checks using tools like FastQC to identify any issues such as low base quality or adapter contamination. If necessary, trimming software such as Trim Galore or Cutadapt is applied to remove adapters and low-quality ends. Clean reads are then aligned to a reference genome (e.g., hg38 for human) using aligners such as Bowtie2 or BWA. The resulting aligned reads are stored in Binary Alignment/Map (BAM) format
230 Bioinformatics of Autoimmune Diseases
FIGURE 7.2 Illustration of ChIP-Seq pipeline workow.
and indexed to facilitate downstream analyses. Duplicate reads, which may arise from PCR ampli­cation, are often marked or removed using tools like Picard or SAMtools to avoid biases in peak identication.
Once alignment is complete, the next step is peak calling, which identies regions of the genome where reads are signicantly enriched (aligned), indicating potential protein–DNA interactions. This step differs slightly depending on whether the ChIP-Seq target is a TF or a histone modica­tion. For TFs, which bind discrete DNA sequences, peak calling identies sharp, narrow peaks, whereas histone modications often yield broader regions of enrichment. MACS2 (Model-based Analysis of ChIP-Seq) is a widely used peak-calling algorithm that models the background distri­bution of reads and detects statistically signicant peaks. It can also incorporate control datasets such as input DNA (a control sample that includes fragmented genomic DNA processed without immunoprecipitation) or immunoglobulin G (IgG) samples (negative controls using non-specic antibodies) to subtract background noise and improve specicity. The output is typically a set of genomic intervals (in BED or narrowPeak format) representing regions where the protein of interest binds or modies chromatin.
After peak calling, the next step involves annotation of the identied peaks to nearby genes and regulatory elements. Tools such as HOMER, ChIPseeker, and GREAT can annotate peaks based on proximity to gene TSSs, classify peaks into functional genomic regions (e.g., promoter, intron, inter­genic), and link them to known gene functions. This step is critical for understanding how protein– DNA interactions regulate gene expression and which biological processes they may inuence. Peaks can also be intersected with publicly available datasets, such as ENCODE or ROADMAP Epigenomics, to assess overlap with known regulatory elements, enhancers, or TF binding motifs.
The nal stage of the pipeline is functional interpretation, where researchers examine the biolog­ical meaning of the ChIP-Seq results. This typically includes pathway enrichment analysis and gene ontology (GO) analysis of the genes associated with the peaks, using tools like DAVID, Enrichr, or GSEA. Motif analysis can also be performed to identify enriched DNA sequences within the peaks, revealing potential co-binding TFs or sequence-specic regulators. In autoimmune disease studies, functional interpretation may focus on identifying immune-related pathways, cytokine signaling networks, or epigenetically regulated genes involved in immune tolerance. When integrated with RNA-Seq or ATAC-Seq data, ChIP-Seq results provide powerful insights into how epigenetic and transcriptional programs are dysregulated in disease, guiding the discovery of novel therapeutic targets and biomarkers.
231 DNA-Protein Interactions and Autoimmune Diseases
7.8.1 ACQUIRING RAW DATA FOR DEMONSTRATION
For demonstration purposes, we will use raw data from the NCBI BioProject PRJNA1020118 (National Center for Biotechnology Information (NCBI), 2023), which investigates the effects of doxycycline (Dox) treatment on human cell samples, with a particular focus on the expression of the autoimmune regulator (AIRE) gene and associated chromatin modications. The study centers on the H3K27ac histone mark, a well-known indicator of active enhancers and promoters.
The experimental design consisted of two groups:
Dox-Treated Group: Cells in this group were treated with doxycycline to induce AIRE expression. The concentration, duration of exposure, and the use of a doxycycline-inducible system were optimized to ensure effective gene induction.
Control Group: These cells were cultured under identical conditions but without doxy­cycline, serving as a baseline for comparison and allowing the assessment of AIRE­dependent changes.
Following treatment, ChIP-Seq (chromatin immunoprecipitation sequencing) was performed on both groups to map the genome-wide distribution of H3K27ac. This allowed researchers to compare chromatin accessibility and gene regulation between Dox-induced and control conditions, providing insight into the role of AIRE in gene expression and its relevance to autoimmune disease mechanisms.
The study design of the selected samples will be saved in a Comma-Separated Values (CSV) le named “meta/metadata.csv”, which contains the following information:
runID,condition SRR26147696,control SRR26147697,control SRR26147702,control SRR26147714,treated SRR26147715,treated SRR26147716,treated
Given the Sequence Read Archive (SRA) run identiers (IDs) (each ID in a line) saved in a text le named sr a _ id s.t x t within the data/raw directory, we can use the following Bash script (executed from within this directory) to download the paired-end FASTQ les from the NCBI SRA database into the same directory:
#!/bin/bash # Check if input file is provided if [ "$#" -ne 1 ]; then
echo "Usage: $0 sra_ids.txt"
exit 1 fi input_file="$1" # Check if fastq-dump is available if ! command -v fastq-dump &> /dev/null; then
echo "Error: fastq-dump is not installed or not in PATH."
exit 1 fi # Download paired-end FASTQ files while IFS= read -r sra_id; do
if [ -n "$sra_id" ]; then
echo "Downloading $sra_id..." fastq-dump --split-files --gzip "$sra_id"
232 Bioinformatics of Autoimmune Diseases
fi done < "$input_file" echo "Download completed."
Save the above script data/raw/do wnload _ fastq.sh, change to d ata/raw/, and make
the le executable:
chmod +x download_fastq.sh
Run it with your le containing SRA run IDs:
./download_fastq.sh sra_ids.txt
The FASTQ les will be downloaded into the data/raw directory, with each sample compris­ing two les corresponding to forward and reverse reads. Once the data is in place, the pipeline can be executed, beginning with QC, reference genome download and indexing, and read mapping. These initial steps are identical to those described in the RNA-Seq pipeline presented in the previ­ous chapter. To avoid redundancy, they will not be discussed again here; readers are encouraged to refer to the earlier chapter for detailed procedures. In the following sections, we focus specically on the process of peak calling, annotation, and functional interpretation.
7.8.2 PEAK CALLING
Peak calling is a critical step in ChIP-Seq data analysis, as it identies regions of the genome that are signicantly enriched for reads, representing potential binding sites of the protein of interest. After aligning the sequencing reads to a reference genome, the next goal is to detect locations where the signal from ChIP samples is notably higher than the background noise, which typically comes from a corresponding input control or mock immunoprecipitation sample. These enriched regions,
or “peaks”, suggest sites where DNA-protein interactions occurred and where the protein being
studied may be binding to the genome.
The process of peak calling involves statistical modeling to differentiate the true signal from noise. Algorithms examine the read coverage across the genome and look for local maxima (regions where the density of reads is higher than expected by chance). Most peak callers use a sliding window approach, scanning the genome to evaluate the signicance of the signal in each window compared to the local or global background. Signicance is often determined through statistical tests, such as the Poisson or negative binomial distributions, and p-values or q-values are assigned to each detected peak to represent the condence level of the enrichment.
Different peak calling tools have been developed, each with specic strengths suited to dif­ferent ChIP-Seq data types. For TFs, which typically produce sharp and narrow peaks, tools like MACS3 are commonly used. MACS3 improves peak calling accuracy by shifting reads to account for strand-specic biases and modeling the background noise using control data. For histone modi­cations that produce broader regions of enrichment, peak callers like SICER or broad peak modes of MACS3 are more appropriate, as they are tailored to detect diffuse signals over larger genomic spans.
Normalization is an essential component of peak calling to ensure that technical differences, such as sequencing depth or sample quality, do not confound the detection of biologically relevant peaks. Peak callers often normalize read counts between ChIP and control samples and may employ input subtraction to reduce background signals. Additionally, some tools use sophisticated models to account for sequence mappability, GC content, and local chromatin structure, which can inu­ence read distribution.
The quality of the identied peaks depends on various factors, including antibody specicity, sequencing depth, and biological variability. High-condence peaks are typically reproducible
233 DNA-Protein Interactions and Autoimmune Diseases
across replicates and show strong enrichment relative to input. Post-peak calling, the results are usually saved in standardized formats such as BED or narrowPeak les, which are used for down­stream analysis like motif discovery, gene annotation, and integration with other epigenomic datas­ets. Overall, peak calling is a powerful computational approach that transforms raw ChIP-Seq reads into biologically meaningful insights about protein–DNA interactions.
def call_peaks(run_id, bam_file, condition):
peaks_dir = os.path.join(out_dir, "peaks") os.makedirs(peaks_dir, exist_ok=True) peak_output = os.path.join(peaks_dir, f"{run_id}_peaks.narrowPeak") cmd = [
"macs3", "callpeak", "-t", bam_file, "-n", run_id, "--outdir", peaks_dir, "-f", "BAMPE", "-g", "hs", "--keep-dup", "all", "-q", "0.01",
"--nomodel" ] if condition == "control":
cmd += ["--nolambda"] subprocess.run(cmd, check=True) return peak_output
The ChIP-Seq data analysis pipeline is implemented in the Python script chipseq _ pipeline.
py. Peak calling is carried out by the call _ peaks function, within the program, which is
responsible for identifying genomic regions where proteins such as TFs or histone modications are signicantly enriched, using the MACS3 tool. It takes in three parameters: the sample’s run ID, the path to the aligned BAM le, and the sample’s condition, which can be either “treated” or “control.” The function begins by ensuring the output directory for peak les exists. It then constructs the full path for the resulting peak le, which will be a .narrowPeak le; this format lists the genomic coordinates of enriched peaks along with additional metrics like fold change, p-value, and q-value.
To call peaks, the function assembles a command to run macs3 callpeak, specifying the input BAM le with the -t ag and naming the output using the run ID. The option --outdir determines where to place the peak results. The -f B AM PE argument tells MACS3 that the input is paired-end BAM format, and -g hs indicates the genome size for Homo sapiens. The --keep- dup all ag instructs MACS3 to keep all duplicate reads, which can be useful in ChIP-Seq to avoid discarding biologically relevant enrichment. The - q 0.01 ag sets the q-value cutoff, con­trolling for FDR and ensuring only statistically condent peaks are retained. The --nomodel option disables automatic model building based on shift size, assuming the data is already well­fragmented and ready for peak calling.
If the sample is labeled as a control, an additional option --nolam b da is included. This sup­presses local lambda background modeling, which may be appropriate when dealing with samples where a high-quality input control is not available or not suitable for normalization. After construct­ing the command, the function uses Python’s su bpr o cess.r u n to execute it, ensuring that peak calling completes without errors. The path to the output .narrowPeak le is returned at the end of the function, making it available for downstream analysis steps such as annotation, visualization, or functional enrichment analysis.
def run_pipeline():
download_and_index_reference() metadata = read_metadata(meta_file)