Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
224 Bioinformatics of Autoimmune Diseases
FIGURE 7.1 Representative human protein–DNA complexes illustrating four major DNA-binding motifs:
(A) HSF1 with a helix–turn–helix domain (PDB: 5D5U); (B) KLF4 zinc nger domain (PDB: 6VTX);
(C) Max leucine zipper domain (PDB: 1HLO); and (D) PAX3 homeodomain (PDB: 3CMY). All structures
were rendered by the author using PyMOL and publicly available PDB les.
physiological or pathological conditions, including in autoimmune diseases, where their activity is
often dysregulated.
In the immune system, TFs play indispensable roles in the development, differentiation, and
function of immune cells. They control lineage-specic gene expression programs, govern cytokine
production, and maintain the delicate balance between immune activation and immune suppression.
Examples of such TFs include FOXP3, which is required for the function of regulatory T cells, and
STAT family members, which mediate cytokine signaling and inuence immune cell fate decisions.
TFs operate within complex regulatory networks that integrate multiple signaling pathways,
including those activated by cytokines, pattern recognition receptors, and environmental stimuli.
These networks are modulated by epigenetic modications such as histone acetylation and DNA
methylation, which can enhance or restrict TF binding to DNA. In autoimmune diseases, this regulatory balance is often disrupted. Overactive TFs may drive the sustained expression of inammatory genes, while loss of function or reduced expression can impair the formation or function of
regulatory immune cells. Because of their central roles, TFs have emerged as both biomarkers of
disease activity and potential targets for therapeutic intervention.
Table 7.1 includes a list of TFs known to play key roles in the development or regulation of
autoimmune diseases, either through genetic mutations, altered expression, or dysregulated pathways. This table was compiled by the author based on ndings from multiple peer-reviewed studies, including Smith et al. (2020), Zhang et al. (2018), and others. Together, these TFs form an
intricate network of gene regulation that maintains immune homeostasis. When their expression
is disrupted, whether through genetic mutation, epigenetic modication, or chronic inammatory
signaling, the immune system may lose tolerance to self and initiate or perpetuate autoimmune
disease. Understanding their expression proles and regulatory pathways is critical for developing
targeted therapies to restore immune balance.
7.3 HISTONE MODIFICATIONS AND CHROMATIN DYNAMICS
Histones are core structural proteins that help package DNA into chromatin, a compact and organized form that ts within the cell nucleus. Each nucleosome, the fundamental unit of chromatin
consists of approximately 147 base pairs of DNA wrapped around a histone octamer (composed of
two copies each of H2A, H2B, H3, and H4) and histone H1 (linker histone) which binds to the DNA

225 DNA-Protein Interactions and Autoimmune Diseases
TA BL E 7.1
Transcription Factors (TFs) Known to Play Roles in Autoimmune Diseases
Transcription Autoimmune Diseases
Factor Involved Function/Role
AIRE APS-1, Type 1 diabetes Central tolerance in thymus; promotes expression of tissue-specic
antigens.
FOXP3 IPEX syndrome Master regulator of Tregs (regulatory T cells); maintains peripheral
tolerance.
STAT3 Lupus, RA, MS, Crohn’s Regulates cytokine signaling and T-cell differentiation; hyperactivation
leads to inammation.
STAT1 Lupus, MS, RA IFN signaling; contributes to chronic inammation.
NF-κB (RELA,
NFKB1)
IRF5 SLE, RA, Sjögren’s Promotes expression of type I interferons and inammatory cytokines.
IRF7 SLE Major regulator of type I interferon pathway; important in antiviral and
BACH2 Type 1 diabetes, MS, Maintains T-cell balance (effector versus regulatory); suppresses
RORγt (RORC)
T-bet (TBX21) IBD, MS, RA
GATA3 Allergy, autoimmunity Promotes Th2 differentiation; imbalance with Th1/Th17 contributes to
BLIMP1 SLE, IBD Controls B- and T-cell differentiation; involved in antibody-secreting
(PRDM1) plasma cell development.
c-MAF Lupus, RA Regulates IL-10 expression in regulatory T cells and other immune cells.
IKZF1 (Ikaros) Lupus, leukemia, RA Controls lymphocyte development; mutations can lead to dysregulated
PAX5 Autoimmunity (via B cell Essential for B cell development and identity.
RUNX1 Autoimmune cytopenias Regulates hematopoiesis and T-cell lineage development.
CREM SLE Represses IL-2 expression; altered function promotes SLE-associated
EGR2/3 Lupus, T1D Induced during T-cell anergy and tolerance; dysregulation leads to
RA, SLE, psoriasis, IBD Pro-inammatory signaling; regulates cytokine gene expression.
autoimmune responses.
Crohn’s excessive inammation.
MS, RA, psoriasis Drives Th17 cell differentiation; Th17 cells are pro-inammatory.
Regulates Th1 cell differentiation; involved in IFN-γ production.
disease.
immune responses.
dysfunction)
T-cell defects.
activation.
between nucleosomes and helps compact chromatin into higher-order structures, playing a key role
in stabilizing chromatin architecture and regulating gene accessibility (see Section 3.2.6). Each histone in the nucleosome octamer possesses a histone tail, which is a exible, unstructured extension
emerging primarily from the N-terminal region (and in some cases the C-terminal region) of the
protein. These tails protrude from the nucleosome core and are enriched in positively charged amino
acids such as lysine and arginine. They serve as key platforms for post-translational modications,
including acetylation, methylation, phosphorylation, ubiquitination, and sumoylation. These chemical modications play a vital role in regulating chromatin dynamics and gene expression by altering
nucleosome stability, modulating the accessibility of DNA to transcriptional machinery, or acting
as binding sites for chromatin remodeling factors and other regulatory proteins.
These post-translational chemical modications serve as molecular signals that recruit or repel
transcriptional regulators and chromatin remodeling complexes. Histone acetylation, particularly
at lysine residues on histones H3 and H4, neutralizes the positive charge of histones and decreases
their afnity for DNA. This results in a more open chromatin conformation that is permissive to
transcription. Methylation, depending on the site and number of methyl groups, can either activate

226 Bioinformatics of Autoimmune Diseases
or repress gene transcription. For example, trimethylation of histone H3 at lysine 4 (H3K4me3) is
commonly associated with active gene promoters, whereas trimethylation at lysine 27 (H3K27me3)
correlates with gene repression.
In the context of autoimmune diseases, normal patterns of histone modication are often altered.
These changes can result in the misexpression of genes involved in immune signaling, antigen presentation, and inammatory responses (Allis & Jenuwein, 2016; Hedrich et al., 2014). In systemic
lupus erythematosus (SLE), for instance, reduced histone acetylation at the FOXP3 locus has been
linked to impaired regulatory T-cell function. In rheumatoid arthritis (RA), increased histone acetylation in synovial immune cells facilitates the transcription of genes encoding pro-inammatory
cytokines such as TNF-alpha and IL-6. Mapping these changes in histone marks offers insight into
disease mechanisms and points to new therapeutic approaches, including the use of HDAC inhibitors to restore appropriate gene expression patterns.
7.4 EPIGENETIC REGULATION OF IMMUNE IDENTITY
Epigenetic regulation refers to heritable changes in gene function that occur without altering the
DNA sequence itself. These changes include histone modications, DNA methylation, chromatin
remodeling, and the activity of non-coding RNAs. Together, they create a exible but stable regulatory framework that guides immune cell development and function.
During hematopoiesis, epigenetic mechanisms are responsible for activating or silencing lineagespecic gene programs, leading to the differentiation of pluripotent progenitor cells into mature
immune cells such as T cells, B cells, macrophages, and dendritic cells. Once differentiation is
complete, epigenetic marks continue to play a role in determining immune cell responsiveness, plasticity, and memory. For instance, the differentiation of helper T-cell subsets, including Th1, Th2,
Th17, and regulatory T cells, is orchestrated by unique epigenetic signatures that regulate access to
lineage-dening TFs and cytokine genes.
External factors such as infections, diet, stress, and exposure to environmental toxins can inuence the epigenetic landscape of immune cells. These exposures can lead to lasting epigenetic
changes, a phenomenon known as “epigenetic memory” or “trained immunity.” While this process
can enhance immune protection against recurring threats, it can also promote persistent inammation in autoimmune diseases. Long-term epigenetic remodeling may lock immune cells into a
pathogenic state, even in the absence of continued exposure to the original trigger.
Epigenetic control also plays a key role in silencing retrotransposons and non-coding elements
of the genome that could otherwise provoke immune activation. Loss of this silencing can mimic
viral infection by generating double-stranded RNA or unmethylated DNA, thereby activating innate
immune sensors and exacerbating inammation. These ndings highlight the importance of epigenetic regulation in preserving immune tolerance and preventing autoimmune disease (Sun et al.,
2015).
7.5 ChIP-Seq AND THE EPIGENOMIC LANDSCAPE OF AUTOIMMUNITY
Chromatin immunoprecipitation followed by sequencing, or ChIP-Seq, is a powerful technique
used to study protein–DNA interactions on a genome-wide scale. By using antibodies to isolate
DNA fragments bound by specic proteins, such as TFs or histone modications, ChIP-Seq allows
researchers to map regulatory elements across the genome and understand how gene expression is
controlled.
In the study of autoimmune diseases, ChIP-Seq has been instrumental in identifying changes
in the binding patterns of TFs and the distribution of histone modications. For example, in SLE
and RA, ChIP-Seq has revealed abnormal binding of TFs such as STAT1, IRF5, and NF-κB at key
immune genes. These changes often correspond with altered gene expression proles and reect the
rewiring of regulatory circuits that promote chronic inammation and loss of immune tolerance.

227 DNA-Protein Interactions and Autoimmune Diseases
ChIP-Seq also allows for the identication of active and repressed chromatin regions through
histone marks such as H3K27ac and H3K4me3. These marks dene enhancers and promoters that
are engaged in transcriptional activity. In autoimmune diseases, the locations and intensities of
these marks are often shifted, revealing previously unrecognized regulatory elements, sometimes
referred to as latent or cryptic enhancers that become active in disease states.
When combined with other genomic technologies, such as RNA sequencing and ATAC-Seq
, ChIP-Seq provides a multidimensional view of gene regulation. These integrative approaches
help pinpoint causal relationships between epigenomic changes and transcriptional outcomes.
Furthermore, many genetic risk variants identied through genome-wide association studies are
found in non-coding regions that coincide with ChIP-Seq peaks. This overlap suggests that disease susceptibility is often mediated through changes in gene regulation rather than protein-coding
mutations.
Overall, ChIP-Seq offers valuable insights into the molecular mechanisms underlying immune
dysregulation in autoimmunity (Barski et al., 2007; Scharer et al., 2019).
By employing ChIP-Seq across a variety of immune cell types, researchers have signicantly
advanced our understanding of the epigenetic mechanisms driving autoimmune diseases. These
studies typically focus on immune cells that play pivotal roles in immune regulation and disease
pathogenesis, using them as the primary sources for DNA extraction in ChIP-Seq experiments. The
key immune cell types examined are outlined below.
CD4+ T cells are a critical component of adaptive immunity and differentiate into distinct functional subsets, including Th1, Th17, and regulatory T cells (Tregs). Each subset is dened by unique
transcriptional and epigenetic signatures. ChIP-Seq has been instrumental in identifying histone
modications such as H3K27ac and H3K4me3 at key regulatory loci in these subsets. Furthermore,
TFs such as FOXP3, T-bet, and RORγt have been mapped using ChIP-Seq to reveal alterations in
binding patterns associated with autoimmune diseases like multiple sclerosis (MS) and RA. These
studies have highlighted changes in enhancer activity and silencing of critical regulatory regions,
offering insight into cytokine dysregulation and loss of immune tolerance in affected individuals.
B cells, responsible for producing antibodies and presenting antigens, also exhibit epigenomic
alterations in autoimmunity. In SLE, ChIP-Seq has uncovered disease-specic super-enhancers that
drive expression of genes involved in autoantibody production and hyperactive B cell states. TFs
such as IRF4 and PAX5 display aberrant binding in these cells, reecting rewired regulatory networks that support a pathogenic phenotype.
Dendritic cells serve as key antigen-presenting cells and gatekeepers of peripheral tolerance.
When ChIP-Seq is applied to dendritic cells from autoimmune models or patient samples, it reveals
transcriptional shifts characterized by increased enhancer activity at pro-inammatory genes.
Enhanced binding of NF-κB and STAT proteins has been observed at promoters of cytokine genes,
which may explain the overactivation of T cells and the breakdown of immune regulation.
Macrophages contribute to autoimmune inammation through cytokine production, tissue
remodeling, and antigen clearance. In chronic autoimmune conditions, macrophages often adopt
a persistently activated phenotype. ChIP-Seq has shown that these cells display enriched histone
marks such as H3K4me3 and H3K27ac at inammatory loci. These epigenetic signatures support
the concept of trained immunity, where previous exposures leave long-lasting effects on the chromatin landscape, perpetuating inammation even after the initial trigger is removed.
ChIP-Seq analysis across key immune cell types—including CD4+ T cells, B cells, dendritic
cells, and macrophages—enables the construction of a cell-type-specic atlas of regulatory alterations in autoimmune disease. This comprehensive approach reveals how intrinsic transcriptional
programs and extrinsic environmental cues converge on chromatin to shape immune dysregulation
and drive disease progression. By mapping the genomic binding patterns of TFs and proling histone modications, ChIP-Seq uncovers critical regulatory nodes that govern immune cell function
and highlights potential therapeutic targets for restoring immune balance or selectively inhibiting
pathogenic responses.

228 Bioinformatics of Autoimmune Diseases
ChIP-Seq can be applied to both bulk cell populations and adapted for single-cell applications,
offering methodological exibility based on experimental goals and sample availability. In autoimmune disease research, bulk ChIP-Seq remains widely used to investigate histone landscapes
and TF occupancy across immune cell populations such as CD4+ T cells, B cells, and monocytes.
Depending on sample preparation, bulk analyses may focus solely on puried immune cells or may
include heterogeneous mixtures of immune and non-immune cell types, particularly when tissue
biopsies are used without prior enrichment. When sample input is limited or when investigating rare
cell types—such as regulatory T cells—single-cell ChIP-Seq provides an indispensable alternative.
It allows researchers to resolve epigenetic heterogeneity, pinpoint cell-specic regulatory elements,
and detect subtle, disease-relevant chromatin changes that may be obscured in bulk datasets. As
such, single-cell ChIP-Seq expands the scope of chromatin research in autoimmune disease by
enabling high-resolution insights into cell-type-specic regulatory dynamics.
7.6 STUDY DESIGN
Designing a ChIP-Seq study to investigate autoimmune diseases requires a methodologically rigorous approach to ensure that meaningful and reproducible biological insights are obtained. The rst
step in designing such a study is the clear denition of the study objective. This includes determining whether the goal is to identify differential TF binding, histone modication patterns, or
enhancer usage associated with disease onset, progression, or therapeutic response. The objective
should be biologically justied, supported by prior literature or preliminary data, and tailored to
specic immune cell types implicated in the autoimmune condition of interest—such as regulatory T cells in type 1 diabetes or B cells in SLE. The choice of TF or histone mark (e.g., H3K27ac,
H3K4me3, FOXP3, STAT1) is determined by the underlying hypothesis, such as aberrant enhancer
activity or dysregulated immune tolerance.
Sample acquisition is a critical component of the study design. Samples are typically obtained
from peripheral blood mononuclear cells (PBMCs), tissue biopsies (e.g., synovium in RA or brain
lesions in MS), or sorted immune cell subsets. It is essential to implement rigorous protocols for
sample preservation, cell sorting, and chromatin preparation to maintain consistency across samples. The study population should be grouped based on disease status (e.g., healthy controls versus
autoimmune patients), disease severity (e.g., mild versus severe), treatment response (e.g., responders versus non-responders), or other relevant clinical variables. In longitudinal designs, paired samples from the same individuals before and after treatment or at different disease stages can provide
additional power to detect dynamic chromatin changes.
Before proceeding to sequencing, several key QC measures must be addressed. These include
assessing chromatin fragmentation quality, antibody specicity and efciency (often validated via
spike-in or known targets), and the input DNA quantity and quality. Library preparation must be
consistent and reproducible, with randomized sample processing to reduce batch effects. Sequencing
depth is determined based on the target: TF binding requires higher depth than broad histone modications, and rare cell populations may necessitate single-cell approaches. Experimental replicates,
both biological and technical, are essential for downstream statistical robustness.
In ChIP-Seq data analysis, bioinformatic preprocessing includes read alignment (e.g., using
Bowtie2), peak calling (e.g., using MACS2), and normalization. For differential analysis, statistical
tools such as DiffBind, csaw, or DESeq2 (adapted for ChIP-Seq count data) are commonly employed
to compare binding intensities or histone modication levels between groups. Multiple testing correction methods, such as the Benjamini-Hochberg procedure, are applied to control the false discovery rate (FDR). In cases where the design includes paired or longitudinal samples, mixed-effects
models or paired statistical tests are preferred. Functional enrichment analysis using tools like
Genomic Regions Enrichment of Annotations Tool (GREAT), Hypergeometric Optimization of
Motif EnRichment (HOMER), or ChIPseeker can identify biological pathways and gene networks
associated with differential regulatory regions. Integrating ChIP-Seq data with transcriptomic (e.g.,

229 DNA-Protein Interactions and Autoimmune Diseases
RNA-Seq) or genetic (e.g., GWAS SNPs) datasets further strengthens mechanistic interpretations.
Ultimately, a well-structured study design tailored to the biological hypothesis, with appropriate
statistical rigor, is fundamental to successfully leveraging ChIP-Seq for uncovering epigenetic
mechanisms in autoimmune diseases.
7.7 QUALITY CONTROL
QC is a critical step in ChIP-Seq analysis to ensure that the raw sequencing data are of sufcient
quality to support reliable downstream interpretation. The QC process begins with the evaluation
of raw FAST Quality (FASTQ) les, which contain the sequence reads generated by the sequencer.
Tools such as FastQC are commonly used to assess key quality metrics, including per-base sequence
quality, germinal center (GC) content, levels of sequence duplication, and the presence of adapter
contamination or overrepresented sequences. Ideally, high-quality ChIP-Seq data should exhibit
uniformly high base quality scores and minimal technical artifacts. Poor-quality reads can introduce alignment errors and false-positive peak calls, compromising the biological validity of the
results.
After initial assessment, adapter trimming and low-quality read ltering are typically performed
using tools such as Trimmomatic or Cutadapt. This preprocessing step helps remove sequencing
artifacts and retains only high-condence reads for genome alignment, thereby reducing background noise in the dataset. Cleaned reads are then aligned to a reference genome using aligners
like Bowtie2 or Burrows–Wheeler Aligner (BWA), and alignment quality is evaluated by examining the proportion of uniquely mapped reads. A low unique mapping rate may indicate technical
problems such as degraded DNA, nonspecic antibody binding, or an overabundance of repetitive
elements in the sample. High mappability is essential for accurate peak detection and reliable interpretation of regulatory regions.
In addition to general mapping statistics, ChIP-Seq-specic quality metrics are assessed to
determine the success of the chromatin immunoprecipitation. One key metric is the fraction of
reads in peaks (FRiP), which measures the proportion of total reads that fall within condently
called peak regions. A high FRiP score suggests effective enrichment of target DNA and a strong
signal-to-noise ratio. Other important metrics include the normalized strand cross-correlation coefcient (NSC) and the relative strand cross-correlation coefcient (RSC), which evaluate enrichment
strength and chromatin fragmentation quality. These metrics are typically computed using tools
such as phantompeakqualtools or integrated pipelines like the ENCODE ChIP-Seq pipeline.
Duplicate reads are also carefully examined during QC, as excessive duplication may result from
PCR amplication bias or indicate poor library complexity. While some duplication is expected,
especially in regions of strong enrichment, very high duplication rates may reduce data quality
and necessitate protocol renement or deeper sequencing. Together, these quality metrics provide
a comprehensive assessment of ChIP-Seq data integrity and are essential for determining whether
the dataset meets the thresholds for robust, reproducible analysis. This is particularly important in
autoimmune disease research, where small but biologically signicant differences in TF binding or
histone modication patterns can have major implications for understanding disease mechanisms.
7.8 ChIP-Seq DATA ANALYSIS PIPELINE
The general ChIP-Seq pipeline (Figure 7.2) begins with data processing, which includes cleaning
and aligning the raw sequencing reads to a reference genome. After sequencing, reads are typically
stored in FASTQ format. These les are rst subjected to quality checks using tools like FastQC
to identify any issues such as low base quality or adapter contamination. If necessary, trimming
software such as Trim Galore or Cutadapt is applied to remove adapters and low-quality ends.
Clean reads are then aligned to a reference genome (e.g., hg38 for human) using aligners such as
Bowtie2 or BWA. The resulting aligned reads are stored in Binary Alignment/Map (BAM) format

230 Bioinformatics of Autoimmune Diseases
FIGURE 7.2 Illustration of ChIP-Seq pipeline workow.
and indexed to facilitate downstream analyses. Duplicate reads, which may arise from PCR amplication, are often marked or removed using tools like Picard or SAMtools to avoid biases in peak
identication.
Once alignment is complete, the next step is peak calling, which identies regions of the genome
where reads are signicantly enriched (aligned), indicating potential protein–DNA interactions.
This step differs slightly depending on whether the ChIP-Seq target is a TF or a histone modication. For TFs, which bind discrete DNA sequences, peak calling identies sharp, narrow peaks,
whereas histone modications often yield broader regions of enrichment. MACS2 (Model-based
Analysis of ChIP-Seq) is a widely used peak-calling algorithm that models the background distribution of reads and detects statistically signicant peaks. It can also incorporate control datasets
such as input DNA (a control sample that includes fragmented genomic DNA processed without
immunoprecipitation) or immunoglobulin G (IgG) samples (negative controls using non-specic
antibodies) to subtract background noise and improve specicity. The output is typically a set of
genomic intervals (in BED or narrowPeak format) representing regions where the protein of interest
binds or modies chromatin.
After peak calling, the next step involves annotation of the identied peaks to nearby genes and
regulatory elements. Tools such as HOMER, ChIPseeker, and GREAT can annotate peaks based on
proximity to gene TSSs, classify peaks into functional genomic regions (e.g., promoter, intron, intergenic), and link them to known gene functions. This step is critical for understanding how protein–
DNA interactions regulate gene expression and which biological processes they may inuence.
Peaks can also be intersected with publicly available datasets, such as ENCODE or ROADMAP
Epigenomics, to assess overlap with known regulatory elements, enhancers, or TF binding motifs.
The nal stage of the pipeline is functional interpretation, where researchers examine the biological meaning of the ChIP-Seq results. This typically includes pathway enrichment analysis and gene
ontology (GO) analysis of the genes associated with the peaks, using tools like DAVID, Enrichr, or
GSEA. Motif analysis can also be performed to identify enriched DNA sequences within the peaks,
revealing potential co-binding TFs or sequence-specic regulators. In autoimmune disease studies,
functional interpretation may focus on identifying immune-related pathways, cytokine signaling
networks, or epigenetically regulated genes involved in immune tolerance. When integrated with
RNA-Seq or ATAC-Seq data, ChIP-Seq results provide powerful insights into how epigenetic and
transcriptional programs are dysregulated in disease, guiding the discovery of novel therapeutic
targets and biomarkers.

231 DNA-Protein Interactions and Autoimmune Diseases
7.8.1 ACQUIRING RAW DATA FOR DEMONSTRATION
For demonstration purposes, we will use raw data from the NCBI BioProject PRJNA1020118
(National Center for Biotechnology Information (NCBI), 2023), which investigates the effects of
doxycycline (Dox) treatment on human cell samples, with a particular focus on the expression of the
autoimmune regulator (AIRE) gene and associated chromatin modications. The study centers on
the H3K27ac histone mark, a well-known indicator of active enhancers and promoters.
The experimental design consisted of two groups:
• Dox-Treated Group: Cells in this group were treated with doxycycline to induce AIRE
expression. The concentration, duration of exposure, and the use of a doxycycline-inducible
system were optimized to ensure effective gene induction.
• Control Group: These cells were cultured under identical conditions but without doxycycline, serving as a baseline for comparison and allowing the assessment of AIREdependent changes.
Following treatment, ChIP-Seq (chromatin immunoprecipitation sequencing) was performed
on both groups to map the genome-wide distribution of H3K27ac. This allowed researchers to
compare chromatin accessibility and gene regulation between Dox-induced and control conditions,
providing insight into the role of AIRE in gene expression and its relevance to autoimmune disease
mechanisms.
The study design of the selected samples will be saved in a Comma-Separated Values (CSV) le
named “meta/metadata.csv”, which contains the following information:
runID,condition
SRR26147696,control
SRR26147697,control
SRR26147702,control
SRR26147714,treated
SRR26147715,treated
SRR26147716,treated
Given the Sequence Read Archive (SRA) run identiers (IDs) (each ID in a line) saved in a text
le named sr a _ id s.t x t within the data/raw directory, we can use the following Bash script
(executed from within this directory) to download the paired-end FASTQ les from the NCBI SRA
database into the same directory:
#!/bin/bash
# Check if input file is provided
if [ "$#" -ne 1 ]; then
echo "Usage: $0 sra_ids.txt"
exit 1
fi
input_file="$1"
# Check if fastq-dump is available
if ! command -v fastq-dump &> /dev/null; then
echo "Error: fastq-dump is not installed or not in PATH."
exit 1
fi
# Download paired-end FASTQ files
while IFS= read -r sra_id; do
if [ -n "$sra_id" ]; then
echo "Downloading $sra_id..."
fastq-dump --split-files --gzip "$sra_id"

232 Bioinformatics of Autoimmune Diseases
fi
done < "$input_file"
echo "Download completed."
Save the above script data/raw/do wnload _ fastq.sh, change to d ata/raw/, and make
the le executable:
chmod +x download_fastq.sh
Run it with your le containing SRA run IDs:
./download_fastq.sh sra_ids.txt
The FASTQ les will be downloaded into the data/raw directory, with each sample comprising two les corresponding to forward and reverse reads. Once the data is in place, the pipeline can
be executed, beginning with QC, reference genome download and indexing, and read mapping.
These initial steps are identical to those described in the RNA-Seq pipeline presented in the previous chapter. To avoid redundancy, they will not be discussed again here; readers are encouraged to
refer to the earlier chapter for detailed procedures. In the following sections, we focus specically
on the process of peak calling, annotation, and functional interpretation.
7.8.2 PEAK CALLING
Peak calling is a critical step in ChIP-Seq data analysis, as it identies regions of the genome that
are signicantly enriched for reads, representing potential binding sites of the protein of interest.
After aligning the sequencing reads to a reference genome, the next goal is to detect locations where
the signal from ChIP samples is notably higher than the background noise, which typically comes
from a corresponding input control or mock immunoprecipitation sample. These enriched regions,
or “peaks”, suggest sites where DNA-protein interactions occurred and where the protein being
studied may be binding to the genome.
The process of peak calling involves statistical modeling to differentiate the true signal from
noise. Algorithms examine the read coverage across the genome and look for local maxima (regions
where the density of reads is higher than expected by chance). Most peak callers use a sliding
window approach, scanning the genome to evaluate the signicance of the signal in each window
compared to the local or global background. Signicance is often determined through statistical
tests, such as the Poisson or negative binomial distributions, and p-values or q-values are assigned
to each detected peak to represent the condence level of the enrichment.
Different peak calling tools have been developed, each with specic strengths suited to different ChIP-Seq data types. For TFs, which typically produce sharp and narrow peaks, tools like
MACS3 are commonly used. MACS3 improves peak calling accuracy by shifting reads to account
for strand-specic biases and modeling the background noise using control data. For histone modications that produce broader regions of enrichment, peak callers like SICER or broad peak modes
of MACS3 are more appropriate, as they are tailored to detect diffuse signals over larger genomic
spans.
Normalization is an essential component of peak calling to ensure that technical differences,
such as sequencing depth or sample quality, do not confound the detection of biologically relevant
peaks. Peak callers often normalize read counts between ChIP and control samples and may employ
input subtraction to reduce background signals. Additionally, some tools use sophisticated models
to account for sequence mappability, GC content, and local chromatin structure, which can inuence read distribution.
The quality of the identied peaks depends on various factors, including antibody specicity,
sequencing depth, and biological variability. High-condence peaks are typically reproducible

233 DNA-Protein Interactions and Autoimmune Diseases
across replicates and show strong enrichment relative to input. Post-peak calling, the results are
usually saved in standardized formats such as BED or narrowPeak les, which are used for downstream analysis like motif discovery, gene annotation, and integration with other epigenomic datasets. Overall, peak calling is a powerful computational approach that transforms raw ChIP-Seq reads
into biologically meaningful insights about protein–DNA interactions.
def call_peaks(run_id, bam_file, condition):
peaks_dir = os.path.join(out_dir, "peaks")
os.makedirs(peaks_dir, exist_ok=True)
peak_output = os.path.join(peaks_dir, f"{run_id}_peaks.narrowPeak")
cmd = [
"macs3", "callpeak",
"-t", bam_file,
"-n", run_id,
"--outdir", peaks_dir,
"-f", "BAMPE",
"-g", "hs",
"--keep-dup", "all",
"-q", "0.01",
"--nomodel"
]
if condition == "control":
cmd += ["--nolambda"]
subprocess.run(cmd, check=True)
return peak_output
The ChIP-Seq data analysis pipeline is implemented in the Python script chipseq _ pipeline.
py. Peak calling is carried out by the call _ peaks function, within the program, which is
responsible for identifying genomic regions where proteins such as TFs or histone modications are
signicantly enriched, using the MACS3 tool. It takes in three parameters: the sample’s run ID, the
path to the aligned BAM le, and the sample’s condition, which can be either “treated” or “control.”
The function begins by ensuring the output directory for peak les exists. It then constructs the full
path for the resulting peak le, which will be a .narrowPeak le; this format lists the genomic
coordinates of enriched peaks along with additional metrics like fold change, p-value, and q-value.
To call peaks, the function assembles a command to run macs3 callpeak, specifying the
input BAM le with the -t ag and naming the output using the run ID. The option --outdir
determines where to place the peak results. The -f B AM PE argument tells MACS3 that the input
is paired-end BAM format, and -g hs indicates the genome size for Homo sapiens. The --keep-
dup all ag instructs MACS3 to keep all duplicate reads, which can be useful in ChIP-Seq to
avoid discarding biologically relevant enrichment. The - q 0.01 ag sets the q-value cutoff, controlling for FDR and ensuring only statistically condent peaks are retained. The --nomodel
option disables automatic model building based on shift size, assuming the data is already wellfragmented and ready for peak calling.
If the sample is labeled as a control, an additional option --nolam b da is included. This suppresses local lambda background modeling, which may be appropriate when dealing with samples
where a high-quality input control is not available or not suitable for normalization. After constructing the command, the function uses Python’s su bpr o cess.r u n to execute it, ensuring that peak
calling completes without errors. The path to the output .narrowPeak le is returned at the end
of the function, making it available for downstream analysis steps such as annotation, visualization,
or functional enrichment analysis.
def run_pipeline():
download_and_index_reference()
metadata = read_metadata(meta_file)
Соседние файлы в папке Библиотека им академика М.И. Перельмана
