Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
254 Bioinformatics of Autoimmune Diseases
From a methodological perspective, the chapter meticulously outlines the experimental design
for ChIP-Seq in autoimmune disease research. It includes considerations for sample selection, antibody specicity, control strategies, and sequencing depth. A step-by-step workow is described,
starting from raw data QC and alignment, through peak calling using tools like MACS3, to downstream annotation and functional interpretation. Key statistical metrics such as signalValue, p-value,
q-value, and peak score are clearly explained, offering readers the conceptual tools to assess the
reliability and biological relevance of detected binding events.
Integration with other datasets such as RNA-Seq or ATAC-Seq is discussed as a crucial next
step, enabling the construction of multi-layered regulatory networks that reect the complexity of
immune regulation. The chapter also presents a working example using publicly available data to
demonstrate practical application, including a Python pipeline for peak calling and annotation, as
well as the use of HOMER for genomic feature association and motif discovery.
Altogether, this chapter serves as a technically detailed yet conceptually rich resource that
bridges molecular biology, immunology, and bioinformatics. It situates ChIP-Seq not just as a laboratory technique, but as a window into the epigenomic mechanisms that drive immune identity and
misregulation in disease. The clarity of exposition, combined with practical guidance and relevance
to human pathology, makes this a valuable reference for both experimentalists and computational
researchers working in epigenetics and autoimmune biology.
BIBLIOGRAPHY
Allis, C. D., & Jenuwein, T. (2016). The molecular hallmarks of epigenetic control. Nature Reviews Genetics,
17(8), 487–500. https://doi.org/10.1038/nrg.2016.59
Barski, A., et al. (2007). High-resolution proling of histone methylations in the human genome. Cell, 129(4),
82 3– 837. https://doi.org/10.1016/j.cell.2007.05.009
Ismail, H. D. (2022). Bioinformatics: A practical guide to NCBI databases and sequence alignments (1st ed.).
Chapman and Hall/CRC. https://doi.org/10.1201/9781003226611
Ismail, H. D. (2023). Bioinformatics: A practical guide to next generation sequencing data analysis (1st ed.).
Chapman and Hall/CRC. https://doi.org/10.1201/9781003355205
Luger, K., Mäder, A. W., Richmond, R. K., Sargent, D. F., & Richmond, T. J. (1997). Crystal structure of the
nucleosome core particle at 2.8 Å resolution. Nature, 389(6648), 251–260. https://doi.org/10.1038/38444
National Center for Biotechnology Information (NCBI). (2023). BioProject PRJNA1020118: H3K27ac ChIP-
Seq after doxycycline induction of AIRE [Data set]. NCBI Sequence Read Archive. https://www.ncbi.
nlm.nih.gov/bioproject/PRJNA1020118
Ptashne, M. (2005). Regulation of transcription: From lambda to eukaryotes. Trends in Biochemical Sciences,
30(6), 275–279. https://doi.org/10.1016/j.tibs.2005.04.001
Scharer, C. D., et al. (2019). Epigenetic programming underpins B cell dysfunction in human SLE. Nature
Im munology, 20(8), 1071–1082. https://doi.org/10.1038/s41590-019-0419-9
Sun, Y., Li, Y., Luo, D., Liao, D. J., & Wang, X. (2015). Epigenetic regulation of autoimmunity by DNA meth-
ylation. Journal of Autoimmunity, 64, 1–12. https://doi.org/10.1016/j.jaut.2015.07.002
Zhang, Q., Vignali, D. A. A., & Chang, M. (2018). Epigenetic regulation of autoimmune disease-associated
gene expression. Trends in Immunology, 39(7), 573–588. https://doi.org/10.1016/j.it.2018.04.004

ATAC-Seq for
8
Autoimmune Diseases
8.1 GENE ACCESSIBILITY
In the previous chapter, we discussed ChIP-Seq and its utility in detecting transcription factor
(TF) binding sites and histone modications. We also examined the structure of chromatin and its
critical role in gene regulation, particularly through its inuence on DNA accessibility. Chromatin
can exist in a more open or closed state, and this dynamic structure determines whether certain
genes, often otherwise silent, can become accessible for transcription and thereby inuence gene
expression.
In recent years, there has been growing recognition that alterations in chromatin accessibility
(changes that affect how tightly DNA is packaged within the nucleus) play a central role in shaping
immune cell function and immune dysregulation. Chromatin architecture dictates which regions of
the genome are open to TF binding and, consequently, which genes are activated or repressed. By
studying these patterns of accessibility, researchers can gain insights into the regulatory networks
that orchestrate immune responses and identify epigenetic disturbances associated with autoimmune diseases.
This chapter explores the central role of chromatin accessibility in autoimmune pathogenesis
through the lens of a powerful epigenomic tool: ATAC-Seq (Assay for Transposase-Accessible
Chromatin using sequencing). ATAC-Seq enables high-resolution, genome-wide proling of
open chromatin regions by utilizing a hyperactive transposase enzyme that preferentially inserts
sequencing adapters into accessible DNA. This technique has become particularly valuable in
immunological studies, where chromatin remodeling plays a key role in processes such as T cell
activation, B-cell differentiation, and the establishment of immunological memory. By applying
ATAC-Seq to immune cells from patients with autoimmune disorders, researchers can identify
disease-specic regulatory elements and TF binding motifs that may drive aberrant immune
responses.
When compared to ChIP-Seq, a widely used method for detecting specic protein–DNA interactions such as TF occupancy or histone modications, ATAC-Seq offers several distinct advantages.
First, ATAC-Seq is less labor-intensive and requires signicantly fewer cells, making it well-suited
for studies involving rare immune cell subsets or limited patient material. Unlike ChIP-Seq, which
depends on antibodies against known proteins and thus focuses on predened targets, ATAC-Seq
is an unbiased technique that reveals all accessible regions of the genome regardless of the proteins
involved. This provides a broader, more global view of the chromatin landscape. Moreover, ATACSeq can be performed on frozen samples and is compatible with single-cell analysis, allowing for
the resolution of cell type–specic regulatory patterns within heterogeneous populations (capabilities that are difcult to achieve with ChIP-Seq).
In addition to introducing the fundamental principles of ATAC-Seq, this chapter highlights its
applications in immunology and autoimmunity. It explores how ATAC-Seq can be used to map
enhancer elements, identify shifts in chromatin accessibility during immune activation, and uncover
regulatory circuits involved in disease. The chapter also provides a detailed overview of computational strategies for analyzing ATAC-Seq data, including peak calling, motif discovery, integration
with RNA-Seq and ChIP-Seq datasets, and functional annotation of regulatory regions. Together,
these approaches position ATAC-Seq as a transformative tool for decoding the non-coding genome
and illuminating the regulatory underpinnings of autoimmune disease.
255 D OI: 10.1201/ 978 10 03 68 5 432- 8

256 Bioinformatics of Autoimmune Diseases
8.2 CHROMATIN ACCESSIBILITY AND IMMUNE REGULATION
The human genome, with its vast 3 billion base pairs, is intricately folded within the microscopic
boundaries of the cell nucleus. This packaging is accomplished through the organization of DNA
into chromatin, a complex structure in which DNA is wound around histone proteins to form nucleosomes (Luger et al., 1997). Chromatin is not a passive framework; it is a dynamic regulatory system
that governs the accessibility of genetic information. The degree to which chromatin is open or
closed (referred to as chromatin accessibility) determines whether genes can be actively transcribed
or remain silent (Klemm et al., 2019). This accessibility serves as a foundational mechanism of epi-
genetic regulation, particularly critical in the immune system, where precise and context-dependent
gene expression denes the function and fate of immune cells (Lara-Astiaso et al., 2014).
Chromatin exists in two primary states (see Section 3.2.6): tightly packed heterochromatin,
which is largely inaccessible to transcriptional machinery and associated with gene silencing,
and loosely arranged euchromatin, which facilitates active gene expression (Allis & Jenuwein,
2016). This regulatory balance is modulated by a diverse array of epigenetic modications that
alter the chemical and structural conguration of chromatin without altering the underlying DNA
sequence. Histone modications such as acetylation and methylation are central to this regulation. Histone acetylation, catalyzed by histone acetyltransferases (HATs), is generally associated with transcriptional activation, as it reduces the electrostatic interaction between DNA and
histones, thereby increasing accessibility (Bannister & Kouzarides, 2011). In contrast, histone
methylation can either activate or repress gene expression depending on the specic lysine residue modied and the degree of methylation. For instance, trimethylation of H3K4 is linked to
active transcription, whereas trimethylation of H3K27 is typically associated with gene repression (Barski et al., 2007).
Chromatin accessibility specically refers to the physical openness of DNA regions that allows
TFs and other regulatory proteins to access and bind their target sequences. This openness is often
a prerequisite for transcriptional activation and is frequently observed at promoters, enhancers,
and other cis-regulatory elements (Buenrostro et al., 2013). High-throughput sequencing technolo-
gies such as ATAC-Seq have revolutionized the ability to map open chromatin genome-wide with
high resolution. These maps have revealed the regulatory architecture of distinct cell types and
states, proving especially informative in deciphering the differentiation and functional response of
immune cells (Corces et al., 2016).
In immune cells, chromatin accessibility is nely tuned to orchestrate lineage-specic gene
expression programs. The differentiation of hematopoietic progenitor cells into various immune
subsets (including T cells, B cells, macrophages, and dendritic cells) requires coordinated changes
in chromatin organization. These changes involve both nucleosome repositioning and the deposition
of specic histone modications. For example, cell type–specic enhancers marked by H3K4me1
and H3K27ac become accessible in a manner that drives the transcription of immunoregulatory
genes essential for cellular identity and immune function (Heinz et al., 2010). These epigenetic
mechanisms ensure that each immune cell type activates the appropriate gene networks only under
the right circumstances, preserving the specicity and timing of immune responses.
In autoimmune diseases, however, the precise regulation of chromatin accessibility can become
aberrant. Disruption in chromatin landscapes may lead to the inappropriate activation of genes
involved in inammation, cytokine signaling, and immune cell proliferation (Scharer et al., 2019).
For instance, in diseases such as systemic lupus erythematosus (SLE) and rheumatoid arthritis
(RA), open chromatin regions have been identied at the regulatory elements of genes like IL6,
Interferon gamma (IFNG), and tumor necrosis factor (TNF), contributing to chronic and pathogenic inammation. Moreover, defects in regulatory T cells, such as diminished accessibility at
the FOXP3 locus, can impair their suppressive function and compromise peripheral tolerance,
allowing autoreactive T cells to escape regulation and drive autoimmunity (Floess et al., 2007).
Environmental factors (including infections, microbiome composition, and dietary exposures) can

257 ATAC-Seq for Autoimmune Diseases
also induce epigenetic modications that alter chromatin architecture, increase exposure to selfantigens, and ultimately contribute to autoimmune pathogenesis (Bach, 2018).
These dynamic changes in chromatin are not merely downstream consequences of autoimmune
pathology; rather, they are often integral to the onset and progression of disease. As such, understanding chromatin remodeling in immune cells provides new avenues for diagnosing, monitoring,
and treating autoimmune disorders. Epigenetic therapeutics aimed at modifying chromatin accessibility or reversing aberrant enhancer activation are under exploration as promising strategies to
restore immune tolerance and suppress pathological immune responses (Peters et al., 2015). By
reshaping the chromatin landscape, it may be possible to reprogram immune cells toward a nonpathogenic state, offering transformative potential in the treatment of complex immune-mediated
diseases.
8.3 ATAC-Seq IN IMMUNE CELLS
Within the intricate orchestration of immune responses, chromatin accessibility plays a decisive
role in shaping the transcriptional fate of immune cells. ATAC-Seq has emerged as a powerful
method for proling genome-wide chromatin accessibility at high resolution, revealing the dynamic
remodeling of the epigenetic landscape during immune cell activation, differentiation, and dysfunction (Buenrostro et al., 2013). By applying ATAC-Seq to diverse immune cell lineages (including
Tcells, B cells, and innate immune cells), researchers have begun to decipher the regulatory logic
that governs immune homeostasis and the disruptions that underlie autoimmune pathology.
In T lymphocytes, chromatin accessibility undergoes profound alterations as cells transition from
naïve to activated and effector states. Naïve CD4+ T cells, upon encountering antigen, differentiate
into lineage-committed subsets such as Th1 and Th17, each dened by distinct transcriptional and
epigenomic signatures. ATAC-Seq has shown that this functional divergence is marked by the selective opening of enhancer regions linked to lineage-dening TFs. For example, Th1 cells exhibit
increased accessibility at the Ifng and Tbx21 loci, supporting the expression of interferon-gamma
and the TF T-bet, respectively, whereas Th17 cells demonstrate accessible chromatin at Il17a, Il17f,
and Rorc, consistent with their pro-inammatory identity (Shih et al., 2014; Yosef et al., 2013).
These chromatin landscapes precede full transcriptional commitment, indicating that epigenomic
remodeling acts as a priming mechanism for lineage determination.
Regulatory T cells (Tregs), critical for maintaining immune tolerance, also display a distinctive chromatin architecture. The TF FoxP3 orchestrates Treg development and function, in part by
promoting accessibility at enhancers of genes such as Ctla4, Tnfrsf18 (GITR), and Ikzf2 (Samstein
et al., 2012). However, in autoimmune conditions, Tregs may exhibit altered chromatin states that
compromise their suppressive capacity or promote phenotypic plasticity toward effector lineages.
This epigenetic instability is particularly pronounced in inamed tissues, suggesting that a hostile
microenvironment can subvert Treg identity and function by reshaping their chromatin landscape
(Schmidl et al., 2009).
B cells similarly undergo extensive epigenomic remodeling during maturation and antigen-driven
responses. Throughout the germinal center (GC) reaction, B cells proliferate, undergo somatic
hypermutation, and initiate class switch recombination, processes regulated by TFs and chromatin
dynamics. ATAC-Seq analyses reveal that memory B cells and antibody-secreting plasma cells possess distinct accessibility patterns. Plasma cells, for instance, exhibit reduced accessibility at loci
governing antigen presentation and B-cell receptor signaling, while showing increased accessibility
near genes involved in antibody secretion such as Prdm1 (BLIMP-1) and Xbp1 (Zhou et al., 2015).
In systemic autoimmune diseases such as lupus, B cells display aberrant enhancer activation and
sustained accessibility at pro-survival loci, contributing to their pathogenic persistence.
Innate immune cells (including dendritic cells, macrophages, and natural killer (NK) cells)
also demonstrate epigenetic plasticity in response to environmental stimuli. Macrophages exposed
to lipopolysaccharide (LPS) or interferons exhibit rapid increases in chromatin accessibility at

258 Bioinformatics of Autoimmune Diseases
inammatory gene loci such as Il1b, Tnf, and Nos2 (Lara-Astiaso et al., 2014). These changes can
be transient or enduring, the latter giving rise to “trained immunity”, wherein innate cells acquire
memory-like features through persistent epigenomic modications. While this adaptation enhances
host defense, in autoimmune contexts it may exacerbate chronic inammation and tissue damage.
During chronic stimulation or persistent antigen exposure, T cells can enter a state of exhaustion
characterized by distinct transcriptional and epigenomic changes. Exhausted T cells, commonly
studied in cancer and chronic infections, exhibit increased chromatin accessibility at loci bound
by TOX and the NR4A family of TFs (Kurtulus et al., 2019). In autoimmune settings, emerging
evidence suggests that prolonged autoantigen exposure can induce similar exhaustion-associated
signatures. This phenomenon may serve as a regulatory mechanism to restrain hyperactive immunity, although it may also contribute to ineffective immune surveillance and incomplete resolution
of inammation.
A key unifying concept is epigenetic plasticity, the capacity of chromatin to reorganize in
response to cues from the environment or internal signaling pathways. In healthy individuals, this
plasticity allows immune cells to transition uidly between activation, regulation, and memory.
However, in autoimmune disease, this process becomes dysregulated. T cells from individuals
predisposed to autoimmunity often display hybrid epigenomic signatures (co-accessibility of both
effector and regulatory elements) indicating a conicted or unstable cellular identity (Van Der Byl
et al., 2024). Such ambiguity disrupts proper immune regulation and contributes to the emergence
of autoreactive phenotypes.
ATAC-Seq has also become central to the study of autoimmune diseases themselves, offering
insights into the epigenetic underpinnings of conditions such as SLE, RA, multiple sclerosis (MS),
and type 1 diabetes (T1D). In lupus, ATAC-Seq of peripheral blood mononuclear cells (PBMCs)
and CD4+ T cells has identied widespread enhancer activation at interferon-regulated genes and
TF binding sites for IRF5, STAT1, and NF-κB (Scharer et al., 2019). These disease-specic acces-
sible regions often reside near critical genes such as IL2RA and TNFAIP3, implicating enhancer
dysregulation in disease pathogenesis.
In RA, studies of synovial broblasts and inltrating immune cells have revealed increased
chromatin accessibility at enhancers linked to STAT4, PTPN22, and CTLA4, genes implicated
in RA susceptibility and inammatory signaling (Zhang et al., 2019). These ndings support the
concept that the autoimmune microenvironment actively recongures the epigenome to sustain
inammation.
MS further illustrates how ATAC-Seq can reveal epigenomic changes in disease-relevant T-cell
subsets. MS risk loci identied by genome-wide association studies (GWAS), such as those near
IL7R, TYK2, and EOMES, overlap with accessible chromatin regions in CD4+ T cells, suggesting that genetic risk is mediated through changes in chromatin architecture (Calderon et al., 2019).
These accessible sites are enriched in Th17 and regulatory T-cell lineages, linking T-cell plasticity
to MS pathophysiology.
In T1D, disease-specic accessible regions have been identied near IL2RA and FOXP3 in both
peripheral and islet-inltrating immune cells (Garg et al., 2012; Fasolino et al., 2020). Many of these
regions harbor T1D-associated single nucleotide polymorphisms (SNPs) that disrupt TF binding or
enhancer activity, highlighting how epigenetic regulation mediates genetic risk.
One of the most powerful applications of ATAC-Seq in autoimmunity is its integration with
GWAS data. By overlaying accessible chromatin maps with disease-associated SNPs, researchers
have linked non-coding variants, formerly regarded as nonfunctional, to specic regulatory elements. These efforts have revealed enhancer–promoter interactions involving key immunoregulatory genes, including STAT4, IL2RA, and CTLA4 (Corces et al., 2016). Functional validation using
techniques like CRISPR interference has conrmed that many of these enhancers directly inuence
gene expression and immune cell function.
Together, these studies demonstrate that autoimmunity arises not only from coding mutations
or environmental insults but from subtle and dynamic disruptions in gene regulation embedded in

259 ATAC-Seq for Autoimmune Diseases
the non-coding genome. As the eld advances, integrating ATAC-Seq with transcriptomics, proteomics, and single-cell resolution technologies will deepen our understanding of immune cell heterogeneity and identify new therapeutic targets based on regulatory architecture.
8.4 THE ATAC-Seq METHOD
In the ongoing effort to understand how the structure of chromatin inuences gene regulation,
the development of technologies that reveal chromatin accessibility has become central to modern
molecular biology. Among these, ATAC-Seq has emerged as a transformative method. ATAC-Seq
allows researchers to map open chromatin regions across the genome with exceptional sensitivity
and speed, offering a window into the epigenetic regulation of gene expression. But to fully appreciate the power of ATAC-Seq, it is important to understand how the method evolved, how it works,
and why it now supersedes older technologies like DNase-Seq and formaldehyde-assisted isolation
of regulatory elements (FAIRE-Seq).
The story of ATAC-Seq begins with the broader scientic challenge of mapping chromatin
accessibility. For decades, scientists relied on techniques such as DNase I hypersensitivity assays
to detect regions of the genome where DNA was exposed and potentially transcriptionally active.
DNase-Seq, which adapted this enzymatic method for high-throughput sequencing, allowed for the
identication of accessible chromatin at a genome-wide scale. Around the same time, FAIRE-Seq
was developed, using chemical cross-linking followed by phenol–chloroform extraction to isolate
nucleosome-depleted regions. Both methods signicantly advanced the eld but came with technical demands: DNase-Seq required large amounts of starting material and involved cumbersome
digestion and purication steps, while FAIRE-Seq suffered from lower resolution and limited sensitivity in certain cell types.
ATAC-Seq was introduced in 2013 by Jason Buenrostro and colleagues at Stanford University
as a simpler, faster, and more sensitive alternative. The method is built on an elegant principle: it
exploits a hyperactive Tn5 transposase enzyme, which has been engineered to simultaneously cut
open chromatin and insert sequencing adapters into accessible regions of DNA. This process, often
termed “tagmentation”, combines tagging and fragmentation in a single step (Buenrostro et al.,
2013). In practice, cells or nuclei are gently lysed, and the Tn5 transposase is introduced. Because
the enzyme preferentially targets nucleosome-free regions (NFRs), it inserts adapters into sites of
open chromatin. The resulting DNA fragments (enriched for regulatory elements such as promoters,
enhancers, and TF binding sites) can be directly amplied and sequenced. This not only reveals the
locations of accessible chromatin but also provides information about nucleosome positioning and
TF occupancy, all within a single assay.
Compared to DNase-Seq and FAIRE-Seq, ATAC-Seq offers several clear advantages. First, it
requires remarkably low input material. While DNase-Seq typically demands millions of cells,
ATAC-Seq protocols can be performed with as few as 500–50,000 cells, making it ideal for primary tissues, rare cell populations, and even single-cell applications. Second, the protocol is signicantly faster, often completed in less than a day, and it does not require complex digestion curves
or organic extractions. Third, ATAC-Seq yields high-resolution data, capturing precise chromatin
features with minimal background noise. Moreover, because the same method can be adapted for
bulk tissue, sorted populations, or single-cell analysis, it provides a scalable platform suitable for
diverse experimental designs.
To perform ATAC-Seq, a series of carefully controlled laboratory steps are followed. The process
begins with the preparation of viable, intact cells (either fresh or cryopreserved) which are washed
and pelleted to remove contaminants. The cells are then gently lysed using a non-ionic detergent in a
cold lysis buffer to release intact nuclei while minimizing chromatin disruption. Following lysis, the
nuclei are incubated with a transposition reaction mix containing the hyperactive Tn5 transposase
and pre-loaded sequencing adapters, which are used to tag accessible regions of DNA during transposition, enabling subsequent polymerase chain reaction (PCR) amplication and high-throughput

260 Bioinformatics of Autoimmune Diseases
sequencing of those regions. This incubation is typically carried out at 37°C for 30 minutes. During
this step, the transposase inserts the adapters into regions of accessible DNA, effectively fragmenting and tagging these sites in a single reaction.
After the transposition reaction, the tagged DNA fragments are puried (usually with a silica
column or magnetic beads) to remove proteins, salts, and residual enzymes. The puried DNA is
then subjected to PCR amplication using primers that target the adapter sequences added by Tn5.
This amplication step enriches for successfully tagged fragments and adds sequencing platformspecic indices and barcodes. Often, a preliminary quantitative PCR (qPCR) is performed on a
small aliquot to determine the optimal number of additional amplication cycles, avoiding overamplication that could skew fragment distribution or increase duplication rates.
At this stage, the DNA library is referred to as “sequencing-ready”, but it must undergo two
critical steps before loading onto a sequencer: quantication and normalization. Quantication
measures the concentration of DNA in the library to ensure there is sufcient material for sequencing. This is typically performed using uorometric assays such as Qubit, which specically measure double-stranded DNA, or with qPCR-based methods that also verify the presence of properly
ligated adapters. Accurate quantication ensures efcient cluster generation on the ow cell during
sequencing. Normalization refers to adjusting the concentration of the DNA library to a standard
input amount compatible with the sequencing instrument. If multiple libraries are being multiplexed
(sequenced together in the same run), normalization ensures that each contributes proportionally
to the total read count. This avoids imbalances where one sample might dominate the sequencing
output, which could compromise data quality and comparability across samples.
The nal, quantied, and normalized ATAC-Seq library is then loaded onto a high-throughput
sequencing platform such as Illumina’s NovaSeq or NextSeq. These instruments employ sequencing-by-synthesis chemistry to determine the nucleotide sequence of millions of DNA fragments in
parallel. Once loaded onto the ow cell, each DNA fragment is amplied into clusters and sequenced
base-by-base. The output consists of short reads corresponding to regions of open chromatin (precisely those areas where the Tn5 transposase was able to access and tag DNA).
Once sequencing is complete, the reads are mapped to a reference genome using alignment
tools such as Bowtie2 or Burrows–Wheeler Aligner (BWA). This is followed by the identication
of peaks (regions of the genome enriched for sequencing reads) using peak calling algorithms like
MACS2, which are also commonly used in ChIP-Seq data analysis. In this respect, the bioinformatics workow of ATAC-Seq shares many similarities with the ChIP-Seq analysis pipeline described
in the previous chapter. Both approaches involve quality control (QC) of raw sequencing reads,
adapter trimming, genome alignment, duplicate removal, and peak detection.
However, there are important differences between ATAC-Seq and ChIP-Seq data analysis that
stem from the nature of the data each method generates. While ChIP-Seq identies DNA regions
bound by specic proteins (such as TFs or modied histones) by immunoprecipitating DNA–protein
complexes, ATAC-Seq identies regions of accessible chromatin without requiring antibodies. This
distinction means that ATAC-Seq peaks represent a combination of regulatory elements, including promoters, enhancers, insulators, and TF binding sites, rather than binding sites for a single,
targeted protein. Consequently, ATAC-Seq data often requires additional annotation steps to functionally categorize the peaks, such as overlapping them with known genomic features or integrating
them with TF motif analysis and footprinting tools like TOBIAS or HINT-ATAC.
Another key difference lies in fragment size analysis. ATAC-Seq generates a characteristic size
distribution pattern that reects nucleosome positioning (short fragments (~100 bp) correspond
to NFRs, while longer fragments (~200–500 bp) correspond to mono- and di-nucleosomes). This
nucleosome periodicity offers an additional layer of biological interpretation not typically obtained
from ChIP-Seq data. Therefore, ATAC-Seq analysis often includes fragment length decomposition
and nucleosome occupancy proling to map both open chromatin and nucleosome architecture.
In summary, although the computational analysis of ATAC-Seq shares many foundational
steps with ChIP-Seq, the differences in biological input and data complexity necessitate additional

261 ATAC-Seq for Autoimmune Diseases
analytical layers specic to ATAC-Seq. These include motif footprinting, nucleosome positioning,
and chromatin state integration, which provide a richer view of genome-wide regulatory dynamics
and make ATAC-Seq uniquely suited for studying the epigenetic landscape in health and disease.
8.5 EXPERIMENTAL DESIGN AND SAMPLE PREPARATION FOR ATAC-Seq
The successful implementation of ATAC-Seq begins long before sequencing, with the careful
design of experiments and thoughtful preparation of biological samples. In studies of autoimmune
diseases, where the immune system’s dysregulation is rooted in subtle epigenetic changes, these preparatory steps are especially critical. The reproducibility and interpretability of results hinge upon
the biological context of the samples, the integrity of the nuclei, and the uniformity of the chromatin
accessibility landscape captured during the assay (Buenrostro et al., 2013).
A fundamental decision in ATAC-Seq experimental design is the choice of tissue or cell type.
PBMCs are commonly used in autoimmune research due to their accessibility and their representation of diverse immune cell populations (Corces et al., 2016). They provide a practical window
into systemic immune responses, capturing signals from T cells, B cells, monocytes, and NK cells.
However, PBMCs represent a mixed cell population, and in diseases where specic subtypes are
implicated, such as CD4+ T helper cells in MS or regulatory T cells in lupus, this cellular heterogeneity can obscure subtle, cell-specic chromatin accessibility changes (Mumbach et al., 2017).
To overcome this, researchers often employ uorescence-activated cell sorting (FACS) or magnetic bead–based separation techniques to isolate puried immune subsets prior to ATAC-Seq. This
approach enhances signal resolution and allows for the identication of transcriptional regulatory
elements unique to distinct immune lineages (Buenrostro et al., 2015).
In certain cases, particularly when the autoimmune pathology is localized, such as synovial tissue in RA or intestinal mucosa in Crohn’s disease, biopsy samples provide a more direct view of the
diseased environment. However, these tissues present their own challenges. Autoimmune lesions
are often inltrated by a variety of immune and stromal cells, leading to profound cellular heterogeneity (Scharer et al., 2019). This complexity can confound interpretation unless further fractionation
or single-nucleus approaches are used. Moreover, biopsies are typically low in cell number and may
undergo substantial stress during isolation, impacting chromatin integrity (Preissl et al., 2018).
A related consideration is the state of the sample: fresh versus frozen. Fresh samples are ideal,
as they preserve chromatin accessibility proles with minimal artifacts. However, for many clinical
studies and biobanked materials, frozen samples are the only option. Freezing introduces potential
complications such as membrane rupture and chromatin degradation, which can affect transposase
accessibility and reduce data quality (Corces et al., 2017). Despite these drawbacks, protocols have
been adapted to enable effective ATAC-Seq from cryopreserved cells and even frozen tissue sections. Cryopreservation protocols must be optimized to minimize ice crystal formation and maintain nuclear structure, while downstream steps, particularly nuclei isolation, must be carefully
controlled to ensure the quality and consistency of the data (Fujiwara et al., 2019).
The isolation of intact nuclei is a crucial step in ATAC-Seq, especially in tissues with high
cellular heterogeneity or dense extracellular matrices. Detergent-based lysis buffers are typically
used to release nuclei, with optimization required to avoid over-digestion or loss of nuclear content.
Following isolation, nuclei are often stained and counted to ensure input consistency. QC metrics
such as nuclear integrity, fragment size distribution, and transcription start site (TSS) enrichment
scores are essential to determine the suitability of each sample for sequencing Buenrostro et al.
(2015). A characteristic ATAC-Seq library should exhibit a nucleosomal ladder pattern on a fragment analyzer, reecting mono-, di-, and tri-nucleosome fragments, and a high TSS enrichment
score indicating successful capture of regulatory regions (Meers et al., 2019).
Autoimmune research presents particular challenges in sample acquisition. Tissues of interest
are often inamed, brotic, or scarce, making standardized processing difcult. Moreover, ethical and logistical constraints can limit the availability of samples from active disease sites. The

262 Bioinformatics of Autoimmune Diseases
dynamic and uctuating nature of autoimmune conditions also complicates timing; epigenomic
states can change rapidly in response to disease activity or treatment, making temporal standard-
ization critical (Mazzone et al., 2019). In some cases, longitudinal sampling from the same patient,
although technically and ethically demanding, can help mitigate inter-individual variability and
provide insight into disease progression.
Finally, the inherent cellular heterogeneity in autoimmune diseases cannot be overstated. The
immune landscape in affected tissues is shaped not only by resident cells but also by inltrating
immune populations, whose proportions and states can vary signicantly between patients and
disease stages. Bulk ATAC-Seq provides averaged signals that may obscure critical regulatory differences between rare but functionally important cell subsets. This limitation has spurred growing interest in single-cell ATAC-Seq (scATAC-Seq), which offers higher resolution at the cost of
increased complexity and data sparsity (Satpathy et al., 2019).
In sum, ATAC-Seq experimental design for autoimmune research demands meticulous attention
to biological context, cell purity, and technical parameters. From the selection of tissue and sample
state to the isolation of nuclei and application of quality metrics, each step inuences the delity
with which chromatin accessibility landscapes are captured. Addressing the practical and biological
challenges unique to autoimmune disease studies is essential for generating meaningful and reproducible insights into the epigenetic mechanisms that govern immune dysfunction.
8.6 ATAC-Seq DATA PROCESSING PIPELINE
The analysis of ATAC-Seq data begins with a meticulous and highly structured bioinformatics
pipeline, designed to transform raw sequencing reads into biologically meaningful insights about
chromatin accessibility, nucleosome organization, and TF binding. This section details the critical
steps involved in this pipeline, from preprocessing to data visualization and interpretation.
Following sequencing, the rst essential step is read trimming and QC. The reads generated by
ATAC-Seq contain adapter sequences introduced by the transposase enzyme used during library
preparation. These adapters must be accurately removed to prevent interference with downstream
alignment and peak calling. Tools such as Trimmomatic or Cutadapt are commonly employed to
trim these adapters, as well as to remove low-quality bases from the ends of reads. After trimming,
QC is performed using programs like FastQC to assess metrics such as per-base quality scores, GC
content, and sequence duplication levels. Only high-quality reads that pass these initial checks are
retained for further analysis.
The next phase involves aligning the trimmed reads to a reference genome, typically using a
high-performance aligner such as Bowtie2 or BWA. Given the short and variable length of ATACSeq fragments (ranging from those reecting NFRs to mono-, di-, or tri-nucleosome structures),
the aligner must be able to handle such diversity efciently. It is also common practice to lter out
reads that map to mitochondrial DNA, as these often represent a disproportionate fraction of the
sequencing data and may obscure signals from nuclear chromatin regions. Additionally, duplicate
reads resulting from PCR amplication are marked or removed to prevent biases in downstream
analyses.
Once the reads are successfully mapped, peak calling is performed to identify regions of open
chromatin where the transposase enzyme inserted sequencing adapters. These regions represent
accessible DNA and are often indicative of regulatory activity. One of the most widely used tools for
this purpose is MACS2 (Model-based Analysis of ChIP-Seq), which has been adapted for ATACSeq by treating the aligned fragments as single-end reads centered on the estimated transposase
cut sites. MACS2 identies statistically signicant peaks (clusters of reads that are enriched over
background noise) suggesting areas of open chromatin, such as enhancers, promoters, or insulators. The resulting peak les can then be annotated with genomic features, such as gene promoters
or enhancers, using tools like HOMER (Hypergeometric Optimization of Motif EnRichment) or
ChIPseeker to contextualize their regulatory relevance.

263 ATAC-Seq for Autoimmune Diseases
To interpret these results visually and explore specic genomic regions in detail, genome browsers like the Integrative Genomics Viewer (IGV) or the University of California, Santa Cruz (UCSC)
Genome Browser are indispensable. These platforms allow researchers to load their aligned read
les (typically in Binary Alignment/Map (BAM) format) and peak les (in Browser Extensible
Data (BED) format) to examine patterns of chromatin accessibility across the genome. In IGV, for
example, users can zoom into promoter regions of key immune genes to assess the depth and distribution of ATAC-Seq signals, compare accessibility across samples, or overlay other epigenomic
datasets, such as ChIP-Seq for histone modications or TFs. UCSC, on the other hand, provides
additional annotation layers, including conservation scores, gene models, and ENCODE datasets,
which help place the observed chromatin accessibility into a broader functional and evolutionary
context.
Beyond identifying open chromatin regions, ATAC-Seq data enables deeper analyses of chromatin architecture, particularly nucleosome positioning and TF footprinting. The fragment size distribution of ATAC-Seq reads inherently reects the structure of chromatin. Short fragments, typically
less than 100 bp, indicate NFRs and correspond to sites of direct DNA–protein interaction, such
as active TF binding. In contrast, fragments approximately 200 bp in length correspond to mononucleosomes, while longer fragments represent di- and tri-nucleosomes, revealing the higher-order
structure of chromatin around accessible regions. Tools such as NucleoATAC can model these patterns to infer precise nucleosome positions, generating maps that highlight the phased organization
of nucleosomes anking active regulatory elements.
TF footprinting, another powerful application of ATAC-Seq, leverages the ne-scale resolution
of transposase insertions. When a TF is bound to DNA, it protects a short region from transposase
cleavage, creating a local depletion or “footprint” in the accessibility signal anked by regions of
high transposase activity. Computational tools such as HINT-ATAC and TOBIAS can detect these
footprints and match them to known TF binding motifs, revealing the identity and activity of regulatory proteins operating within specic cell types or disease contexts. In autoimmune diseases,
where immune cell function and identity are often dysregulated, such footprinting analyses can
uncover key TFs that drive pathological gene expression programs.
Taken together, the ATAC-Seq data processing pipeline not only facilitates the identication of
open chromatin landscapes but also enables a nuanced understanding of the underlying regulatory
architecture. By integrating quality-controlled sequencing data with genome alignment, peak calling, visualization, and advanced modeling of nucleosome positions and TF footprints, researchers
can reconstruct the epigenetic states that govern immune cell function and contribute to autoimmune pathogenesis. This layered approach provides a powerful framework for linking chromatin
accessibility to gene regulation in health and disease.
8.6.1 RAW DATA ACQUISITION
In a typical ATAC-Seq experiment, raw data acquisition begins in the laboratory, where accessible
regions of chromatin are tagged using a hyperactive Tn5 transposase, followed by high-throughput
sequencing of the resulting fragments. These raw sequencing reads (usually in FASTQ (FAST
Quality) format) are generated directly from the sequencer and represent the initial input for downstream bioinformatics analysis. However, for demonstration purposes, rather than generating data
in a wet-lab setting, we will obtain publicly available raw ATAC-Seq data from the NCBI Sequence
Read Archive (SRA). The SRA is a comprehensive repository of sequencing data submitted by
researchers worldwide and offers a valuable resource for practicing computational workows,
reproducing published ndings, or benchmarking analysis pipelines. Through careful selection of
relevant datasets, such as those derived from autoimmune disease studies, we can simulate a complete ATAC-Seq analysis starting from real, biologically meaningful raw data.
You can use the “search _ ATAC _ seq.py” Python script, provided as supplementary mate-
rial, to search the NCBI BioProject database for raw ATAC-Seq data. This script utilizes the NCBI
Соседние файлы в папке Библиотека им академика М.И. Перельмана
