Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
26 Мб
Скачать
254 Bioinformatics of Autoimmune Diseases
From a methodological perspective, the chapter meticulously outlines the experimental design for ChIP-Seq in autoimmune disease research. It includes considerations for sample selection, anti­body specicity, control strategies, and sequencing depth. A step-by-step workow is described, starting from raw data QC and alignment, through peak calling using tools like MACS3, to down­stream annotation and functional interpretation. Key statistical metrics such as signalValue, p-value, q-value, and peak score are clearly explained, offering readers the conceptual tools to assess the reliability and biological relevance of detected binding events.
Integration with other datasets such as RNA-Seq or ATAC-Seq is discussed as a crucial next step, enabling the construction of multi-layered regulatory networks that reect the complexity of immune regulation. The chapter also presents a working example using publicly available data to demonstrate practical application, including a Python pipeline for peak calling and annotation, as well as the use of HOMER for genomic feature association and motif discovery.
Altogether, this chapter serves as a technically detailed yet conceptually rich resource that bridges molecular biology, immunology, and bioinformatics. It situates ChIP-Seq not just as a labo­ratory technique, but as a window into the epigenomic mechanisms that drive immune identity and misregulation in disease. The clarity of exposition, combined with practical guidance and relevance to human pathology, makes this a valuable reference for both experimentalists and computational researchers working in epigenetics and autoimmune biology.
BIBLIOGRAPHY
Allis, C. D., & Jenuwein, T. (2016). The molecular hallmarks of epigenetic control. Nature Reviews Genetics,
17(8), 487–500. https://doi.org/10.1038/nrg.2016.59
Barski, A., et al. (2007). High-resolution proling of histone methylations in the human genome. Cell, 129(4),
82 3– 837. https://doi.org/10.1016/j.cell.2007.05.009
Ismail, H. D. (2022). Bioinformatics: A practical guide to NCBI databases and sequence alignments (1st ed.).
Chapman and Hall/CRC. https://doi.org/10.1201/9781003226611
Ismail, H. D. (2023). Bioinformatics: A practical guide to next generation sequencing data analysis (1st ed.).
Chapman and Hall/CRC. https://doi.org/10.1201/9781003355205
Luger, K., Mäder, A. W., Richmond, R. K., Sargent, D. F., & Richmond, T. J. (1997). Crystal structure of the
nucleosome core particle at 2.8 Å resolution. Nature, 389(6648), 251–260. https://doi.org/10.1038/38444
National Center for Biotechnology Information (NCBI). (2023). BioProject PRJNA1020118: H3K27ac ChIP-
Seq after doxycycline induction of AIRE [Data set]. NCBI Sequence Read Archive. https://www.ncbi.
nlm.nih.gov/bioproject/PRJNA1020118
Ptashne, M. (2005). Regulation of transcription: From lambda to eukaryotes. Trends in Biochemical Sciences,
30(6), 275–279. https://doi.org/10.1016/j.tibs.2005.04.001
Scharer, C. D., et al. (2019). Epigenetic programming underpins B cell dysfunction in human SLE. Nature
Im munology, 20(8), 1071–1082. https://doi.org/10.1038/s41590-019-0419-9
Sun, Y., Li, Y., Luo, D., Liao, D. J., & Wang, X. (2015). Epigenetic regulation of autoimmunity by DNA meth-
ylation. Journal of Autoimmunity, 64, 1–12. https://doi.org/10.1016/j.jaut.2015.07.002
Zhang, Q., Vignali, D. A. A., & Chang, M. (2018). Epigenetic regulation of autoimmune disease-associated
gene expression. Trends in Immunology, 39(7), 573–588. https://doi.org/10.1016/j.it.2018.04.004
ATAC-Seq for
8
Autoimmune Diseases
8.1 GENE ACCESSIBILITY
In the previous chapter, we discussed ChIP-Seq and its utility in detecting transcription factor (TF) binding sites and histone modications. We also examined the structure of chromatin and its critical role in gene regulation, particularly through its inuence on DNA accessibility. Chromatin can exist in a more open or closed state, and this dynamic structure determines whether certain genes, often otherwise silent, can become accessible for transcription and thereby inuence gene expression.
In recent years, there has been growing recognition that alterations in chromatin accessibility (changes that affect how tightly DNA is packaged within the nucleus) play a central role in shaping immune cell function and immune dysregulation. Chromatin architecture dictates which regions of the genome are open to TF binding and, consequently, which genes are activated or repressed. By studying these patterns of accessibility, researchers can gain insights into the regulatory networks that orchestrate immune responses and identify epigenetic disturbances associated with autoim­mune diseases.
This chapter explores the central role of chromatin accessibility in autoimmune pathogenesis through the lens of a powerful epigenomic tool: ATAC-Seq (Assay for Transposase-Accessible Chromatin using sequencing). ATAC-Seq enables high-resolution, genome-wide proling of open chromatin regions by utilizing a hyperactive transposase enzyme that preferentially inserts sequencing adapters into accessible DNA. This technique has become particularly valuable in immunological studies, where chromatin remodeling plays a key role in processes such as T cell activation, B-cell differentiation, and the establishment of immunological memory. By applying ATAC-Seq to immune cells from patients with autoimmune disorders, researchers can identify disease-specic regulatory elements and TF binding motifs that may drive aberrant immune responses.
When compared to ChIP-Seq, a widely used method for detecting specic protein–DNA interac­tions such as TF occupancy or histone modications, ATAC-Seq offers several distinct advantages. First, ATAC-Seq is less labor-intensive and requires signicantly fewer cells, making it well-suited for studies involving rare immune cell subsets or limited patient material. Unlike ChIP-Seq, which depends on antibodies against known proteins and thus focuses on predened targets, ATAC-Seq is an unbiased technique that reveals all accessible regions of the genome regardless of the proteins involved. This provides a broader, more global view of the chromatin landscape. Moreover, ATAC­Seq can be performed on frozen samples and is compatible with single-cell analysis, allowing for the resolution of cell type–specic regulatory patterns within heterogeneous populations (capabili­ties that are difcult to achieve with ChIP-Seq).
In addition to introducing the fundamental principles of ATAC-Seq, this chapter highlights its applications in immunology and autoimmunity. It explores how ATAC-Seq can be used to map enhancer elements, identify shifts in chromatin accessibility during immune activation, and uncover regulatory circuits involved in disease. The chapter also provides a detailed overview of computa­tional strategies for analyzing ATAC-Seq data, including peak calling, motif discovery, integration with RNA-Seq and ChIP-Seq datasets, and functional annotation of regulatory regions. Together, these approaches position ATAC-Seq as a transformative tool for decoding the non-coding genome and illuminating the regulatory underpinnings of autoimmune disease.
255 D OI: 10.1201/ 978 10 03 68 5 432- 8
256 Bioinformatics of Autoimmune Diseases
8.2 CHROMATIN ACCESSIBILITY AND IMMUNE REGULATION
The human genome, with its vast 3 billion base pairs, is intricately folded within the microscopic boundaries of the cell nucleus. This packaging is accomplished through the organization of DNA into chromatin, a complex structure in which DNA is wound around histone proteins to form nucleo­somes (Luger et al., 1997). Chromatin is not a passive framework; it is a dynamic regulatory system that governs the accessibility of genetic information. The degree to which chromatin is open or closed (referred to as chromatin accessibility) determines whether genes can be actively transcribed or remain silent (Klemm et al., 2019). This accessibility serves as a foundational mechanism of epi- genetic regulation, particularly critical in the immune system, where precise and context-dependent gene expression denes the function and fate of immune cells (Lara-Astiaso et al., 2014).
Chromatin exists in two primary states (see Section 3.2.6): tightly packed heterochromatin, which is largely inaccessible to transcriptional machinery and associated with gene silencing, and loosely arranged euchromatin, which facilitates active gene expression (Allis & Jenuwein,
2016). This regulatory balance is modulated by a diverse array of epigenetic modications that
alter the chemical and structural conguration of chromatin without altering the underlying DNA sequence. Histone modications such as acetylation and methylation are central to this regula­tion. Histone acetylation, catalyzed by histone acetyltransferases (HATs), is generally associ­ated with transcriptional activation, as it reduces the electrostatic interaction between DNA and histones, thereby increasing accessibility (Bannister & Kouzarides, 2011). In contrast, histone methylation can either activate or repress gene expression depending on the specic lysine resi­due modied and the degree of methylation. For instance, trimethylation of H3K4 is linked to active transcription, whereas trimethylation of H3K27 is typically associated with gene repres­sion (Barski et al., 2007).
Chromatin accessibility specically refers to the physical openness of DNA regions that allows TFs and other regulatory proteins to access and bind their target sequences. This openness is often a prerequisite for transcriptional activation and is frequently observed at promoters, enhancers, and other cis-regulatory elements (Buenrostro et al., 2013). High-throughput sequencing technolo- gies such as ATAC-Seq have revolutionized the ability to map open chromatin genome-wide with high resolution. These maps have revealed the regulatory architecture of distinct cell types and states, proving especially informative in deciphering the differentiation and functional response of immune cells (Corces et al., 2016).
In immune cells, chromatin accessibility is nely tuned to orchestrate lineage-specic gene expression programs. The differentiation of hematopoietic progenitor cells into various immune subsets (including T cells, B cells, macrophages, and dendritic cells) requires coordinated changes in chromatin organization. These changes involve both nucleosome repositioning and the deposition of specic histone modications. For example, cell type–specic enhancers marked by H3K4me1 and H3K27ac become accessible in a manner that drives the transcription of immunoregulatory genes essential for cellular identity and immune function (Heinz et al., 2010). These epigenetic mechanisms ensure that each immune cell type activates the appropriate gene networks only under the right circumstances, preserving the specicity and timing of immune responses.
In autoimmune diseases, however, the precise regulation of chromatin accessibility can become aberrant. Disruption in chromatin landscapes may lead to the inappropriate activation of genes involved in inammation, cytokine signaling, and immune cell proliferation (Scharer et al., 2019). For instance, in diseases such as systemic lupus erythematosus (SLE) and rheumatoid arthritis (RA), open chromatin regions have been identied at the regulatory elements of genes like IL6, Interferon gamma (IFNG), and tumor necrosis factor (TNF), contributing to chronic and patho­genic inammation. Moreover, defects in regulatory T cells, such as diminished accessibility at the FOXP3 locus, can impair their suppressive function and compromise peripheral tolerance,
allowing autoreactive T cells to escape regulation and drive autoimmunity (Floess et al., 2007).
Environmental factors (including infections, microbiome composition, and dietary exposures) can
257 ATAC-Seq for Autoimmune Diseases
also induce epigenetic modications that alter chromatin architecture, increase exposure to self­antigens, and ultimately contribute to autoimmune pathogenesis (Bach, 2018).
These dynamic changes in chromatin are not merely downstream consequences of autoimmune pathology; rather, they are often integral to the onset and progression of disease. As such, under­standing chromatin remodeling in immune cells provides new avenues for diagnosing, monitoring, and treating autoimmune disorders. Epigenetic therapeutics aimed at modifying chromatin acces­sibility or reversing aberrant enhancer activation are under exploration as promising strategies to
restore immune tolerance and suppress pathological immune responses (Peters et al., 2015). By
reshaping the chromatin landscape, it may be possible to reprogram immune cells toward a non­pathogenic state, offering transformative potential in the treatment of complex immune-mediated diseases.
8.3 ATAC-Seq IN IMMUNE CELLS
Within the intricate orchestration of immune responses, chromatin accessibility plays a decisive role in shaping the transcriptional fate of immune cells. ATAC-Seq has emerged as a powerful method for proling genome-wide chromatin accessibility at high resolution, revealing the dynamic remodeling of the epigenetic landscape during immune cell activation, differentiation, and dysfunc­tion (Buenrostro et al., 2013). By applying ATAC-Seq to diverse immune cell lineages (including Tcells, B cells, and innate immune cells), researchers have begun to decipher the regulatory logic that governs immune homeostasis and the disruptions that underlie autoimmune pathology.
In T lymphocytes, chromatin accessibility undergoes profound alterations as cells transition from naïve to activated and effector states. Naïve CD4+ T cells, upon encountering antigen, differentiate into lineage-committed subsets such as Th1 and Th17, each dened by distinct transcriptional and epigenomic signatures. ATAC-Seq has shown that this functional divergence is marked by the selec­tive opening of enhancer regions linked to lineage-dening TFs. For example, Th1 cells exhibit increased accessibility at the Ifng and Tbx21 loci, supporting the expression of interferon-gamma and the TF T-bet, respectively, whereas Th17 cells demonstrate accessible chromatin at Il17a, Il17f, and Rorc, consistent with their pro-inammatory identity (Shih et al., 2014; Yosef et al., 2013). These chromatin landscapes precede full transcriptional commitment, indicating that epigenomic remodeling acts as a priming mechanism for lineage determination.
Regulatory T cells (Tregs), critical for maintaining immune tolerance, also display a distinc­tive chromatin architecture. The TF FoxP3 orchestrates Treg development and function, in part by promoting accessibility at enhancers of genes such as Ctla4, Tnfrsf18 (GITR), and Ikzf2 (Samstein
et al., 2012). However, in autoimmune conditions, Tregs may exhibit altered chromatin states that
compromise their suppressive capacity or promote phenotypic plasticity toward effector lineages. This epigenetic instability is particularly pronounced in inamed tissues, suggesting that a hostile microenvironment can subvert Treg identity and function by reshaping their chromatin landscape
(Schmidl et al., 2009).
B cells similarly undergo extensive epigenomic remodeling during maturation and antigen-driven responses. Throughout the germinal center (GC) reaction, B cells proliferate, undergo somatic hypermutation, and initiate class switch recombination, processes regulated by TFs and chromatin dynamics. ATAC-Seq analyses reveal that memory B cells and antibody-secreting plasma cells pos­sess distinct accessibility patterns. Plasma cells, for instance, exhibit reduced accessibility at loci governing antigen presentation and B-cell receptor signaling, while showing increased accessibility near genes involved in antibody secretion such as Prdm1 (BLIMP-1) and Xbp1 (Zhou et al., 2015). In systemic autoimmune diseases such as lupus, B cells display aberrant enhancer activation and sustained accessibility at pro-survival loci, contributing to their pathogenic persistence.
Innate immune cells (including dendritic cells, macrophages, and natural killer (NK) cells) also demonstrate epigenetic plasticity in response to environmental stimuli. Macrophages exposed to lipopolysaccharide (LPS) or interferons exhibit rapid increases in chromatin accessibility at
258 Bioinformatics of Autoimmune Diseases
inammatory gene loci such as Il1b, Tnf, and Nos2 (Lara-Astiaso et al., 2014). These changes can be transient or enduring, the latter giving rise to “trained immunity”, wherein innate cells acquire memory-like features through persistent epigenomic modications. While this adaptation enhances host defense, in autoimmune contexts it may exacerbate chronic inammation and tissue damage.
During chronic stimulation or persistent antigen exposure, T cells can enter a state of exhaustion characterized by distinct transcriptional and epigenomic changes. Exhausted T cells, commonly studied in cancer and chronic infections, exhibit increased chromatin accessibility at loci bound by TOX and the NR4A family of TFs (Kurtulus et al., 2019). In autoimmune settings, emerging evidence suggests that prolonged autoantigen exposure can induce similar exhaustion-associated signatures. This phenomenon may serve as a regulatory mechanism to restrain hyperactive immu­nity, although it may also contribute to ineffective immune surveillance and incomplete resolution of inammation.
A key unifying concept is epigenetic plasticity, the capacity of chromatin to reorganize in response to cues from the environment or internal signaling pathways. In healthy individuals, this plasticity allows immune cells to transition uidly between activation, regulation, and memory. However, in autoimmune disease, this process becomes dysregulated. T cells from individuals predisposed to autoimmunity often display hybrid epigenomic signatures (co-accessibility of both
effector and regulatory elements) indicating a conicted or unstable cellular identity (Van Der Byl et al., 2024). Such ambiguity disrupts proper immune regulation and contributes to the emergence of autoreactive phenotypes.
ATAC-Seq has also become central to the study of autoimmune diseases themselves, offering insights into the epigenetic underpinnings of conditions such as SLE, RA, multiple sclerosis (MS), and type 1 diabetes (T1D). In lupus, ATAC-Seq of peripheral blood mononuclear cells (PBMCs) and CD4+ T cells has identied widespread enhancer activation at interferon-regulated genes and TF binding sites for IRF5, STAT1, and NF-κB (Scharer et al., 2019). These disease-specic acces- sible regions often reside near critical genes such as IL2RA and TNFAIP3, implicating enhancer dysregulation in disease pathogenesis.
In RA, studies of synovial broblasts and inltrating immune cells have revealed increased chromatin accessibility at enhancers linked to STAT4, PTPN22, and CTLA4, genes implicated
in RA susceptibility and inammatory signaling (Zhang et al., 2019). These ndings support the
concept that the autoimmune microenvironment actively recongures the epigenome to sustain inammation.
MS further illustrates how ATAC-Seq can reveal epigenomic changes in disease-relevant T-cell subsets. MS risk loci identied by genome-wide association studies (GWAS), such as those near IL7R, TYK2, and EOMES, overlap with accessible chromatin regions in CD4+ T cells, suggest­ing that genetic risk is mediated through changes in chromatin architecture (Calderon et al., 2019). These accessible sites are enriched in Th17 and regulatory T-cell lineages, linking T-cell plasticity to MS pathophysiology.
In T1D, disease-specic accessible regions have been identied near IL2RA and FOXP3 in both peripheral and islet-inltrating immune cells (Garg et al., 2012; Fasolino et al., 2020). Many of these regions harbor T1D-associated single nucleotide polymorphisms (SNPs) that disrupt TF binding or enhancer activity, highlighting how epigenetic regulation mediates genetic risk.
One of the most powerful applications of ATAC-Seq in autoimmunity is its integration with GWAS data. By overlaying accessible chromatin maps with disease-associated SNPs, researchers have linked non-coding variants, formerly regarded as nonfunctional, to specic regulatory ele­ments. These efforts have revealed enhancer–promoter interactions involving key immunoregula­tory genes, including STAT4, IL2RA, and CTLA4 (Corces et al., 2016). Functional validation using techniques like CRISPR interference has conrmed that many of these enhancers directly inuence gene expression and immune cell function.
Together, these studies demonstrate that autoimmunity arises not only from coding mutations or environmental insults but from subtle and dynamic disruptions in gene regulation embedded in
259 ATAC-Seq for Autoimmune Diseases
the non-coding genome. As the eld advances, integrating ATAC-Seq with transcriptomics, pro­teomics, and single-cell resolution technologies will deepen our understanding of immune cell het­erogeneity and identify new therapeutic targets based on regulatory architecture.
8.4 THE ATAC-Seq METHOD
In the ongoing effort to understand how the structure of chromatin inuences gene regulation, the development of technologies that reveal chromatin accessibility has become central to modern molecular biology. Among these, ATAC-Seq has emerged as a transformative method. ATAC-Seq allows researchers to map open chromatin regions across the genome with exceptional sensitivity and speed, offering a window into the epigenetic regulation of gene expression. But to fully appreci­ate the power of ATAC-Seq, it is important to understand how the method evolved, how it works, and why it now supersedes older technologies like DNase-Seq and formaldehyde-assisted isolation of regulatory elements (FAIRE-Seq).
The story of ATAC-Seq begins with the broader scientic challenge of mapping chromatin accessibility. For decades, scientists relied on techniques such as DNase I hypersensitivity assays to detect regions of the genome where DNA was exposed and potentially transcriptionally active. DNase-Seq, which adapted this enzymatic method for high-throughput sequencing, allowed for the identication of accessible chromatin at a genome-wide scale. Around the same time, FAIRE-Seq was developed, using chemical cross-linking followed by phenol–chloroform extraction to isolate nucleosome-depleted regions. Both methods signicantly advanced the eld but came with techni­cal demands: DNase-Seq required large amounts of starting material and involved cumbersome digestion and purication steps, while FAIRE-Seq suffered from lower resolution and limited sen­sitivity in certain cell types.
ATAC-Seq was introduced in 2013 by Jason Buenrostro and colleagues at Stanford University as a simpler, faster, and more sensitive alternative. The method is built on an elegant principle: it exploits a hyperactive Tn5 transposase enzyme, which has been engineered to simultaneously cut open chromatin and insert sequencing adapters into accessible regions of DNA. This process, often termed “tagmentation”, combines tagging and fragmentation in a single step (Buenrostro et al.,
2013). In practice, cells or nuclei are gently lysed, and the Tn5 transposase is introduced. Because
the enzyme preferentially targets nucleosome-free regions (NFRs), it inserts adapters into sites of open chromatin. The resulting DNA fragments (enriched for regulatory elements such as promoters, enhancers, and TF binding sites) can be directly amplied and sequenced. This not only reveals the locations of accessible chromatin but also provides information about nucleosome positioning and TF occupancy, all within a single assay.
Compared to DNase-Seq and FAIRE-Seq, ATAC-Seq offers several clear advantages. First, it requires remarkably low input material. While DNase-Seq typically demands millions of cells, ATAC-Seq protocols can be performed with as few as 500–50,000 cells, making it ideal for pri­mary tissues, rare cell populations, and even single-cell applications. Second, the protocol is signi­cantly faster, often completed in less than a day, and it does not require complex digestion curves or organic extractions. Third, ATAC-Seq yields high-resolution data, capturing precise chromatin features with minimal background noise. Moreover, because the same method can be adapted for bulk tissue, sorted populations, or single-cell analysis, it provides a scalable platform suitable for diverse experimental designs.
To perform ATAC-Seq, a series of carefully controlled laboratory steps are followed. The process begins with the preparation of viable, intact cells (either fresh or cryopreserved) which are washed and pelleted to remove contaminants. The cells are then gently lysed using a non-ionic detergent in a cold lysis buffer to release intact nuclei while minimizing chromatin disruption. Following lysis, the nuclei are incubated with a transposition reaction mix containing the hyperactive Tn5 transposase and pre-loaded sequencing adapters, which are used to tag accessible regions of DNA during trans­position, enabling subsequent polymerase chain reaction (PCR) amplication and high-throughput
260 Bioinformatics of Autoimmune Diseases
sequencing of those regions. This incubation is typically carried out at 37°C for 30 minutes. During this step, the transposase inserts the adapters into regions of accessible DNA, effectively fragment­ing and tagging these sites in a single reaction.
After the transposition reaction, the tagged DNA fragments are puried (usually with a silica column or magnetic beads) to remove proteins, salts, and residual enzymes. The puried DNA is then subjected to PCR amplication using primers that target the adapter sequences added by Tn5. This amplication step enriches for successfully tagged fragments and adds sequencing platform­specic indices and barcodes. Often, a preliminary quantitative PCR (qPCR) is performed on a small aliquot to determine the optimal number of additional amplication cycles, avoiding over­amplication that could skew fragment distribution or increase duplication rates.
At this stage, the DNA library is referred to as “sequencing-ready”, but it must undergo two critical steps before loading onto a sequencer: quantication and normalization. Quantication measures the concentration of DNA in the library to ensure there is sufcient material for sequenc­ing. This is typically performed using uorometric assays such as Qubit, which specically mea­sure double-stranded DNA, or with qPCR-based methods that also verify the presence of properly ligated adapters. Accurate quantication ensures efcient cluster generation on the ow cell during sequencing. Normalization refers to adjusting the concentration of the DNA library to a standard input amount compatible with the sequencing instrument. If multiple libraries are being multiplexed (sequenced together in the same run), normalization ensures that each contributes proportionally to the total read count. This avoids imbalances where one sample might dominate the sequencing output, which could compromise data quality and comparability across samples.
The nal, quantied, and normalized ATAC-Seq library is then loaded onto a high-throughput sequencing platform such as Illumina’s NovaSeq or NextSeq. These instruments employ sequenc­ing-by-synthesis chemistry to determine the nucleotide sequence of millions of DNA fragments in parallel. Once loaded onto the ow cell, each DNA fragment is amplied into clusters and sequenced base-by-base. The output consists of short reads corresponding to regions of open chromatin (pre­cisely those areas where the Tn5 transposase was able to access and tag DNA).
Once sequencing is complete, the reads are mapped to a reference genome using alignment tools such as Bowtie2 or Burrows–Wheeler Aligner (BWA). This is followed by the identication of peaks (regions of the genome enriched for sequencing reads) using peak calling algorithms like MACS2, which are also commonly used in ChIP-Seq data analysis. In this respect, the bioinformat­ics workow of ATAC-Seq shares many similarities with the ChIP-Seq analysis pipeline described in the previous chapter. Both approaches involve quality control (QC) of raw sequencing reads, adapter trimming, genome alignment, duplicate removal, and peak detection.
However, there are important differences between ATAC-Seq and ChIP-Seq data analysis that stem from the nature of the data each method generates. While ChIP-Seq identies DNA regions bound by specic proteins (such as TFs or modied histones) by immunoprecipitating DNA–protein complexes, ATAC-Seq identies regions of accessible chromatin without requiring antibodies. This distinction means that ATAC-Seq peaks represent a combination of regulatory elements, includ­ing promoters, enhancers, insulators, and TF binding sites, rather than binding sites for a single, targeted protein. Consequently, ATAC-Seq data often requires additional annotation steps to func­tionally categorize the peaks, such as overlapping them with known genomic features or integrating them with TF motif analysis and footprinting tools like TOBIAS or HINT-ATAC.
Another key difference lies in fragment size analysis. ATAC-Seq generates a characteristic size distribution pattern that reects nucleosome positioning (short fragments (~100 bp) correspond to NFRs, while longer fragments (~200–500 bp) correspond to mono- and di-nucleosomes). This nucleosome periodicity offers an additional layer of biological interpretation not typically obtained from ChIP-Seq data. Therefore, ATAC-Seq analysis often includes fragment length decomposition and nucleosome occupancy proling to map both open chromatin and nucleosome architecture.
In summary, although the computational analysis of ATAC-Seq shares many foundational steps with ChIP-Seq, the differences in biological input and data complexity necessitate additional
261 ATAC-Seq for Autoimmune Diseases
analytical layers specic to ATAC-Seq. These include motif footprinting, nucleosome positioning, and chromatin state integration, which provide a richer view of genome-wide regulatory dynamics and make ATAC-Seq uniquely suited for studying the epigenetic landscape in health and disease.
8.5 EXPERIMENTAL DESIGN AND SAMPLE PREPARATION FOR ATAC-Seq
The successful implementation of ATAC-Seq begins long before sequencing, with the careful design of experiments and thoughtful preparation of biological samples. In studies of autoimmune diseases, where the immune system’s dysregulation is rooted in subtle epigenetic changes, these pre­paratory steps are especially critical. The reproducibility and interpretability of results hinge upon the biological context of the samples, the integrity of the nuclei, and the uniformity of the chromatin accessibility landscape captured during the assay (Buenrostro et al., 2013).
A fundamental decision in ATAC-Seq experimental design is the choice of tissue or cell type. PBMCs are commonly used in autoimmune research due to their accessibility and their represen­tation of diverse immune cell populations (Corces et al., 2016). They provide a practical window into systemic immune responses, capturing signals from T cells, B cells, monocytes, and NK cells. However, PBMCs represent a mixed cell population, and in diseases where specic subtypes are implicated, such as CD4+ T helper cells in MS or regulatory T cells in lupus, this cellular hetero­geneity can obscure subtle, cell-specic chromatin accessibility changes (Mumbach et al., 2017). To overcome this, researchers often employ uorescence-activated cell sorting (FACS) or mag­netic bead–based separation techniques to isolate puried immune subsets prior to ATAC-Seq. This approach enhances signal resolution and allows for the identication of transcriptional regulatory elements unique to distinct immune lineages (Buenrostro et al., 2015).
In certain cases, particularly when the autoimmune pathology is localized, such as synovial tis­sue in RA or intestinal mucosa in Crohn’s disease, biopsy samples provide a more direct view of the diseased environment. However, these tissues present their own challenges. Autoimmune lesions are often inltrated by a variety of immune and stromal cells, leading to profound cellular heteroge­neity (Scharer et al., 2019). This complexity can confound interpretation unless further fractionation or single-nucleus approaches are used. Moreover, biopsies are typically low in cell number and may undergo substantial stress during isolation, impacting chromatin integrity (Preissl et al., 2018).
A related consideration is the state of the sample: fresh versus frozen. Fresh samples are ideal, as they preserve chromatin accessibility proles with minimal artifacts. However, for many clinical studies and biobanked materials, frozen samples are the only option. Freezing introduces potential complications such as membrane rupture and chromatin degradation, which can affect transposase accessibility and reduce data quality (Corces et al., 2017). Despite these drawbacks, protocols have been adapted to enable effective ATAC-Seq from cryopreserved cells and even frozen tissue sec­tions. Cryopreservation protocols must be optimized to minimize ice crystal formation and main­tain nuclear structure, while downstream steps, particularly nuclei isolation, must be carefully
controlled to ensure the quality and consistency of the data (Fujiwara et al., 2019).
The isolation of intact nuclei is a crucial step in ATAC-Seq, especially in tissues with high cellular heterogeneity or dense extracellular matrices. Detergent-based lysis buffers are typically used to release nuclei, with optimization required to avoid over-digestion or loss of nuclear content. Following isolation, nuclei are often stained and counted to ensure input consistency. QC metrics such as nuclear integrity, fragment size distribution, and transcription start site (TSS) enrichment scores are essential to determine the suitability of each sample for sequencing Buenrostro et al. (2015). A characteristic ATAC-Seq library should exhibit a nucleosomal ladder pattern on a frag­ment analyzer, reecting mono-, di-, and tri-nucleosome fragments, and a high TSS enrichment score indicating successful capture of regulatory regions (Meers et al., 2019).
Autoimmune research presents particular challenges in sample acquisition. Tissues of interest are often inamed, brotic, or scarce, making standardized processing difcult. Moreover, ethi­cal and logistical constraints can limit the availability of samples from active disease sites. The
262 Bioinformatics of Autoimmune Diseases
dynamic and uctuating nature of autoimmune conditions also complicates timing; epigenomic states can change rapidly in response to disease activity or treatment, making temporal standard-
ization critical (Mazzone et al., 2019). In some cases, longitudinal sampling from the same patient,
although technically and ethically demanding, can help mitigate inter-individual variability and provide insight into disease progression.
Finally, the inherent cellular heterogeneity in autoimmune diseases cannot be overstated. The immune landscape in affected tissues is shaped not only by resident cells but also by inltrating immune populations, whose proportions and states can vary signicantly between patients and disease stages. Bulk ATAC-Seq provides averaged signals that may obscure critical regulatory dif­ferences between rare but functionally important cell subsets. This limitation has spurred grow­ing interest in single-cell ATAC-Seq (scATAC-Seq), which offers higher resolution at the cost of increased complexity and data sparsity (Satpathy et al., 2019).
In sum, ATAC-Seq experimental design for autoimmune research demands meticulous attention to biological context, cell purity, and technical parameters. From the selection of tissue and sample state to the isolation of nuclei and application of quality metrics, each step inuences the delity with which chromatin accessibility landscapes are captured. Addressing the practical and biological challenges unique to autoimmune disease studies is essential for generating meaningful and repro­ducible insights into the epigenetic mechanisms that govern immune dysfunction.
8.6 ATAC-Seq DATA PROCESSING PIPELINE
The analysis of ATAC-Seq data begins with a meticulous and highly structured bioinformatics pipeline, designed to transform raw sequencing reads into biologically meaningful insights about chromatin accessibility, nucleosome organization, and TF binding. This section details the critical steps involved in this pipeline, from preprocessing to data visualization and interpretation.
Following sequencing, the rst essential step is read trimming and QC. The reads generated by ATAC-Seq contain adapter sequences introduced by the transposase enzyme used during library preparation. These adapters must be accurately removed to prevent interference with downstream alignment and peak calling. Tools such as Trimmomatic or Cutadapt are commonly employed to trim these adapters, as well as to remove low-quality bases from the ends of reads. After trimming, QC is performed using programs like FastQC to assess metrics such as per-base quality scores, GC content, and sequence duplication levels. Only high-quality reads that pass these initial checks are retained for further analysis.
The next phase involves aligning the trimmed reads to a reference genome, typically using a high-performance aligner such as Bowtie2 or BWA. Given the short and variable length of ATAC­Seq fragments (ranging from those reecting NFRs to mono-, di-, or tri-nucleosome structures), the aligner must be able to handle such diversity efciently. It is also common practice to lter out reads that map to mitochondrial DNA, as these often represent a disproportionate fraction of the sequencing data and may obscure signals from nuclear chromatin regions. Additionally, duplicate reads resulting from PCR amplication are marked or removed to prevent biases in downstream analyses.
Once the reads are successfully mapped, peak calling is performed to identify regions of open chromatin where the transposase enzyme inserted sequencing adapters. These regions represent accessible DNA and are often indicative of regulatory activity. One of the most widely used tools for this purpose is MACS2 (Model-based Analysis of ChIP-Seq), which has been adapted for ATAC­Seq by treating the aligned fragments as single-end reads centered on the estimated transposase cut sites. MACS2 identies statistically signicant peaks (clusters of reads that are enriched over background noise) suggesting areas of open chromatin, such as enhancers, promoters, or insula­tors. The resulting peak les can then be annotated with genomic features, such as gene promoters or enhancers, using tools like HOMER (Hypergeometric Optimization of Motif EnRichment) or ChIPseeker to contextualize their regulatory relevance.
263 ATAC-Seq for Autoimmune Diseases
To interpret these results visually and explore specic genomic regions in detail, genome brows­ers like the Integrative Genomics Viewer (IGV) or the University of California, Santa Cruz (UCSC) Genome Browser are indispensable. These platforms allow researchers to load their aligned read les (typically in Binary Alignment/Map (BAM) format) and peak les (in Browser Extensible Data (BED) format) to examine patterns of chromatin accessibility across the genome. In IGV, for example, users can zoom into promoter regions of key immune genes to assess the depth and dis­tribution of ATAC-Seq signals, compare accessibility across samples, or overlay other epigenomic datasets, such as ChIP-Seq for histone modications or TFs. UCSC, on the other hand, provides additional annotation layers, including conservation scores, gene models, and ENCODE datasets, which help place the observed chromatin accessibility into a broader functional and evolutionary context.
Beyond identifying open chromatin regions, ATAC-Seq data enables deeper analyses of chroma­tin architecture, particularly nucleosome positioning and TF footprinting. The fragment size distri­bution of ATAC-Seq reads inherently reects the structure of chromatin. Short fragments, typically less than 100 bp, indicate NFRs and correspond to sites of direct DNA–protein interaction, such as active TF binding. In contrast, fragments approximately 200 bp in length correspond to mono­nucleosomes, while longer fragments represent di- and tri-nucleosomes, revealing the higher-order structure of chromatin around accessible regions. Tools such as NucleoATAC can model these pat­terns to infer precise nucleosome positions, generating maps that highlight the phased organization of nucleosomes anking active regulatory elements.
TF footprinting, another powerful application of ATAC-Seq, leverages the ne-scale resolution of transposase insertions. When a TF is bound to DNA, it protects a short region from transposase cleavage, creating a local depletion or “footprint” in the accessibility signal anked by regions of high transposase activity. Computational tools such as HINT-ATAC and TOBIAS can detect these footprints and match them to known TF binding motifs, revealing the identity and activity of regu­latory proteins operating within specic cell types or disease contexts. In autoimmune diseases, where immune cell function and identity are often dysregulated, such footprinting analyses can uncover key TFs that drive pathological gene expression programs.
Taken together, the ATAC-Seq data processing pipeline not only facilitates the identication of open chromatin landscapes but also enables a nuanced understanding of the underlying regulatory architecture. By integrating quality-controlled sequencing data with genome alignment, peak call­ing, visualization, and advanced modeling of nucleosome positions and TF footprints, researchers can reconstruct the epigenetic states that govern immune cell function and contribute to autoim­mune pathogenesis. This layered approach provides a powerful framework for linking chromatin accessibility to gene regulation in health and disease.
8.6.1 RAW DATA ACQUISITION
In a typical ATAC-Seq experiment, raw data acquisition begins in the laboratory, where accessible regions of chromatin are tagged using a hyperactive Tn5 transposase, followed by high-throughput sequencing of the resulting fragments. These raw sequencing reads (usually in FASTQ (FAST Quality) format) are generated directly from the sequencer and represent the initial input for down­stream bioinformatics analysis. However, for demonstration purposes, rather than generating data in a wet-lab setting, we will obtain publicly available raw ATAC-Seq data from the NCBI Sequence Read Archive (SRA). The SRA is a comprehensive repository of sequencing data submitted by researchers worldwide and offers a valuable resource for practicing computational workows, reproducing published ndings, or benchmarking analysis pipelines. Through careful selection of relevant datasets, such as those derived from autoimmune disease studies, we can simulate a com­plete ATAC-Seq analysis starting from real, biologically meaningful raw data.
You can use the “search _ ATAC _ seq.py” Python script, provided as supplementary mate- rial, to search the NCBI BioProject database for raw ATAC-Seq data. This script utilizes the NCBI