Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
84 Bioinformatics of Autoimmune Diseases
FIGURE 3.3 Barcoding process of the single-cell mRNA.
molecule identity. The captured mRNA is then subjected to reverse transcription, during which
cDNA is synthesized using the oligo(dT) primers, incorporating the barcodes and UMIs. This
results in barcoded cDNA libraries that retain information about the cell and the original mRNA
molecule from which each sequence was derived. These libraries are then pooled and subjected to
high-throughput sequencing, enabling the reconstruction of individual transcriptomes from thousands of cells in a single experiment.
In scRNA-Seq, barcoding is a critical step that enables the tracing of gene expression proles
back to individual cells. Each cell is uniquely labeled using a distinct molecular barcode embedded
in synthetic oligonucleotides afxed to microbeads. When a single cell is encapsulated in a nanoliter
droplet alongside a bead, the bead’s barcoded oligos capture the full complement of mRNA transcripts released upon cell lysis, tagging each transcript with a cell-specic identier (Zheng etal.,
2017). This tagging preserves the identity of transcripts throughout downstream processing and
ensures accurate assignment of gene expression data to individual cells.
To achieve scalability, droplet-based microuidic systems such as those developed by 10×
Genomics are employed. These systems rely on a vast pool of pre-synthesized barcodes—generated
through combinatorial synthesis, allowing millions of distinct barcodes to be used in a single experiment. This high barcode complexity ensures that each droplet likely receives a unique barcode,
enabling simultaneous proling of thousands to millions of cells.
After barcoding, the captured mRNA molecules are reverse transcribed into cDNA using reverse
transcriptase enzymes. The resulting cDNA is then amplied using PCR to increase yield. During
this step, sequencing adapters are added to facilitate compatibility with high-throughput sequencing platforms. Rigorous library preparation is crucial for reducing background noise, minimizing
amplication bias, and enhancing sensitivity in transcript detection (Picelli etal., 2014).
3.2.5.4 High-Throughput Sequencing
Following library preparation, single-cell cDNA libraries are sequenced-using high-throughput
NGS platforms, such as Illumina NovaSeq, HiSeq, or NextSeq. The selection of a sequencing platform is typically guided by study objectives, including the desired read length, read depth, and
overall project budget (Zheng etal., 2017).
Read depth is a critical parameter in scRNA-Seq, directly inuencing the resolution and sensitivity
of gene expression detection. For standard transcriptomic proling, approximately 50,000–100,000
reads per cell are generally sufcient. However, applications requiring detection of low-abundance
transcripts, alternative splicing events, or rare cell types may necessitate deeper sequencing, sometimes exceeding 1 million reads per cell. While higher sequencing depths improve the detection of
rare or weakly expressed genes, they also signicantly increase computational and nancial costs.

85 Bioinformatics Design for Autoimmune Research
Sequencing may be performed using either paired-end or single-end reads, with read lengths
typically ranging from 50 to 150 bp. Paired-end sequencing offers enhanced resolution of transcript
isoforms and more accurate alignment, while single-end sequencing is more cost-effective but may
provide limited information for complex transcript structures.
The raw output is stored in FASTQ format, which includes both the nucleotide sequence and the
associated base quality scores. Before downstream analysis, these data undergo preprocessing steps
such as adapter trimming, quality ltering, and removal of low-complexity reads to ensure high data
integrity and reduce technical noise.
3.2.5.5 Data Processing and Quality Control
Once sequencing data is obtained, a series of preprocessing steps are performed to ensure the generation of high-quality, artifact-free transcriptomic proles. The rst step typically involves quality
assessment using tools such as FastQC, which evaluate sequencing metrics, including per-base quality scores, read length distributions, GC content, and potential adapter contamination (Andrews,
2010). Reads containing adapter sequences or low-quality bases are trimmed using tools such as
Cutadapt or Trim Galore, which help reduce alignment errors and improve downstream analyses
(Mar tin, 2011).
Following trimming, ltered reads are aligned to a reference genome (e.g., GRCh38 for humans)
using spliced aligners such as STAR or HISAT2, which are designed to accommodate exon–exon
junctions in RNA-Seq data (Dobin etal., 2013). Alternatively, pseudoalignment tools like Kallisto
can be used for transcript-level quantication with improved computational efciency. In scRNASeq workows, UMIs are used to distinguish original transcripts from PCR duplicates, thereby
enhancing the accuracy of transcript quantication.
After alignment and quantication, cell-level QC is performed to identify and exclude lowquality cells. Metrics commonly assessed include the proportion of mitochondrial gene expression,
total number of detected genes, and overall read complexity. Cells exhibiting high mitochondrial
content, low gene counts, or excessive dropout rates (indicative of RNA degradation or insufcient
capture) are ltered out to ensure the integrity of the dataset and reliability of subsequent analyses.
3.2.5.6 Clustering, Differential Expression, and Functional Analysis
Following preprocessing, scRNA-Seq data undergoes a series of computational analyses to identify cellular subpopulations and elucidate transcriptional alterations associated with disease states.
Dimensionality reduction techniques, such as PCA and Uniform Manifold Approximation and
Projection (UMAP), are employed to project high-dimensional gene expression data into lowerdimensional space for effective visualization of cellular heterogeneity. Clustering algorithms, implemented in widely used platforms such as Seurat, SCANPY, and Monocle, categorize individual cells
into transcriptionally distinct subpopulations based on shared expression patterns.
To uncover disease-relevant gene regulation, differential gene expression analysis is conducted
using statistical tools such as DESeq2, edgeR, or limma. These frameworks identify genes that
are signicantly upregulated or downregulated in specic immune cell subsets, thereby revealing
molecular signatures of autoimmune pathogenesis (Love etal., 2014). The resulting gene lists are
subjected to pathway and functional enrichment analysis using databases like GO and KEGG to
identify disrupted biological pathways, including those related to cytokine signaling, antigen processing, or lymphocyte activation.
Beyond gene-level expression, scRNA-Seq enables the reconstruction of cell–cell communication networks through tools such as CellPhoneDB and NicheNet, which infer ligand–receptor interactions across immune cell types. These analyses provide critical insight into how immune cells
coordinate responses and drive inammation in autoimmune conditions. Furthermore, pseudotime
trajectory analysis using methods like Monocle allows researchers to model dynamic cellular transitions and differentiation processes, thereby shedding light on the progression of immune activation
and dysfunction over time.

86 Bioinformatics of Autoimmune Diseases
3.2.6 EPIGENOME-WIDE ASSOCIATION STUDIES
Epigenome-wide association studies (EWAS) are large-scale analyses that explore genome-wide
epigenetic modications to identify associations with disease phenotypes, including autoimmune
disorders. Unlike GWAS, which investigate inherited genetic variants such as SNPs, EWAS focuses
on reversible and dynamic epigenetic changes (including DNA methylation, histone modications,
chromatin accessibility, and non-coding RNA regulation) that modulate gene expression without
altering the underlying DNA sequence (Michels etal., 2013).
Autoimmune diseases arise from multifaceted interactions between genetic susceptibility, environmental exposures, and immune dysregulation. Epigenetic mechanisms serve as key intermediaries that integrate environmental signals into stable but modiable changes in gene expression,
especially within immune cells. Dysregulation of epigenetic pathways has been implicated in the
pathogenesis of diseases such as SLE, RA, and MS. EWAS enables researchers to pinpoint epigenetic loci that differ between patients and healthy individuals, providing insight into disease-specic
regulatory mechanisms.
Environmental factors such as infections, smoking, diet, and psychosocial stress have been
shown to inuence epigenetic landscapes and, consequently, immune function (Feil & Fraga, 2012).
By capturing the interface between gene regulation and environment, EWAS holds promise for
identifying predictive biomarkers, understanding disease heterogeneity, and informing personalized therapeutic strategies.
However, EWAS poses unique challenges due to the cell-type specicity and temporal variability of epigenetic marks. Thus, rigorous study design (including proper cohort selection, cell-type
enrichment or deconvolution, and correction for confounding variables) is critical. The standard
EWAS workow includes sample collection, high-quality DNA extraction, epigenetic proling via
methods such as bisulte sequencing or array-based platforms, data normalization, statistical modeling, and biological interpretation using functional annotation tools and pathway analysis (Michels
etal., 2013).
3.2.6.1 Study Design and Sample Selection
The validity and interpretability of EWAS are critically dependent on rigorous study design and
appropriate cohort selection. A well-constructed EWAS typically involves comparing individuals
diagnosed with autoimmune diseases (cases) to matched healthy controls. Given that epigenetic
modications are modulated by numerous confounding factors (such as age, sex, ancestry, environmental exposures, and immune cell composition), these variables must be carefully controlled
or statistically adjusted to minimize bias and increase the specicity of epigenetic associations
(Michels etal., 2013).
Tissue and cell-type specicity represents a fundamental consideration in EWAS design,
as epigenetic marks exhibit distinct proles across different biological contexts. In autoimmune
research, commonly studied sample types include PBMCs, puried immune cell subsets (e.g.,
+
T cells, B cells, and monocytes), as well as disease-relevant tissues such as synovial biop-
CD4
sies in RA, intestinal biopsies in IBD, and cerebrospinal uid in MS. Recent advances in singlecell epigenomics have further enhanced the resolution of these studies, enabling the dissection of
cell-type-specic epigenetic variation and the identication of rare pathogenic cell populations
(Buenrostro etal., 2015; Grosselin etal., 2019).
Precise matching of cases and controls by demographic and clinical factors is essential to reduce
residual confounding. Additionally, detailed environmental exposure data (such as smoking history,
diet, medication use, infections, and psychosocial stressors) should be collected and incorporated into
analytical models, as these factors are known to exert profound and often reversible effects on the
epigenome (Feil & Fraga, 2012). Longitudinal EWAS designs, where participants are followed over
time, are particularly advantageous for distinguishing causal epigenetic changes from those arising
secondarily to disease activity or treatment, thus improving the inference of mechanistic relationships.

87 Bioinformatics Design for Autoimmune Research
3.2.6.2 DNA Extraction and Quality Control
After sample collection, gDNA is extracted from blood or tissue using column-based purication kits or traditional phenol–chloroform extraction methods. The quality and concentration of
the extracted DNA are evaluated using spectrophotometry (NanoDrop), uorometry (Qubit), and
electrophoresis (Bioanalyzer or TapeStation) to ensure high integrity. High-quality DNA is critical
for obtaining accurate and reproducible epigenetic data. DNA degradation can introduce biases in
methylation analysis, so proper storage and handling of samples are necessary to preserve integrity.
3.2.6.3 Epigenetic Proling and High-Throughput Sequencing
Epigenetic proling in EWAS employs a range of high-throughput methodologies to investigate
DNA methylation, histone modications, and other chromatin-based mechanisms that regulate
gene expression. Among these, DNA methylation is the most extensively characterized epigenetic
mark due to its stability, regulatory signicance, and relevance in autoimmune pathogenesis.
Whole-genome bisulte sequencing (WGBS) is considered the gold standard for DNA methylation proling, as it provides comprehensive, single nucleotide resolution across the entire genome.
However, WGBS is cost-intensive and computationally demanding, limiting its routine use in largescale studies. As a cost-effective alternative, reduced representation bisulte sequencing (RRBS)
enriches for CpG-dense regions such as promoters and regulatory elements, enabling focused analysis of functionally relevant loci while reducing sequencing load.
Microarray-based approaches remain widely used for large EWAS cohorts. The Innium
MethylationEPIC BeadChip array (Illumina) interrogates over 850,000 CpG sites genome-wide,
covering gene promoters, enhancers, and intergenic regions. This platform offers high reproducibility, scalability, and cost efciency, making it particularly well-suited for population-based studies
of complex diseases, including autoimmune disorders.
Beyond DNA methylation, EWAS can incorporate histone modication proling to capture additional regulatory dimensions. Chromatin immunoprecipitation followed by sequencing (ChIP-Seq)
is the standard method for mapping post-translational histone marks, such as H3K4me3 (associated
with transcriptional activation) and H3K27ac (associated with active enhancers). ChIP-Seq enables
the genome-wide identication of active and repressive chromatin states, offering insight into the
epigenetic regulation of immune-related genes in autoimmune conditions.
ChIP (Figure 3.4) is a molecular biology technique used to investigate interactions between proteins and DNA within the chromatin context of living cells. The process begins with the crosslinking of proteins to DNA using formaldehyde, which preserves protein–DNA interactions as they
occur in the native cellular environment. The chromatin is then sheared into smaller fragments,
typically by sonication or enzymatic digestion. An antibody specic to the protein of interest (such
as a transcription factor or a modied histone) is used to selectively immunoprecipitate the protein–
DNA complexes. These complexes are then isolated, and the crosslinks are reversed to release the
DNA. The puried DNA fragments represent the genomic regions bound by the protein and can be
identied using qPCR, microarray (ChIP-chip), or high-throughput sequencing (ChIP-Seq). ChIP
allows researchers to map the binding sites of regulatory proteins across the genome and study the
epigenetic landscape, including histone modications and transcription factor occupancy, providing
critical insights into gene regulation and chromatin dynamics.
These histone marks help elucidate chromatin dynamics and gene regulation in autoimmune diseases. Another technique, Assay for Transposase-Accessible Chromatin using sequencing (ATACSeq ), is used to study chromatin accessibility, which identies active regulatory regions in immune
cells.
Transposase-accessible chromatin (Figure 3.5) refers to regions of the genome that are open
and not tightly packed by nucleosomes, making them more accessible to regulatory proteins such
as transcription factors. These regions are functionally important because they often correspond to
active promoters, enhancers, and other cis-regulatory elements that control gene expression. The

88 Bioinformatics of Autoimmune Diseases
FIGURE 3.4 Chromatin immunoprecipitation.
concept is central to the Assay for Transposase-Accessible Chromatin using sequencing (ATACSeq), which employs a hyperactive Tn5 transposase to simultaneously cut and tag accessible DNA
regions with sequencing adapters. The resulting fragments are then amplied and sequenced, allowing researchers to map open chromatin landscapes genome-wide. Because accessible chromatin is a
hallmark of active regulatory elements, studying these regions provides insight into cellular identity,
transcriptional regulation, and the dynamic changes that occur in development or disease.
FIGURE 3.5 Transposase-accessible chromatin.

89 Bioinformatics Design for Autoimmune Research
Following immunoprecipitation in a ChIP-Seq experiment, the next crucial step involves reversing the cross-links between DNA and proteins. This is typically achieved by heating the sample, which breaks the formaldehyde-induced covalent bonds that were originally used to preserve
protein–DNA interactions. Once the cross-links are reversed, the sample is treated with proteinase
K to digest the proteins and liberate the DNA fragments. The DNA is then puried using phenolchloroform extraction or column-based purication methods. At this stage, the puried DNA represents genomic regions that were bound by the protein of interest. This enriched DNA is then
used to construct a sequencing library, which involves end repair, A-tailing, adaptor ligation, and
PCR amplication. The resulting DNA library is nally subjected to high-throughput sequencing to
determine the genomic locations where the protein was bound.
3.2.6.4 Peak Enrichment
ChIP-Seq is a widely used technique for mapping the genome-wide binding sites of DNA-associated
proteins and histone modications. Peak enrichment in ChIP-Seq refers to the identication of
genomic regions with a statistically signicant accumulation of aligned sequencing reads, indicating probable sites of protein–DNA interactions or histone modications.
The experimental workow begins by crosslinking protein–DNA complexes in cells, followed
by chromatin fragmentation and immunoprecipitation using an antibody specic to the target protein or histone mark. The associated DNA fragments are puried and subjected to high-throughput
sequencing. The resulting short reads are mapped to a reference genome (e.g., GRCh38), and regions
with clusters of aligned reads (termed peaks) represent loci where the protein or histone mark is
most likely enriched.
To determine the signicance of enrichment (see Figure 3.6), read densities in the ChIP sample
are compared against background signals obtained from control samples, such as input DNA (which
accounts for genomic representation) or IgG controls (which account for non-specic antibody
binding). Computational tools like MACS (Model-based Analysis for ChIP-Seq) use a dynamic
Poisson distribution to distinguish true peaks from background noise, adjusting for local biases in
FIGURE 3.6 ChIP-Seq peak enrichment.

90 Bioinformatics of Autoimmune Diseases
=
(
µij˙
)
ij
µ=˙+˙
()
i 1
0 i
the genome. Alternative algorithms like SICER are better suited for identifying diffuse signals,
such as broad histone modications.
The height and sharpness of a peak correlate with the degree of protein–DNA binding or modication frequency. Highly enriched peaks suggest strong and stable interactions, whereas broader
or weaker peaks may indicate more transient or dispersed binding. Accurate peak calling is critical
for downstream analyses, including motif discovery, co-binding analysis, and integration with gene
expression or chromatin accessibility data.
Enriched peaks are characterized not only by their height (reecting the number of overlapping reads) but also by their shape and width, which can give insights into the nature of the binding event. For instance, transcription factors typically generate sharp, narrow peaks near promoter
regions or enhancers, while histone modications like H3K27me3 often produce broader, more diffuse peaks across large genomic domains. The identication and characterization of these enriched
peaks allow researchers to infer the regulatory landscape of the genome, uncovering transcriptional
networks, enhancer-promoter interactions, epigenetic regulation, and chromatin states in various
biological conditions.
3.2.6.5 Statistical Analysis
Following peak identication in ChIP-Seq experiments, a critical next step is differential binding
analysis, which aims to detect statistically signicant differences in peak signal intensities between
experimental conditions (e.g., treatment vs. control). This comparison is analogous to differential
gene expression analysis in RNA-Seq and relies on modeling the distribution of read counts mapped
to each peak across samples.
Read counts, dened as the number of sequencing reads aligned to a particular peak region, are
treated as discrete count data. However, unlike a simple Poisson model, which assumes that the
mean equals the variance, ChIP-Seq data often exhibits overdispersion, where the variance exceeds
the mean. To accommodate this biological variability, the negative binomial distribution is commonly employed. This model introduces a dispersion parameter that accounts for extra variability
not captured by the Poisson framework, making it more appropriate for high-throughput sequencing
data (Love etal., 2014).
In a differential binding analysis workow:
• Each peak is treated as a feature, similar to a gene in RNA-Seq analysis.
• For each biological replicate, the number of reads overlapping a peak is quantied.
• A negative binomial model is tted to the count data, estimating the expected read count
for each peak under each experimental condition.
• The dispersion parameter is estimated for each peak individually, capturing variability
across replicates within the same group (Anders & Huber, 2010).
• Statistical testing is performed to assess whether observed differences in read counts
across conditions are signicant, often using software packages such as DESeq2, edgeR,
or csaw, which are adapted for ChIP-Seq data.
Mathematically, for a given peak i and sample j, the model assumes
NB
ij
i
where ij is the observed read count, ij is the expected mean count (dependent on the experimental
condition and sequencing depth), and i is the dispersion parameter for peak i, capturing the biological and technical variance.
The expected counts are modeled using a GLM framework with a log link function:
g
X (3.12)
jij
(3.11)

FIGURE 3.7 The owchart of statistical analysis and modeling in ChIP-Seq.
0i
1i
X
j
91 Bioinformatics Design for Autoimmune Research
where
treatment), and
is the intercept (baseline count for the peak),
is a binary indicator for the condition (0 for control, 1 for treated).
represents the effect of the condition (e.g.,
To determine if a peak is differentially enriched, the null hypothesis:
H0: β1i = 0 (no difference between conditions)
This hypothesis is tested in DESeq2 using Wald test to compute p-values for each peak or in
edgeR using a LRT that compares the full model to a reduced model without the condition term.
The p-values are then adjusted using the Benjamini-Hochberg procedure to control the FDR,
producing q-values. Peaks with adjusted p-values below a certain threshold (commonly 0.05) are
considered signicantly differentially bound.
This model-based statistical framework (Figure 3.7) enables researchers to go beyond peak presence/absence and assess quantitative differences in binding intensity between conditions. This is
especially critical in experiments investigating stimulus-dependent changes in transcription factor
binding, histone modication states, or epigenetic regulators under different biological contexts.
3.2.6.6 Functional Enrichment after Peak Enrichment in ChIP-Seq
After signicant peaks are identied through differential binding analysis, the next step involves
interpreting their biological relevance using functional enrichment analysis. This process typically
begins by associating peaks with nearby genes, often using genomic proximity to transcription start
sites or chromatin conformation data such as Hi-C or ChIA-PET to account for distal regulatory
interactions. Once peaks are annotated to genes, the associated gene list is analyzed for enrichment
of biological functions, pathways, or regulatory motifs.
Tools such as HOMER, GREAT (Genomic Regions Enrichment of Annotations Tool), and
ChIPseeker facilitate this mapping and enrichment process. These tools evaluate whether genes
associated with peak regions are overrepresented in predened functional categories, including GO
terms, KEGG pathways, and Reactome pathways. In addition, motif enrichment analysis can identify overrepresented DNA-binding motifs within peaks, offering insights into the sequence specicity and potential co-factors of the immunoprecipitated protein.
This layer of analysis provides a systems-level understanding of how chromatin-associated
proteins—such as transcription factors, histone modiers, or chromatin remodelers—regulate gene
expression in a given biological context. In the study of autoimmune diseases, such functional insights
can reveal critical immune pathways or cellular differentiation programs inuenced by epigenetic regulation. Ultimately, functional enrichment adds interpretive depth to ChIP-Seq datasets by converting
raw genomic coordinates into biologically meaningful hypotheses and candidate regulatory networks.

92 Bioinformatics of Autoimmune Diseases
3.2.7 METAGENOMIC AND MICROBIOME SEQUENCING
Metagenomic and microbiome sequencing are high-throughput technologies that enable comprehensive analysis of microbial communities residing in diverse environments, including the human
body. In particular, the human gut microbiome has emerged as a critical regulator of immune system development and homeostasis. Increasing evidence suggests that disruption of microbial balance, or dysbiosis, is associated with the pathogenesis of various autoimmune diseases, such as RA,
SLE, MS, T1D, and IBDs, including Crohn’s disease and ulcerative colitis.
The microbiome modulates immune responses through several interrelated mechanisms.
Microbial metabolites, such as short-chain fatty acids (SCFAs), regulate inammatory pathways
and promote the differentiation of regulatory T cells (Arpaia etal., 2015). In some cases, microbial
proteins exhibit molecular mimicry, where structural similarities with host proteins can lead to the
activation of autoreactive immune cells (Scher etal.,2012). Dysbiosis-induced gut barrier dysfunc-
tion (commonly referred to as “leaky gut”) increases intestinal permeability, allowing microbial
antigens to translocate into systemic circulation and trigger immune activation. Furthermore, host
genetic susceptibility inuences microbiome composition and its immunomodulatory capacity,
underscoring the complex host–microbe interplay in autoimmune pathogenesis.
Through metagenomic and metatranscriptomic sequencing, researchers can assess taxonomic
diversity, gene content, and metabolic potential of microbial communities in disease versus healthy
states. These insights facilitate the identication of microbial biomarkers for early diagnosis, inform
dietary and probiotic strategies, and support therapeutic approaches such as fecal microbiota transplantation (FMT) and microbiome-based immunomodulation. As our understanding deepens,
microbiome proling is increasingly integrated into precision medicine efforts aimed at mitigating
immune dysregulation in autoimmune diseases.
3.2.7.1 Study Design and Sample Collection
A rigorous and thoughtfully designed study is essential to generate reliable and biologically meaningful insights in microbiome research. The rst step involves the selection of well-dened study
cohorts, including individuals diagnosed with autoimmune diseases and appropriately matched
healthy controls. Matching participants based on age, sex, ancestry, and environmental exposures
minimizes confounding variables that can obscure true microbial differences. In some studies, a
family-based design is adopted to identify heritable microbial signatures by comparing microbiome
proles between affected individuals and their unaffected relatives.
The choice of sample type is dictated by the specic autoimmune condition under investigation.
Fecal samples are most commonly used for proling the gut microbiome, which plays a central role
in systemic autoimmune conditions such as IBD and RA (Scher etal., 2012). Oral swabs are relevant
for diseases like Sjögren’s syndrome, which involves mucosal immune dysfunction. For examining
systemic microbial inuence, blood microbiome analysis can reveal microbial translocation and its
immunological consequences. In dermatological autoimmune diseases, such as psoriasis and cutaneous lupus erythematosus, skin swabs are used to study localized microbial alterations.
Preserving the integrity of microbial communities during and after sample collection is critical.
Standard protocols involve immediate freezing at −80°C or stabilization in nucleic acid-preserving
reagents, such as RNAlater or OMNIgene, to prevent microbial degradation and preserve the native
composition of the microbiome. These precautions ensure reproducibility and consistency across
samples, enabling robust downstream metagenomic analysis.
3.2.7.2 DNA/RNA Extraction and Quality Control
Once samples are collected, microbial DNA or RNA is extracted using specialized kits designed
to efciently recover diverse microbial species while minimizing contamination. High-quality
nucleic acids are crucial for sequencing reliability. DNA quality is assessed using spectrophotometry (NanoDrop) to measure purity, uorometry (Qubit) for concentration quantication, and

93 Bioinformatics Design for Autoimmune Research
electrophoresis (Bioanalyzer or TapeStation) to evaluate integrity. For RNA-based studies, rRNA
depletion is performed to enrich microbial mRNA for metatranscriptomic analysis.
3.2.7.3 Microbiome and Metagenomic Sequencing Approaches
Microbiome sequencing encompasses a range of methodologies tailored to the specic goals of the
study and the resolution of taxonomic or functional information required. One of the most widely
employed methods is 16S rRNA gene sequencing, which targets conserved regions of the bacterial
16S rRNA gene using PCR. Amplicons are sequenced-using platforms such as Illumina MiSeq or
NovaSeq, followed by taxonomic classication using curated databases such as SILVA, Greengenes,
or the Ribosomal Database Project (RDP). This method is cost-effective and highly efcient for
proling bacterial diversity but is limited by its inability to resolve strain-level differences or detect
non-bacterial organisms, such as fungi and viruses.
To obtain broader taxonomic coverage and functional insights, researchers increasingly employ
whole-genome shotgun sequencing (WGSS metagenomic sequencing, which involves random fragmentation and deep sequencing of all microbial DNA in a sample. This technique enables strainlevel resolution and the identication of bacteria, archaea, fungi, viruses, and their genetic functions.
Sequencing platforms such as Illumina NovaSeq, PacBio, or Oxford Nanopore are used, followed
by bioinformatic analyses using tools such as Kraken2, MetaPhlAn, and HUMAnN to determine
taxonomic composition and metabolic pathway activity (Beghini etal., 2021). Although WGS offers
comprehensive insights, it is computationally intensive and requires higher sequencing depth, which
can increase cost and complexity.
Metatranscriptomics, or microbial RNA-Seq, adds a dynamic layer by capturing the active transcriptional activity of microbial communities. After RNA extraction, rRNA is depleted to enrich for
mRNA, which is then reverse transcribed into cDNA and sequenced-using high-throughput platforms. This approach enables the identication of metabolically active taxa and functional pathways
under different physiological or pathological conditions. However, RNA degradation and sample
complexity pose analytical challenges, requiring robust protocols and QC measures.
Specialized techniques are applied to specic microbial domains. Internal Transcribed Spacer
(ITS) sequencing is used for proling fungal communities, while viral metagenomics relies on
WGSS to characterize virome composition and diversity. These additional approaches allow for a
more comprehensive understanding of microbial interactions, including potential roles in immune
modulation and autoimmunity.
3.2.7.4 Data Processing and Bioinformatics Analysis
Once microbiome sequencing is completed, raw sequencing data undergoes a comprehensive bioinformatics pipeline to identify microbial taxa and infer functional capabilities. The process begins
with quality assessment of the raw reads using tools such as FastQC, which evaluates parameters like read length distribution, per-base quality scores, GC content, and adapter contamination
(Andrews, 2010). Low-quality bases and sequencing adapters are removed using preprocessing
tools such as Trim Galore and Cutadapt (Marti n, 2011), ensuring that only high-quality reads proceed to downstream analysis.
Taxonomic classication is performed using sequence alignment or marker gene-based
approaches. Tools like QIIME2 and Mothur utilize 16S rRNA or ITS sequences and classify microbial taxa using curated databases such as SILVA or Greengenes. For metagenomic data, Kraken2
and MetaPhlAn enable rapid and accurate taxonomic proling by comparing reads to reference
databases or clade-specic marker genes (Beghini etal., 2021). Read alignment to microbial refer-
ence genomes may also be performed using Bowtie2, particularly in strain-level analysis or custom
genome investigations.
Functional characterization of the microbiome involves gene annotation using tools such as
HUMAnN (The HMP Unied Metabolic Analysis Network), which links microbial gene content
to metabolic pathways through databases such as KEGG, UniRef, and MetaCyc. These annotations
Соседние файлы в папке Библиотека им академика М.И. Перельмана
