Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
74 Bioinformatics of Autoimmune Diseases
100,000, 000 × 150
3000,000, 000
For example, if we sequence 100 million reads of 150 bp each for a genome of 3 billion bp
(human genome):
Depth = =×5coverage (3.3)
The sequencing depth is usually referred to as “×” coverage as the “×” symbol means “times”;
so 5× coverage means each base has been sequenced ve times on average.
The required depth depends on the study’s goals:
• 30× depth (standard coverage) is sufcient for detecting germline variants.
• 60–100× depth is used for detecting somatic mutations in diseases like cancer.
• 10× depth (low-coverage WGS) is sometimes used in large-scale population studies for
cost-effective variant discovery.
Short-read sequencing (e.g., Illumina) is widely used due to its high accuracy and cost efciency,
while long-read sequencing (e.g., PacBio, Oxford Nanopore) is benecial for detecting structural
variants, repetitive elements, and phased haplotypes.
Raw sequencing data is generated in FASTQ format, containing nucleotide sequences and quality scores. These data undergo preprocessing steps such as adapter trimming, quality ltering, and
base correction to remove low-condence reads.
3.2.3.4 Read Alignment and Variant Calling
After sequencing, the raw reads are aligned to a reference genome (e.g., GRCh38 or GRCh37/hg19)
using alignment tools such as Burrows–Wheeler Aligner (BWA) or Bowtie2. Accurate alignment
ensures proper variant detection, and alignment quality is assessed using mapping rate, duplication
rate, and coverage uniformity metrics.
Once aligned, variant calling is performed to identify genetic differences between the study
participants and the reference genome. Variant callers such as
• GATK (Genome Analysis Toolkit): the gold standard for germline variant calling.
• DeepVariant: a deep learning-based variant caller for improved accuracy.
• Lumpy or Manta: for structural variant detection.
are used to detect different classes of variants, including SNVs, indels, CNVs, and SVs.
Post-calling, variants are ltered based on read depth, genotype quality, and allele frequency
to eliminate false positives. Variants are then annotated using reference databases such as d b SN P,
1000 Genomes, and gnomAD to determine their population frequencies and potential pathogenicity.
3.2.3.5 Functional Annotation and Pathway Analysis
Once genetic variants have been identied through whole-genome or exome sequencing, the next
essential step is to assess their potential biological relevance. Functional annotation tools such as
ANNOVAR, VEP, and SnpEff are widely used to predict the consequences of variants on gene function, protein structure, and regulatory elements (McLaren etal., 2016). Exonic variants are exam-
ined for their effects on protein-coding sequences, including missense, nonsense, and frameshift
mutations, which may alter amino acid composition or lead to truncated, nonfunctional proteins.
Conversely, variants in non-coding regions are evaluated for their inuence on gene regulation,
messenger RNA (mRNA) splicing, and enhancer activity.
To prioritize variants likely to contribute to disease, researchers utilize pathogenicity prediction algorithms such as PolyPhen-2, SIFT, and CADD. These tools integrate evolutionary conservation, biochemical properties, and multiple annotation features to estimate the likelihood that a

75 Bioinformatics Design for Autoimmune Research
given variant is deleterious (Adzhubei etal., 2010). Moreover, variants located in regulatory regions
are interpreted using epigenomic resources such as the ENCODE project and the GTEx database,
which provide insights into tissue-specic gene expression and regulatory element activity, particularly in immune-relevant tissues.
Beyond single-variant interpretation, pathway and network analysis methods are applied to place
genetic ndings into a systems biology context. Tools such as Kyoto Encyclopedia of Genes and
Genomes (KEGG), Reactome, and IPA are used to determine whether disease-associated variants
are enriched in known biological pathways. This is particularly important in autoimmune diseases,
which frequently involve dysregulation in pathways related to cytokine signaling, T-cell activation,
antigen presentation, and NF-κB signaling. By mapping variants to these immune-related pathways, researchers can uncover underlying disease mechanisms and identify potential targets for
therapeutic intervention.
3.2.4 RNA SEQUENCING
RNA-Seq is a high-throughput NGS technology that enables comprehensive proling of gene
expression across the entire transcriptome. Unlike traditional microarray-based approaches, RNASeq offers greater sensitivity, a broader dynamic range, and the ability to detect novel transcripts
without requiring prior sequence information. This technology allows for the quantication of gene
expression, detection of alternative splicing events, and characterization of various RNA species,
including long non-coding RNAs (lncRNAs), miRNAs, and circular RNAs (circRNAs), providing
a multifaceted view of transcriptomic regulation.
In the context of autoimmune diseases, RNA-Seq has become an indispensable tool for uncovering the molecular mechanisms underlying immune dysregulation. Autoimmune conditions such
as SLE, RA, and MS are driven by aberrant immune responses that result in chronic inammation
and tissue destruction. By comparing gene expression proles between affected individuals and
healthy controls, researchers can identify differentially expressed genes (DEGs) and dysregulated
pathways central to disease pathogenesis. RNA-Seq also facilitates the identication of immune
cell-specic signatures, cytokine expression patterns, and transcriptional networks associated with
disease activity and progression.
Moreover, RNA-Seq provides valuable insights into how environmental exposures (such as infections, stress, or dietary factors) interact with genetic predispositions to inuence gene expression
in autoimmune settings. This integrative approach helps to identify biomarkers for early diagnosis,
stratify patients by molecular subtypes, and inform precision medicine strategies by predicting
treatment responses.
The RNA-Seq workow includes critical steps such as sample collection, RNA extraction, quality assessment, library construction, sequencing, alignment to reference genomes, and downstream
analysis. Differential expression analysis is typically performed using statistical tools such as
DESeq2 or edgeR, followed by functional enrichment analysis to interpret the biological signicance of observed transcriptional changes (Love etal., 2014). Each step requires stringent QC and
bioinformatics expertise to ensure the generation of accurate and reproducible data.
3.2.4.1 Study Design and Sample Collection
A well-designed RNA-Seq study begins with the careful selection of an appropriate cohort, typically composed of individuals diagnosed with autoimmune diseases and demographically matched
healthy controls. Because gene expression is affected by numerous variables (such as age, sex, medication history, disease severity, and environmental exposures), these confounding factors must be
rigorously controlled during study design to ensure that observed transcriptional differences reect
disease biology rather than technical or demographic artifacts (Conesa etal., 2016). Matching par-
ticipants and incorporating metadata in the statistical analysis helps mitigate these confounding
effects and improves interpretability.

76 Bioinformatics of Autoimmune Diseases
Standardizing sample collection procedures is essential to minimize batch effects, which can
obscure biological signals and reduce reproducibility. Factors such as collection time, storage
conditions, and RNA extraction protocols must be kept consistent across all samples. Even small
deviations in processing can introduce technical variability that signicantly impacts transcript
quantication.
RNA-Seq can be applied to a variety of biological sources depending on the research objective.
Peripheral blood mononuclear cells (PBMCs), sorted immune cell subsets (e.g., T cells, B cells,
monocytes), and whole blood are commonly used in systemic autoimmune diseases. Site-specic
samples such as synovial uid (for RA), intestinal biopsies (for inammatory bowel disease (IBD)),
or cerebrospinal uid (for MS) provide localized insights into tissue-specic immune responses.
The selection of tissue type should align with the pathophysiological features of the disease being
studied.
To maintain RNA integrity, immediate stabilization is required after sample collection. Reagents
such as TRIzol or RNAlater are commonly used to prevent RNA degradation by RNases. If immediate processing is not feasible, samples are rapidly frozen and stored at −80°C to preserve RNA
quality. Given RNA’s susceptibility to enzymatic degradation, meticulous handling and storage procedures are paramount to ensure the generation of reliable sequencing data.
3.2.4.2 RNA Extraction and Quality Control
Total RNA is extracted from the collected samples using commercial kits (e.g., Qiagen RNeasy,
TRIzol reagent) or column-based purication methods. The extracted RNA is then assessed for
integrity and purity using spectrophotometry (e.g., NanoDrop), uorometry (e.g., Qubit), and electrophoresis (e.g., Agilent Bioanalyzer).
The RNA integrity number (RIN) is used to evaluate RNA quality, with a RIN ≥7 considered
suitable for RNA-Seq. High-quality RNA is essential for generating reliable sequencing data, as
degraded RNA can lead to biased or incomplete transcriptome proles.
Depending on the research objectives, different RNA fractions may be analyzed:
• Total RNA sequencing captures all RNA species, including coding and non-coding RNAs.
• Poly(A)+ RNA sequencing enriches for mRNA, excluding ribosomal RNA (rRNA) and
most non-coding RNAs.
• rRNA-depleted sequencing is used for analyzing non-coding RNAs, such as long non-
coding RNAs (lncRNAs) and circRNAs.
• Small RNA sequencing focuses on small RNAs like microRNAs (miRNAs), which regulate immune responses.
3.2.4.3 Library Preparation and RNA Fragmentation
RNA-Seq library preparation consists of a series of precise molecular steps aimed at converting
RNA molecules into complementary DNA (cDNA) and attaching sequencing adapters to enable
high-throughput sequencing. The specic methodology depends on the type of RNA being targeted
(most commonly mRNA or total RNA).
For mRNA-focused studies, the rst and most critical step is poly(A) selection, which enriches
for mature mRNA transcripts by isolating RNA molecules that possess polyadenylated [poly(A)]
tails. These tails are a hallmark of most eukaryotic mRNAs and can be selectively captured using
oligonucleotides composed of deoxythymidine [oligo(dT)] bound to magnetic beads. After incubation of total RNA with oligo(dT)-conjugated beads, a series of wash and elution steps removes
rRNA and the majority of non-coding RNAs, resulting in an mRNA-enriched fraction ready for
downstream processing.
In contrast, total RNA-Seq enables a more comprehensive view of the transcriptome, including
both coding and non-coding RNA species such as lncRNAs and circRNAs. Because rRNA constitutes over 80% of total RNA, its removal is essential for capturing low-abundance transcripts.

77 Bioinformatics Design for Autoimmune Research
Total numberofreadReadlength
×
Sequencin
)
FIGURE 3.1 Workow of mRNA enrichment via poly(A) selection.
Depletion strategies such as Ribo-Zero or NEBNext rRNA depletion kits are commonly employed
for this purpose.
Following RNA enrichment, either by poly(A) selection or rRNA depletion, the RNA is fragmented into smaller segments, typically ranging from 200 to 300 bp, using enzymatic or heat-based
methods (see Figure 3.1). This fragmentation is necessary to ensure uniform read coverage across
the transcript and to facilitate accurate transcript reconstruction during bioinformatics analysis
(Conesa etal., 2016).
The fragmented RNA is then reverse transcribed into cDNA using random primers and reverse
transcriptase enzymes. Sequencing adapters are ligated to both ends of the cDNA fragments to
enable their amplication and subsequent detection by NGS platforms. In many modern protocols,
UMIs are incorporated during adapter ligation or PCR amplication. UMIs allow researchers to
track individual RNA molecules, helping correct for PCR duplicates and improve transcript quantication accuracy.
3.2.4.4 High-Throughput Sequencing
The prepared RNA-Seq libraries are sequenced-using high-throughput NGS platforms such as
Illumina NovaSeq, HiSeq, or NextSeq. The selection of sequencing platform depends on factors
such as sequencing depth, read depth, read length, and study objectives.
In RNA-Seq, sequencing depth (coverage) is determined by the size of the transcriptome, rather
than the genome size as in WGS.
gDepth X
The human transcriptome size refers to the total number of nucleotides in all RNA molecules
expressed in a cell or tissue at a given time. Unlike the genome, which is xed at about 3.2 billion
bp, the transcriptome is dynamic and varies by cell type, condition, and study purpose. In studies
focusing only on protein-coding genes, the transcriptome size is estimated at around 50 million
bp, based on roughly 20,000 genes with an average transcript length of 2500 bp. When non-coding
RNAs such as lncRNAs and microRNAs (miRNAs) are included, the estimated size increases to
60–80 million bp, depending on the annotation database used (e.g., GENCODE, RefSeq). However,
in typical RNA-Seq experiments, only 10,000–15,000 genes are expressed in a sample, resulting in
an effective transcriptome size of about 25–40 million bp. This estimate is essential for planning
appropriate sequencing depth.
In RNA-Seq experiments, it is more appropriate to report sequencing output as reads per sample rather than average sequencing depth (X coverage). Unlike whole-genome sequencing, where
coverage is uniformly distributed across a xed genome size, RNA-Seq targets the transcriptome
(adynamic and variable subset of the genome that differs between cell types, conditions, and even
=
()
Totallengthofall transcriptsTranscriptome( size
(3.4)

78 Bioinformatics of Autoimmune Diseases
TABLE 3.1
Typical RNA-Seq Read Depth Recommendations
Experimental Goal Recommended Read Depth Comments
Gene expression proling (standard) 20–30 million reads/sample Detecting moderately to highly expressed genes
Differential gene expression (DGE) 30–50 million reads/sample Higher depth improves statistical power
Low-abundance transcript detection
Transcript isoform/splicing analysis
Single-cell RNA-Seq 20,000–100,000 reads per cell Depends on cell type and transcript complexity
De novo transcriptome assembly
≥50 million reads/sample
≥75–100 million reads/sample
≥100 million reads per sample
Increases sensitivity to weakly expressed genes
Required for resolving alternative splicing or
isoforms
Necessary due to lack of a reference and high
diversity
individual samples). Because transcript abundance varies widely, sequencing reads are not evenly
distributed across all genes, making average coverage a less meaningful measure. Reporting the
number of reads per sample (read depth) provides a more direct and practical metric for assessing
data quantity, evaluating experiment consistency, and planning the required depth for capturing
transcript diversity and detecting differential expression. This approach aligns more closely with
how RNA-Seq data is generated and analyzed in practice.
Read depth is a critical factor in RNA-Seq experiments. Standard RNA-Seq requires approximately 10–50 million reads per sample for gene-level expression analysis, while deeper sequencing
(>100 million reads per sample) is needed for detecting alternative splicing, non-coding RNAs, and
rare transcripts. Table 3.1 lists the recommended RNA-Seq read depths for various experimental
goals. Estimates are based on current best practices in RNA-Seq analysis (Conesa etal., 2016).
Paired-end sequencing (e.g., 2 × 100 bp) is commonly used, as it improves transcript assembly
and the detection of splice junctions. Single-end sequencing (e.g., 1 × 50 bp) is more cost-effective
but provides less information on transcript isoforms. The raw sequencing data is stored in FASTQ
format, containing nucleotide sequences and quality scores for each read. These data undergo preprocessing before further analysis.
3.2.4.5 Data Processing and Quality Control
Data processing and QC are foundational steps in RNA-Seq analysis that ensure the reliability and
interpretability of downstream results. After sequencing, raw reads are typically stored in FASTQ
format, containing both nucleotide sequences and associated quality scores. The rst step in data
processing involves assessing the quality of these raw reads using tools such as FastQC. This allows
researchers to evaluate metrics like per-base sequence quality, GC content, adapter contamination,
and duplication levels. Identifying issues at this stage is critical, as poor-quality data can introduce
biases or obscure true biological signals.
3.2.4.6 Read Count
In RNA-Seq, the read count refers to the number of sequencing reads that are aligned to a specic
genomic feature, typically a gene, exon, or transcript. These counts are the raw digital representation of gene expression levels and form the foundation of downstream quantitative analyses. During
the alignment or quantication stage, sequencing reads are mapped to the reference genome or transcriptome, and the number of reads corresponding to each feature is tallied. Tools such as HTSeq,
featureCounts, or pseudoaligners like Salmon and Kallisto are commonly used to generate these
count matrices efciently and accurately.
The magnitude of the read count for a given gene reects its transcriptional activity, with highly
expressed genes generally accumulating more reads. However, these counts are inuenced by

79 Bioinformatics Design for Autoimmune Research
109 × C
RPK
NL
i
×
C
i
L
i
CL
TP
jj=1 j
multiple technical and biological factors, including gene length, GC content, library preparation
biases, and sequencing depth. Therefore, raw read counts require normalization to enable meaningful comparisons between genes and samples. Despite their raw nature, read counts are essential inputs for differential expression analysis, statistical modeling, and various types of functional
enrichment analysis.
3.2.4.7 Count Normalization
Normalization is an essential step in RNA-Seq data analysis, ensuring that gene expression levels
can be meaningfully compared across samples by correcting for technical variability. The raw read
counts generated by RNA-Seq are inuenced by factors such as sequencing depth, gene length, and
compositional differences in RNA samples, all of which can introduce biases into the data (Love
etal., 2014). Without appropriate normalization, these technical effects can obscure true biologi-
cal differences, leading to inaccurate conclusions about differential gene expression. For instance,
a highly expressed gene in one sample may appear upregulated simply due to deeper sequencing,
rather than an actual increase in transcriptional activity. Normalization methods adjust for these
discrepancies, thereby enhancing the accuracy of downstream analyses and enabling valid biological interpretations (Anders & Huber, 2010).
Several normalization methods are used depending on the analysis goal. Reads Per Kilobase of
transcript per Million mapped reads (RPKM) and Fragments Per Kilobase of transcript per Million
mapped reads (FPKM) normalize for both gene length and sequencing depth. These are especially
useful for comparing expression levels of different genes within a sample. The formulas are the
following:
i
where
is the number of reads mapped to gene i, N is the total number of mapped reads in the
sample, and
Mi = (3.5)
is the length of gene i in base pairs.
FPKM is used in paired-end sequencing and is essentially the same as RPKM but accounts for
fragments instead of individual reads.
For between-sample comparisons, transcripts per million (TPM) is often preferred because it
ensures that the sum of normalized values across a sample is constant, facilitating direct comparison of transcript proportions:
M =
i
°
i
N
CL
i
6
× 10 (3.6)
In differential expression analysis, raw counts are typically normalized using scaling methods
such as Trimmed Mean of M-values (TMM), Upper Quartile (UQ), or DESeq2’s median-of-ratios
method, which are better suited for modeling count-based data in statistical frameworks like negative binomial distributions.
3.2.4.8 Differential Gene Expression
Differential gene expression analysis involves identifying and quantifying changes in transcript
levels between distinct biological conditions, such as healthy versus diseased tissues, treated versus
untreated samples, or across different developmental stages. This approach enables researchers to
determine which genes are upregulated or downregulated in response to specic stimuli or pathological states, thereby uncovering key molecular mechanisms and biological pathways underlying these conditions (Love etal., 2014). By comparing expression proles across groups, scientists
can infer regulatory changes and functional gene activity, providing valuable insights into cellular
responses and disease pathogenesis.

80 Bioinformatics of Autoimmune Diseases
y
gi
=
(
µ
˙
)
gi
gi
gi igi
µ=
s +ˆ +ˆ x
(
)
()
0 g1 i
x
i
g1
° 0
RNA-Seq is a widely used technology for measuring gene expression, offering high resolution
and sensitivity for quantifying transcript abundance (Wang etal., 2009). The raw data, represented
as read counts assigned to genes or transcripts, must rst be normalized to correct for biases due to
sequencing depth, gene length, and library size. Once normalized, statistical models—often based
on negative binomial distributions—are applied to determine whether observed expression differences between conditions are statistically signicant, while accounting for both technical variation
and biological replicates (Anders & Huber, 2010; Love etal., 2014). This rigorous analytical frame-
work ensures robust identication of DEGs that are biologically relevant.
Among the most widely used statistical frameworks are the negative binomial generalized lin-
ear models (GLMs) implemented in tools such as DESeq2 and edgeR. These models assume that
RNA-Seq count data
, representing the observed counts for gene g in sample i, follow a negative
binomial distribution:
y
NB
gi
,
gi
g
(3.7 )
where is the expected mean count for gene g in sample i, and α is the gene-specic dispersion
g
parameter that models biological variability.
The mean is typically modeled as:
(3.8)
where i is a size factor representing the library size of sample i, and qgi is the normalized expression
level of gene g in that sample.
The logarithm of the expected counts is then modeled using a linear predictor:
g
log
gi
i g
(3.9)
where g0 is the intercept term for gene g, βg1 captures the effect size (i.e., log fold change) between
conditions, and
is the covariate indicating the group (e.g., 0 for control, 1 for treatment). Hypothesis
testing is typically conducted using a Wald test or likelihood ratio test (LRT) to determine if
indicating signicant differential expression.
Alternatively, limma-voom is another popular approach that transforms count data into log2counts per million (log-CPM) and estimates the mean-variance relationship using precision weights.
It then applies linear modeling:
(3.10)
where Y is the matrix of log-transformed expression values, X is the design matrix encoding experi-
mental conditions, β are the coefcients (including fold changes), and ε is the error term. Limma
uses empirical Bayes methods to shrink the variance estimates, improving statistical power when
sample sizes are small.
The outcome of differential expression analysis is a list of genes ranked by statistical signicance
(adjusted p-values or FDR) and magnitude of change (e.g., log2 fold change). These results can be
further explored through pathway enrichment analysis, Gene Ontology annotation, and networkbased methods to interpret the biological context of the transcriptional changes.
3.2.4.9 Functional Enrichment
Functional enrichment analysis is a critical step following differential gene expression analysis in
RNA-Seq workows. It provides biological context by identifying which cellular processes, molecular pathways, or gene ontologies are signicantly overrepresented in lists of upregulated or downregulated genes (Huang etal., 2009). Once DEGs are identied, they are systematically compared
against curated biological databases such as Gene Ontology (GO), the KEGG, and Reactome. This

81 Bioinformatics Design for Autoimmune Research
comparison enables researchers to detect enriched functional themes that may underlie disease
mechanisms, treatment responses, or experimental perturbations.
Statistical tests such as Fisher’s exact test, the hypergeometric test, and gene set enrichment analysis (GSEA) are commonly employed to evaluate whether the representation of a particular functional category exceeds what would be expected by random chance. For instance, an enrichment
of immune response-related GO terms among upregulated genes may suggest immune pathway
activation in autoimmune disease models. Functional enrichment not only enhances the interpretability of transcriptomic datasets but also helps generate testable hypotheses regarding biological
mechanisms and regulatory networks.
Results of enrichment analyses are often visualized using bar plots, bubble charts, and network
diagrams to aid in identifying key biological processes. These visualizations facilitate communication of results and inform subsequent experimental design or therapeutic targeting. Ultimately,
functional enrichment analysis bridges the gap between gene-level changes and systems-level
understanding, adding mechanistic insight to RNA-Seq data interpretation.
3.2.5 SINGLE-CELL RNA SEQUENCING
Single-cell RNA-Seq (scRNA-Seq) is a cutting-edge NGS technology that enables transcriptomic proling at the resolution of individual cells. Unlike bulk RNA-Seq, which aggregates gene expression
data across mixed cell populations, scRNA-Seq provides granular insights into the transcriptional
activity of individual cells, revealing cellular heterogeneity and rare subpopulations that are often
obscured in bulk analyses. This resolution has transformed immunological research, particularly in the
context of autoimmune diseases where immune cell diversity and function are central to pathogenesis.
Autoimmune disorders are characterized by aberrant immune activation, leading to chronic
inammation and tissue damage. Since immune cells are the primary effectors in these conditions,
characterizing their gene expression patterns at the single-cell level is essential to understanding
disease mechanisms. scRNA-Seq enables the identication of distinct immune cell subsets, lineage
trajectories, and cell-specic transcriptional signatures associated with autoimmunity (Zheng etal.,
2017). It also allows researchers to monitor immune cell plasticity, uncover regulatory networks,
and examine cell–cell communication dynamics within affected tissues. These insights contribute
to the discovery of novel therapeutic targets and facilitate the development of precision medicine
approaches for autoimmune diseases.
The scRNA-Seq workow includes a series of technically demanding steps, such as sample
preparation, cell isolation, droplet-based or microwell barcoding, reverse transcription, cDNA
amplication, library preparation, sequencing, and high-dimensional data analysis (Haque etal.,
2017). Each of these stages must be rigorously optimized to ensure cell viability, minimize technical
artifacts, and preserve the transcriptional delity of each cell.
3.2.5.1 Study Design and Sample Collection
The rst critical step in a scRNA-Seq study is the thoughtful design of the experiment, including the selection of appropriate patient cohorts. Samples for scRNA-Seq are commonly obtained
from PBMCs, affected tissues such as synovial uid in RA or intestinal biopsies in IBD, or from
puried immune cell subsets relevant to the autoimmune condition being investigated. To ensure
that observed transcriptomic differences are disease-specic, it is essential to include well-matched
healthy control samples. Matching based on variables such as age, sex, and ancestry helps reduce
confounding effects and increases the interpretability of results.
Because gene expression at the single-cell level is highly sensitive to external stimuli, it is important to account for biological variability introduced by disease stage, therapeutic intervention, and
environmental exposures (Haque etal., 2017). Furthermore, to preserve cellular viability and main-
tain transcriptomic delity, sample collection and processing must be performed promptly. Ideally,
fresh samples are used, but when immediate processing is not feasible, cryopreservation under

82 Bioinformatics of Autoimmune Diseases
optimized protocols can be employed without compromising RNA integrity. Degraded RNA can
lead to dropout events and biased expression estimates, which signicantly affect downstream analyses and interpretation. Therefore, rigorous handling protocols and QC checks are vital components
of successful scRNA-Seq studies in autoimmune disease research.
3.2.5.2 Single-Cell Isolation and Preparation
Following sample collection, the isolation of individual cells is a critical step in scRNA-Seq that
directly affects data quality and biological interpretation. Successful isolation must preserve both
cell viability and transcriptomic integrity, which are essential for accurate gene expression proling
at the single-cell level. The choice of isolation method depends on the sample type, the cell populations of interest, and the specic goals of the study (Haque etal., 2017).
Fluorescence-activated cell sorting (FACS) is one of the most commonly employed techniques
for isolating specic immune cell subsets. It uses uorescently labeled antibodies to identify and
sort cells based on the expression of surface markers, allowing for high-purity separation of popula-
+
tions such as CD4
T helper cells, CD8+ cytotoxic T cells, CD25+FoxP3+ regulatory T cells (Tregs),
and various monocyte subsets. This approach offers precise selection and enables the study of rare
or functionally distinct immune cells in autoimmune diseases.
Magnetic-activated cell sorting (MACS) offers a simpler, scalable, and cost-effective alternative. In MACS, magnetic beads conjugated with antibodies are used to enrich or deplete specic
cell types by passing them through a magnetic column. Although less precise than FACS, MACS
is effective for bulk enrichment and is particularly useful in studies with limited sample volume or
when ultra-high cell purity is not required.
Each method has its strengths and limitations, and the decision between FACS and MACS often
reects a balance between specicity, throughput, cost, and the sensitivity of downstream transcriptomic applications.
Figure 3.2(a) illustrates the principle of FACS, a specialized form of ow cytometry used to
separate and collect specic cell populations based on their uorescent properties. Cells are rst
FIGURE 3.2 (a) Fluorescence-activated cell sorting and (b) magnetic-activated cell sorting.

83 Bioinformatics Design for Autoimmune Research
suspended in a uid and directed through a ow cell in single le. As they pass through a laser
beam, uorescently labeled cells emit light that is detected by a sensitive photodetector. Based on
the intensity and pattern of uorescence, each cell is assigned an electrical charge by a charge plate.
These charged droplets are then deected by an electric eld generated by deection plates, guiding
them into different collection tubes according to their uorescence signal. Unlabeled cells, which
lack uorescence, follow a default trajectory and are collected separately. This technique enables
high-precision isolation of distinct cell subsets for downstream applications such as transcriptomic
analysis, proteomic proling, or cell culture.
Figure 3.2(b) illustrates the principle of MACS, a technique used to isolate specic cell popula-
tions based on the presence of surface markers recognized by magnetically labeled antibodies. A
heterogeneous cell suspension is rst incubated with magnetic microbeads conjugated to antibodies
that selectively bind the target cells. This mixture is then passed through a column placed within
a strong magnetic eld. Labeled cells are retained in the column due to their magnetic properties,
while unlabeled cells ow through and are collected separately. Once the magnetic eld is removed,
the retained cells can be eluted and recovered. This approach provides a simple, efcient, and gentle
method for enriching or depleting specic cell types, supporting a wide range of downstream applications in immunology, molecular biology, and regenerative medicine.
For solid tissues, single-cell suspensions are generated by enzymatic digestion using collagenase
or trypsin. Tissue dissociation protocols must be carefully optimized to prevent cell damage and
transcriptional artifacts. In cases where tissue-derived cells cannot be processed immediately, cryopreservation is used to store cells without signicant loss of viability.
Once isolated, cell viability is assessed using trypan blue staining or automated cell counters.
High viability (>85%) is essential for obtaining high-quality single-cell transcriptomic data.
3.2.5.3 Single-Cell Barcoding and Library Preparation
After cell isolation, the next essential step in scRNA-Seq is to uniquely label the transcripts of each
individual cell to enable accurate downstream analysis. This is accomplished using single-cell barcoding techniques that attach UMIs and cell-specic barcodes to each transcript. UMIs allow for
the correction of amplication biases and enable precise quantication of gene expression, while
cell barcodes ensure that transcripts can be traced back to their cell of origin (Zheng etal., 2017).
Several scRNA-Seq platforms have been developed to implement this strategy, each with distinct
trade-offs in terms of sensitivity, cost, and throughput. The 10× Genomics Chromium system is the
most widely adopted droplet-based platform. It uses microuidic devices to encapsulate single cells
into nanoliter-scale droplets, each containing a gel bead coated with barcoded oligonucleotides.
These oligos include a cell barcode, UMI, and poly(dT) sequence to capture mRNA, enabling massively parallel sequencing of thousands of cells (Zheng etal., 2017).
Smart-Seq2, an alternative plate-based method, provides full-length transcript coverage and
is highly sensitive, making it ideal for detecting splice variants and low-abundance transcripts.
However, it has lower throughput and higher per-cell cost compared to droplet-based systems
(Picelli etal., 2014). Seq-Well, a microwell-based platform, offers a cost-effective and scalable solu-
tion for high-throughput single-cell proling in settings with limited resources.
The choice of barcoding and library preparation method depends on the specic objectives of the
study, the required resolution, and budgetary constraints.
Figure 3.3 illustrates the single-cell barcoding process in a stepwise manner, beginning with
the preparation of a cell suspension. In this step, individual cells from a heterogeneous sample are
isolated and suspended in a solution, ensuring that they remain viable and evenly distributed for
encapsulation. The next stage involves the formation of microdroplets, where each droplet ideally
contains a single cell along with a uniquely barcoded bead. This bead is coated with oligonucle-
otides containing a unique cell barcode, UMIs, and oligo(dT) sequences for mRNA capture. As the
cell is lysed within the droplet, its released mRNA molecules hybridize to the barcoded oligos on
the bead, allowing for the association of each transcript with both its cell of origin and its individual
Соседние файлы в папке Библиотека им академика М.И. Перельмана
