Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
26 Мб
Скачать
74 Bioinformatics of Autoimmune Diseases
100,000, 000 × 150
3000,000, 000
For example, if we sequence 100 million reads of 150 bp each for a genome of 3 billion bp (human genome):
Depth = 5coverage (3.3)
The sequencing depth is usually referred to as “×” coverage as the “×” symbol means “times”; so 5× coverage means each base has been sequenced ve times on average.
The required depth depends on the study’s goals:
• 30× depth (standard coverage) is sufcient for detecting germline variants.
• 60–100× depth is used for detecting somatic mutations in diseases like cancer.
• 10× depth (low-coverage WGS) is sometimes used in large-scale population studies for cost-effective variant discovery.
Short-read sequencing (e.g., Illumina) is widely used due to its high accuracy and cost efciency, while long-read sequencing (e.g., PacBio, Oxford Nanopore) is benecial for detecting structural variants, repetitive elements, and phased haplotypes.
Raw sequencing data is generated in FASTQ format, containing nucleotide sequences and qual­ity scores. These data undergo preprocessing steps such as adapter trimming, quality ltering, and base correction to remove low-condence reads.
3.2.3.4 Read Alignment and Variant Calling
After sequencing, the raw reads are aligned to a reference genome (e.g., GRCh38 or GRCh37/hg19) using alignment tools such as Burrows–Wheeler Aligner (BWA) or Bowtie2. Accurate alignment ensures proper variant detection, and alignment quality is assessed using mapping rate, duplication rate, and coverage uniformity metrics.
Once aligned, variant calling is performed to identify genetic differences between the study participants and the reference genome. Variant callers such as
• GATK (Genome Analysis Toolkit): the gold standard for germline variant calling.
• DeepVariant: a deep learning-based variant caller for improved accuracy.
• Lumpy or Manta: for structural variant detection.
are used to detect different classes of variants, including SNVs, indels, CNVs, and SVs.
Post-calling, variants are ltered based on read depth, genotype quality, and allele frequency to eliminate false positives. Variants are then annotated using reference databases such as d b SN P, 1000 Genomes, and gnomAD to determine their population frequencies and potential pathogenicity.
3.2.3.5 Functional Annotation and Pathway Analysis
Once genetic variants have been identied through whole-genome or exome sequencing, the next essential step is to assess their potential biological relevance. Functional annotation tools such as ANNOVAR, VEP, and SnpEff are widely used to predict the consequences of variants on gene func­tion, protein structure, and regulatory elements (McLaren etal., 2016). Exonic variants are exam- ined for their effects on protein-coding sequences, including missense, nonsense, and frameshift mutations, which may alter amino acid composition or lead to truncated, nonfunctional proteins. Conversely, variants in non-coding regions are evaluated for their inuence on gene regulation, messenger RNA (mRNA) splicing, and enhancer activity.
To prioritize variants likely to contribute to disease, researchers utilize pathogenicity predic­tion algorithms such as PolyPhen-2, SIFT, and CADD. These tools integrate evolutionary conser­vation, biochemical properties, and multiple annotation features to estimate the likelihood that a
75 Bioinformatics Design for Autoimmune Research
given variant is deleterious (Adzhubei etal., 2010). Moreover, variants located in regulatory regions are interpreted using epigenomic resources such as the ENCODE project and the GTEx database, which provide insights into tissue-specic gene expression and regulatory element activity, particu­larly in immune-relevant tissues.
Beyond single-variant interpretation, pathway and network analysis methods are applied to place genetic ndings into a systems biology context. Tools such as Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and IPA are used to determine whether disease-associated variants are enriched in known biological pathways. This is particularly important in autoimmune diseases, which frequently involve dysregulation in pathways related to cytokine signaling, T-cell activation, antigen presentation, and NF-κB signaling. By mapping variants to these immune-related path­ways, researchers can uncover underlying disease mechanisms and identify potential targets for therapeutic intervention.
3.2.4 RNA SEQUENCING
RNA-Seq is a high-throughput NGS technology that enables comprehensive proling of gene expression across the entire transcriptome. Unlike traditional microarray-based approaches, RNA­Seq offers greater sensitivity, a broader dynamic range, and the ability to detect novel transcripts without requiring prior sequence information. This technology allows for the quantication of gene expression, detection of alternative splicing events, and characterization of various RNA species, including long non-coding RNAs (lncRNAs), miRNAs, and circular RNAs (circRNAs), providing a multifaceted view of transcriptomic regulation.
In the context of autoimmune diseases, RNA-Seq has become an indispensable tool for uncov­ering the molecular mechanisms underlying immune dysregulation. Autoimmune conditions such as SLE, RA, and MS are driven by aberrant immune responses that result in chronic inammation and tissue destruction. By comparing gene expression proles between affected individuals and healthy controls, researchers can identify differentially expressed genes (DEGs) and dysregulated pathways central to disease pathogenesis. RNA-Seq also facilitates the identication of immune cell-specic signatures, cytokine expression patterns, and transcriptional networks associated with disease activity and progression.
Moreover, RNA-Seq provides valuable insights into how environmental exposures (such as infec­tions, stress, or dietary factors) interact with genetic predispositions to inuence gene expression in autoimmune settings. This integrative approach helps to identify biomarkers for early diagnosis, stratify patients by molecular subtypes, and inform precision medicine strategies by predicting treatment responses.
The RNA-Seq workow includes critical steps such as sample collection, RNA extraction, qual­ity assessment, library construction, sequencing, alignment to reference genomes, and downstream analysis. Differential expression analysis is typically performed using statistical tools such as DESeq2 or edgeR, followed by functional enrichment analysis to interpret the biological signi­cance of observed transcriptional changes (Love etal., 2014). Each step requires stringent QC and bioinformatics expertise to ensure the generation of accurate and reproducible data.
3.2.4.1 Study Design and Sample Collection
A well-designed RNA-Seq study begins with the careful selection of an appropriate cohort, typi­cally composed of individuals diagnosed with autoimmune diseases and demographically matched healthy controls. Because gene expression is affected by numerous variables (such as age, sex, medi­cation history, disease severity, and environmental exposures), these confounding factors must be rigorously controlled during study design to ensure that observed transcriptional differences reect disease biology rather than technical or demographic artifacts (Conesa etal., 2016). Matching par- ticipants and incorporating metadata in the statistical analysis helps mitigate these confounding effects and improves interpretability.
76 Bioinformatics of Autoimmune Diseases
Standardizing sample collection procedures is essential to minimize batch effects, which can obscure biological signals and reduce reproducibility. Factors such as collection time, storage conditions, and RNA extraction protocols must be kept consistent across all samples. Even small deviations in processing can introduce technical variability that signicantly impacts transcript quantication.
RNA-Seq can be applied to a variety of biological sources depending on the research objective. Peripheral blood mononuclear cells (PBMCs), sorted immune cell subsets (e.g., T cells, B cells, monocytes), and whole blood are commonly used in systemic autoimmune diseases. Site-specic samples such as synovial uid (for RA), intestinal biopsies (for inammatory bowel disease (IBD)), or cerebrospinal uid (for MS) provide localized insights into tissue-specic immune responses. The selection of tissue type should align with the pathophysiological features of the disease being studied.
To maintain RNA integrity, immediate stabilization is required after sample collection. Reagents such as TRIzol or RNAlater are commonly used to prevent RNA degradation by RNases. If imme­diate processing is not feasible, samples are rapidly frozen and stored at −80°C to preserve RNA quality. Given RNA’s susceptibility to enzymatic degradation, meticulous handling and storage pro­cedures are paramount to ensure the generation of reliable sequencing data.
3.2.4.2 RNA Extraction and Quality Control
Total RNA is extracted from the collected samples using commercial kits (e.g., Qiagen RNeasy, TRIzol reagent) or column-based purication methods. The extracted RNA is then assessed for integrity and purity using spectrophotometry (e.g., NanoDrop), uorometry (e.g., Qubit), and elec­trophoresis (e.g., Agilent Bioanalyzer).
The RNA integrity number (RIN) is used to evaluate RNA quality, with a RIN 7 considered suitable for RNA-Seq. High-quality RNA is essential for generating reliable sequencing data, as degraded RNA can lead to biased or incomplete transcriptome proles.
Depending on the research objectives, different RNA fractions may be analyzed:
Total RNA sequencing captures all RNA species, including coding and non-coding RNAs.
Poly(A)+ RNA sequencing enriches for mRNA, excluding ribosomal RNA (rRNA) and most non-coding RNAs.
rRNA-depleted sequencing is used for analyzing non-coding RNAs, such as long non- coding RNAs (lncRNAs) and circRNAs.
Small RNA sequencing focuses on small RNAs like microRNAs (miRNAs), which regu­late immune responses.
3.2.4.3 Library Preparation and RNA Fragmentation
RNA-Seq library preparation consists of a series of precise molecular steps aimed at converting RNA molecules into complementary DNA (cDNA) and attaching sequencing adapters to enable high-throughput sequencing. The specic methodology depends on the type of RNA being targeted (most commonly mRNA or total RNA).
For mRNA-focused studies, the rst and most critical step is poly(A) selection, which enriches for mature mRNA transcripts by isolating RNA molecules that possess polyadenylated [poly(A)] tails. These tails are a hallmark of most eukaryotic mRNAs and can be selectively captured using oligonucleotides composed of deoxythymidine [oligo(dT)] bound to magnetic beads. After incu­bation of total RNA with oligo(dT)-conjugated beads, a series of wash and elution steps removes rRNA and the majority of non-coding RNAs, resulting in an mRNA-enriched fraction ready for downstream processing.
In contrast, total RNA-Seq enables a more comprehensive view of the transcriptome, including both coding and non-coding RNA species such as lncRNAs and circRNAs. Because rRNA con­stitutes over 80% of total RNA, its removal is essential for capturing low-abundance transcripts.
77 Bioinformatics Design for Autoimmune Research
Total numberofreadReadlength
×
Sequencin
)
FIGURE 3.1 Workow of mRNA enrichment via poly(A) selection.
Depletion strategies such as Ribo-Zero or NEBNext rRNA depletion kits are commonly employed for this purpose.
Following RNA enrichment, either by poly(A) selection or rRNA depletion, the RNA is frag­mented into smaller segments, typically ranging from 200 to 300 bp, using enzymatic or heat-based methods (see Figure 3.1). This fragmentation is necessary to ensure uniform read coverage across the transcript and to facilitate accurate transcript reconstruction during bioinformatics analysis (Conesa etal., 2016).
The fragmented RNA is then reverse transcribed into cDNA using random primers and reverse transcriptase enzymes. Sequencing adapters are ligated to both ends of the cDNA fragments to enable their amplication and subsequent detection by NGS platforms. In many modern protocols, UMIs are incorporated during adapter ligation or PCR amplication. UMIs allow researchers to track individual RNA molecules, helping correct for PCR duplicates and improve transcript quan­tication accuracy.
3.2.4.4 High-Throughput Sequencing
The prepared RNA-Seq libraries are sequenced-using high-throughput NGS platforms such as Illumina NovaSeq, HiSeq, or NextSeq. The selection of sequencing platform depends on factors such as sequencing depth, read depth, read length, and study objectives.
In RNA-Seq, sequencing depth (coverage) is determined by the size of the transcriptome, rather than the genome size as in WGS.
gDepth X
The human transcriptome size refers to the total number of nucleotides in all RNA molecules expressed in a cell or tissue at a given time. Unlike the genome, which is xed at about 3.2 billion bp, the transcriptome is dynamic and varies by cell type, condition, and study purpose. In studies focusing only on protein-coding genes, the transcriptome size is estimated at around 50 million bp, based on roughly 20,000 genes with an average transcript length of 2500 bp. When non-coding RNAs such as lncRNAs and microRNAs (miRNAs) are included, the estimated size increases to 60–80 million bp, depending on the annotation database used (e.g., GENCODE, RefSeq). However, in typical RNA-Seq experiments, only 10,000–15,000 genes are expressed in a sample, resulting in an effective transcriptome size of about 25–40 million bp. This estimate is essential for planning appropriate sequencing depth.
In RNA-Seq experiments, it is more appropriate to report sequencing output as reads per sam­ple rather than average sequencing depth (X coverage). Unlike whole-genome sequencing, where coverage is uniformly distributed across a xed genome size, RNA-Seq targets the transcriptome (adynamic and variable subset of the genome that differs between cell types, conditions, and even
=
()
Totallengthofall transcriptsTranscriptome( size
(3.4)
78 Bioinformatics of Autoimmune Diseases
TABLE 3.1 Typical RNA-Seq Read Depth Recommendations
Experimental Goal Recommended Read Depth Comments
Gene expression proling (standard) 20–30 million reads/sample Detecting moderately to highly expressed genes Differential gene expression (DGE) 30–50 million reads/sample Higher depth improves statistical power Low-abundance transcript detection Transcript isoform/splicing analysis
Single-cell RNA-Seq 20,000–100,000 reads per cell Depends on cell type and transcript complexity De novo transcriptome assembly
50 million reads/sample 75–100 million reads/sample
100 million reads per sample
Increases sensitivity to weakly expressed genes Required for resolving alternative splicing or
isoforms
Necessary due to lack of a reference and high
diversity
individual samples). Because transcript abundance varies widely, sequencing reads are not evenly distributed across all genes, making average coverage a less meaningful measure. Reporting the number of reads per sample (read depth) provides a more direct and practical metric for assessing data quantity, evaluating experiment consistency, and planning the required depth for capturing transcript diversity and detecting differential expression. This approach aligns more closely with how RNA-Seq data is generated and analyzed in practice.
Read depth is a critical factor in RNA-Seq experiments. Standard RNA-Seq requires approxi­mately 10–50 million reads per sample for gene-level expression analysis, while deeper sequencing (>100 million reads per sample) is needed for detecting alternative splicing, non-coding RNAs, and rare transcripts. Table 3.1 lists the recommended RNA-Seq read depths for various experimental goals. Estimates are based on current best practices in RNA-Seq analysis (Conesa etal., 2016).
Paired-end sequencing (e.g., 2 × 100 bp) is commonly used, as it improves transcript assembly and the detection of splice junctions. Single-end sequencing (e.g., 1 × 50 bp) is more cost-effective but provides less information on transcript isoforms. The raw sequencing data is stored in FASTQ format, containing nucleotide sequences and quality scores for each read. These data undergo pre­processing before further analysis.
3.2.4.5 Data Processing and Quality Control
Data processing and QC are foundational steps in RNA-Seq analysis that ensure the reliability and interpretability of downstream results. After sequencing, raw reads are typically stored in FASTQ format, containing both nucleotide sequences and associated quality scores. The rst step in data processing involves assessing the quality of these raw reads using tools such as FastQC. This allows researchers to evaluate metrics like per-base sequence quality, GC content, adapter contamination, and duplication levels. Identifying issues at this stage is critical, as poor-quality data can introduce biases or obscure true biological signals.
3.2.4.6 Read Count
In RNA-Seq, the read count refers to the number of sequencing reads that are aligned to a specic genomic feature, typically a gene, exon, or transcript. These counts are the raw digital representa­tion of gene expression levels and form the foundation of downstream quantitative analyses. During the alignment or quantication stage, sequencing reads are mapped to the reference genome or tran­scriptome, and the number of reads corresponding to each feature is tallied. Tools such as HTSeq, featureCounts, or pseudoaligners like Salmon and Kallisto are commonly used to generate these count matrices efciently and accurately.
The magnitude of the read count for a given gene reects its transcriptional activity, with highly expressed genes generally accumulating more reads. However, these counts are inuenced by
79 Bioinformatics Design for Autoimmune Research
109 × C
RPK
NL
i
×
C
i
L
i
CL
TP
jj=1 j
multiple technical and biological factors, including gene length, GC content, library preparation biases, and sequencing depth. Therefore, raw read counts require normalization to enable mean­ingful comparisons between genes and samples. Despite their raw nature, read counts are essen­tial inputs for differential expression analysis, statistical modeling, and various types of functional enrichment analysis.
3.2.4.7 Count Normalization
Normalization is an essential step in RNA-Seq data analysis, ensuring that gene expression levels can be meaningfully compared across samples by correcting for technical variability. The raw read counts generated by RNA-Seq are inuenced by factors such as sequencing depth, gene length, and compositional differences in RNA samples, all of which can introduce biases into the data (Love
etal., 2014). Without appropriate normalization, these technical effects can obscure true biologi-
cal differences, leading to inaccurate conclusions about differential gene expression. For instance, a highly expressed gene in one sample may appear upregulated simply due to deeper sequencing, rather than an actual increase in transcriptional activity. Normalization methods adjust for these discrepancies, thereby enhancing the accuracy of downstream analyses and enabling valid biologi­cal interpretations (Anders & Huber, 2010).
Several normalization methods are used depending on the analysis goal. Reads Per Kilobase of transcript per Million mapped reads (RPKM) and Fragments Per Kilobase of transcript per Million mapped reads (FPKM) normalize for both gene length and sequencing depth. These are especially useful for comparing expression levels of different genes within a sample. The formulas are the following:
i
where
is the number of reads mapped to gene i, N is the total number of mapped reads in the
sample, and
Mi = (3.5)
is the length of gene i in base pairs.
FPKM is used in paired-end sequencing and is essentially the same as RPKM but accounts for fragments instead of individual reads.
For between-sample comparisons, transcripts per million (TPM) is often preferred because it ensures that the sum of normalized values across a sample is constant, facilitating direct compari­son of transcript proportions:
M =
i
°
i
N
CL
i
6
× 10 (3.6)
In differential expression analysis, raw counts are typically normalized using scaling methods such as Trimmed Mean of M-values (TMM), Upper Quartile (UQ), or DESeq2’s median-of-ratios method, which are better suited for modeling count-based data in statistical frameworks like nega­tive binomial distributions.
3.2.4.8 Differential Gene Expression
Differential gene expression analysis involves identifying and quantifying changes in transcript levels between distinct biological conditions, such as healthy versus diseased tissues, treated versus untreated samples, or across different developmental stages. This approach enables researchers to determine which genes are upregulated or downregulated in response to specic stimuli or patho­logical states, thereby uncovering key molecular mechanisms and biological pathways underly­ing these conditions (Love etal., 2014). By comparing expression proles across groups, scientists can infer regulatory changes and functional gene activity, providing valuable insights into cellular responses and disease pathogenesis.
80 Bioinformatics of Autoimmune Diseases
y
gi
=
(
µ
˙
)
gi
gi
gi igi
µ=
s x
(
)
()
0 g1 i
x
i
g1
° 0
RNA-Seq is a widely used technology for measuring gene expression, offering high resolution and sensitivity for quantifying transcript abundance (Wang etal., 2009). The raw data, represented as read counts assigned to genes or transcripts, must rst be normalized to correct for biases due to sequencing depth, gene length, and library size. Once normalized, statistical models—often based on negative binomial distributions—are applied to determine whether observed expression differ­ences between conditions are statistically signicant, while accounting for both technical variation and biological replicates (Anders & Huber, 2010; Love etal., 2014). This rigorous analytical frame- work ensures robust identication of DEGs that are biologically relevant.
Among the most widely used statistical frameworks are the negative binomial generalized lin- ear models (GLMs) implemented in tools such as DESeq2 and edgeR. These models assume that RNA-Seq count data
, representing the observed counts for gene g in sample i, follow a negative
binomial distribution:
y
NB
gi
,
gi
g
(3.7 )
where is the expected mean count for gene g in sample i, and α is the gene-specic dispersion
g
parameter that models biological variability.
The mean is typically modeled as:
(3.8)
where i is a size factor representing the library size of sample i, and qgi is the normalized expression level of gene g in that sample.
The logarithm of the expected counts is then modeled using a linear predictor:
g
log
gi
i g
(3.9)
where g0 is the intercept term for gene g, βg1 captures the effect size (i.e., log fold change) between conditions, and
is the covariate indicating the group (e.g., 0 for control, 1 for treatment). Hypothesis testing is typically conducted using a Wald test or likelihood ratio test (LRT) to determine if indicating signicant differential expression.
Alternatively, limma-voom is another popular approach that transforms count data into log2­counts per million (log-CPM) and estimates the mean-variance relationship using precision weights. It then applies linear modeling:
(3.10)
where Y is the matrix of log-transformed expression values, X is the design matrix encoding experi- mental conditions, β are the coefcients (including fold changes), and ε is the error term. Limma uses empirical Bayes methods to shrink the variance estimates, improving statistical power when sample sizes are small.
The outcome of differential expression analysis is a list of genes ranked by statistical signicance (adjusted p-values or FDR) and magnitude of change (e.g., log2 fold change). These results can be further explored through pathway enrichment analysis, Gene Ontology annotation, and network­based methods to interpret the biological context of the transcriptional changes.
3.2.4.9 Functional Enrichment
Functional enrichment analysis is a critical step following differential gene expression analysis in RNA-Seq workows. It provides biological context by identifying which cellular processes, molec­ular pathways, or gene ontologies are signicantly overrepresented in lists of upregulated or down­regulated genes (Huang etal., 2009). Once DEGs are identied, they are systematically compared against curated biological databases such as Gene Ontology (GO), the KEGG, and Reactome. This
81 Bioinformatics Design for Autoimmune Research
comparison enables researchers to detect enriched functional themes that may underlie disease mechanisms, treatment responses, or experimental perturbations.
Statistical tests such as Fisher’s exact test, the hypergeometric test, and gene set enrichment anal­ysis (GSEA) are commonly employed to evaluate whether the representation of a particular func­tional category exceeds what would be expected by random chance. For instance, an enrichment of immune response-related GO terms among upregulated genes may suggest immune pathway activation in autoimmune disease models. Functional enrichment not only enhances the interpret­ability of transcriptomic datasets but also helps generate testable hypotheses regarding biological mechanisms and regulatory networks.
Results of enrichment analyses are often visualized using bar plots, bubble charts, and network diagrams to aid in identifying key biological processes. These visualizations facilitate communi­cation of results and inform subsequent experimental design or therapeutic targeting. Ultimately, functional enrichment analysis bridges the gap between gene-level changes and systems-level understanding, adding mechanistic insight to RNA-Seq data interpretation.
3.2.5 SINGLE-CELL RNA SEQUENCING
Single-cell RNA-Seq (scRNA-Seq) is a cutting-edge NGS technology that enables transcriptomic pro­ling at the resolution of individual cells. Unlike bulk RNA-Seq, which aggregates gene expression data across mixed cell populations, scRNA-Seq provides granular insights into the transcriptional activity of individual cells, revealing cellular heterogeneity and rare subpopulations that are often obscured in bulk analyses. This resolution has transformed immunological research, particularly in the context of autoimmune diseases where immune cell diversity and function are central to pathogenesis.
Autoimmune disorders are characterized by aberrant immune activation, leading to chronic inammation and tissue damage. Since immune cells are the primary effectors in these conditions, characterizing their gene expression patterns at the single-cell level is essential to understanding disease mechanisms. scRNA-Seq enables the identication of distinct immune cell subsets, lineage trajectories, and cell-specic transcriptional signatures associated with autoimmunity (Zheng etal.,
2017). It also allows researchers to monitor immune cell plasticity, uncover regulatory networks,
and examine cell–cell communication dynamics within affected tissues. These insights contribute to the discovery of novel therapeutic targets and facilitate the development of precision medicine approaches for autoimmune diseases.
The scRNA-Seq workow includes a series of technically demanding steps, such as sample preparation, cell isolation, droplet-based or microwell barcoding, reverse transcription, cDNA amplication, library preparation, sequencing, and high-dimensional data analysis (Haque etal.,
2017). Each of these stages must be rigorously optimized to ensure cell viability, minimize technical
artifacts, and preserve the transcriptional delity of each cell.
3.2.5.1 Study Design and Sample Collection
The rst critical step in a scRNA-Seq study is the thoughtful design of the experiment, includ­ing the selection of appropriate patient cohorts. Samples for scRNA-Seq are commonly obtained from PBMCs, affected tissues such as synovial uid in RA or intestinal biopsies in IBD, or from puried immune cell subsets relevant to the autoimmune condition being investigated. To ensure that observed transcriptomic differences are disease-specic, it is essential to include well-matched healthy control samples. Matching based on variables such as age, sex, and ancestry helps reduce confounding effects and increases the interpretability of results.
Because gene expression at the single-cell level is highly sensitive to external stimuli, it is impor­tant to account for biological variability introduced by disease stage, therapeutic intervention, and environmental exposures (Haque etal., 2017). Furthermore, to preserve cellular viability and main- tain transcriptomic delity, sample collection and processing must be performed promptly. Ideally, fresh samples are used, but when immediate processing is not feasible, cryopreservation under
82 Bioinformatics of Autoimmune Diseases
optimized protocols can be employed without compromising RNA integrity. Degraded RNA can lead to dropout events and biased expression estimates, which signicantly affect downstream anal­yses and interpretation. Therefore, rigorous handling protocols and QC checks are vital components of successful scRNA-Seq studies in autoimmune disease research.
3.2.5.2 Single-Cell Isolation and Preparation
Following sample collection, the isolation of individual cells is a critical step in scRNA-Seq that directly affects data quality and biological interpretation. Successful isolation must preserve both cell viability and transcriptomic integrity, which are essential for accurate gene expression proling at the single-cell level. The choice of isolation method depends on the sample type, the cell popula­tions of interest, and the specic goals of the study (Haque etal., 2017).
Fluorescence-activated cell sorting (FACS) is one of the most commonly employed techniques for isolating specic immune cell subsets. It uses uorescently labeled antibodies to identify and sort cells based on the expression of surface markers, allowing for high-purity separation of popula-
+
tions such as CD4
T helper cells, CD8+ cytotoxic T cells, CD25+FoxP3+ regulatory T cells (Tregs), and various monocyte subsets. This approach offers precise selection and enables the study of rare or functionally distinct immune cells in autoimmune diseases.
Magnetic-activated cell sorting (MACS) offers a simpler, scalable, and cost-effective alterna­tive. In MACS, magnetic beads conjugated with antibodies are used to enrich or deplete specic cell types by passing them through a magnetic column. Although less precise than FACS, MACS is effective for bulk enrichment and is particularly useful in studies with limited sample volume or when ultra-high cell purity is not required.
Each method has its strengths and limitations, and the decision between FACS and MACS often reects a balance between specicity, throughput, cost, and the sensitivity of downstream transcrip­tomic applications.
Figure 3.2(a) illustrates the principle of FACS, a specialized form of ow cytometry used to
separate and collect specic cell populations based on their uorescent properties. Cells are rst
FIGURE 3.2 (a) Fluorescence-activated cell sorting and (b) magnetic-activated cell sorting.
83 Bioinformatics Design for Autoimmune Research
suspended in a uid and directed through a ow cell in single le. As they pass through a laser beam, uorescently labeled cells emit light that is detected by a sensitive photodetector. Based on the intensity and pattern of uorescence, each cell is assigned an electrical charge by a charge plate. These charged droplets are then deected by an electric eld generated by deection plates, guiding them into different collection tubes according to their uorescence signal. Unlabeled cells, which lack uorescence, follow a default trajectory and are collected separately. This technique enables high-precision isolation of distinct cell subsets for downstream applications such as transcriptomic analysis, proteomic proling, or cell culture.
Figure 3.2(b) illustrates the principle of MACS, a technique used to isolate specic cell popula-
tions based on the presence of surface markers recognized by magnetically labeled antibodies. A heterogeneous cell suspension is rst incubated with magnetic microbeads conjugated to antibodies that selectively bind the target cells. This mixture is then passed through a column placed within a strong magnetic eld. Labeled cells are retained in the column due to their magnetic properties, while unlabeled cells ow through and are collected separately. Once the magnetic eld is removed, the retained cells can be eluted and recovered. This approach provides a simple, efcient, and gentle method for enriching or depleting specic cell types, supporting a wide range of downstream appli­cations in immunology, molecular biology, and regenerative medicine.
For solid tissues, single-cell suspensions are generated by enzymatic digestion using collagenase or trypsin. Tissue dissociation protocols must be carefully optimized to prevent cell damage and transcriptional artifacts. In cases where tissue-derived cells cannot be processed immediately, cryo­preservation is used to store cells without signicant loss of viability.
Once isolated, cell viability is assessed using trypan blue staining or automated cell counters. High viability (>85%) is essential for obtaining high-quality single-cell transcriptomic data.
3.2.5.3 Single-Cell Barcoding and Library Preparation
After cell isolation, the next essential step in scRNA-Seq is to uniquely label the transcripts of each individual cell to enable accurate downstream analysis. This is accomplished using single-cell bar­coding techniques that attach UMIs and cell-specic barcodes to each transcript. UMIs allow for the correction of amplication biases and enable precise quantication of gene expression, while cell barcodes ensure that transcripts can be traced back to their cell of origin (Zheng etal., 2017).
Several scRNA-Seq platforms have been developed to implement this strategy, each with distinct trade-offs in terms of sensitivity, cost, and throughput. The 10× Genomics Chromium system is the most widely adopted droplet-based platform. It uses microuidic devices to encapsulate single cells into nanoliter-scale droplets, each containing a gel bead coated with barcoded oligonucleotides. These oligos include a cell barcode, UMI, and poly(dT) sequence to capture mRNA, enabling mas­sively parallel sequencing of thousands of cells (Zheng etal., 2017).
Smart-Seq2, an alternative plate-based method, provides full-length transcript coverage and is highly sensitive, making it ideal for detecting splice variants and low-abundance transcripts. However, it has lower throughput and higher per-cell cost compared to droplet-based systems (Picelli etal., 2014). Seq-Well, a microwell-based platform, offers a cost-effective and scalable solu- tion for high-throughput single-cell proling in settings with limited resources.
The choice of barcoding and library preparation method depends on the specic objectives of the study, the required resolution, and budgetary constraints.
Figure 3.3 illustrates the single-cell barcoding process in a stepwise manner, beginning with
the preparation of a cell suspension. In this step, individual cells from a heterogeneous sample are isolated and suspended in a solution, ensuring that they remain viable and evenly distributed for encapsulation. The next stage involves the formation of microdroplets, where each droplet ideally contains a single cell along with a uniquely barcoded bead. This bead is coated with oligonucle- otides containing a unique cell barcode, UMIs, and oligo(dT) sequences for mRNA capture. As the cell is lysed within the droplet, its released mRNA molecules hybridize to the barcoded oligos on the bead, allowing for the association of each transcript with both its cell of origin and its individual