Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
64 Bioinformatics of Autoimmune Diseases
Matzaraki, V., Kumar, V., Wijmenga, C., & Zhernakova, A. (2017). The MHC locus and genetic susceptibility
to autoimmune and infectious diseases. Genome Biology, 18, 76. https://doi.org/10.1186/s13059-017-
1207-1
Rawlings, D. J., Schwartz, M. A., Jackson, S. W., & Meyer-Bahlburg, A. (2012). Integration of B cell responses
through Toll-like receptors and antigen receptors. Nature Reviews Immunology, 12(4), 282–294. https://
doi.org /10.1038/nri3150
Raychaudhuri, S. (2010). Recent advances in the genetics of rheumatoid arthritis. Current Opinion in
Rheumatology, 22(2), 109–118. https://doi.org/10.1097/BOR.0b013e328336 474d
Remmers, E. F., Plenge, R. M., Lee, A. T., Graham, R. R., Hom, G., Behrens, T. W., … Gregersen, P. K.
(2007). STAT4 and the risk of rheumatoid arthritis and systemic lupus erythematosus. New England
Journal of Medicine, 357(10), 977–986. https://doi.org/10.1056/NEJ Moa073003
Sands, B. E. (2007). Inammatory bowel disease: Past, present, and future. Journal of Gastroenterology,
42(Suppl 1), 16–25. https://doi.org/10.1007/s00535-006-1995-7
Schubert, D., Bode, C., Kenefeck, R., Hou, T. Z., Wing, J. B., Kennedy, A., … Walker, L. S. (2014). Autosomal
dominant immune dysregulation syndrome in humans with CTLA4 mutations. Nature Medicine, 20(12),
1410 –1416. htt ps://doi.org/10.1038/nm.3746
Snow, A. L., Xiao, W., Stinson, J. R., Lu, W., Chaigne-Delalande, B., Zheng, L., … Lenardo, M. J. (2012).
Congenital B cell lymphocytosis explained by novel germline CARD11 mutations. Journal of
Experimental Medicine, 209(12), 2247–2261. https://doi.org/10.1084/jem.20120831
Suzuki, A., Yamada, R., Chang, X., Tokuhiro, S., Sawada, T., Suzuki, M., … Yamamoto, K. (2003). Functional
haplotypes of PADI4, encoding citrullinating enzyme peptidylarginine deiminase 4, are associated with
rheumatoid arthritis. Nature Genetics, 34(4), 395–402. https://doi.org/10.1038/ng1206
Tang, Y., Luo, X., Cui, H., Ni, X., Yuan, M., Guo, Y., …, Tsokos, G. C. (2009). MicroRNA-146a contributes
to abnormal activation of the type I interferon pathway in human lupus by targeting the key signaling
proteins. Arthritis & Rheumatism, 60(4), 1065–1075. https://doi.org/10.1002/art.24436
Trynka, G., Hunt, K. A., Bockett, N. A., Romanos, J., Mistry, V., Szperl, A., … van Heel, D. A. (2011). Dense
genotyping identies and localizes multiple common and rare variant association signals in celiac dis-
ease. Nature Genetics, 43(12), 1193–1201. https://doi.org/10.1038/ng.998
Ueda, H., Howson, J. M. M., Esposito, L., Heward, J., Snook, H., Chamberlain, G., … Todd, J. A. (2003).
Association of the T-cell regulatory gene CTLA4 with susceptibility to autoimmune disease. Nature,
423(6939), 506–511. https://doi.org/10.1038/nature01621
Yang, Y., Chung, E. K., Wu, Y. L., Savelli, S. L., Nagaraja, H. N., Zhou, B., … Tsao, B. P. (2007). Gene
copy-number variation and associated polymorphisms of complement component C4 in human systemic
lupus erythematosus (SLE): Low copy number is a risk factor for and high copy number is a protective
factor against SLE susceptibility in European Americans. American Journal of Human Genetics, 80(6),
1037–1054. https://doi.org/10.1086/518257
Zhernakova, A., van Diemen, C. C., & Wijmenga, C. (2009). Detecting shared pathogenesis from the shared
genetics of immune-related diseases. Nature Reviews Genetics, 10(1), 43–55. https://doi.org/10.1038/
nrg2489
Zhou, Q., Wang, H., Schwartz, D. M., Stoffels, M., Park, Y. H., Zhang, Y., … Kastner, D. L. (2016). Loss-of-
function mutations in TNFAIP3 leading to A20 haploinsufciency cause an early-onset autoinamma-
tory disease. Nature Genetics, 48(1), 67–73. https://doi.org/10.1038/ng.3459

Bioinformatics Design for
3
Autoimmune Research
3.1 STUDY DESIGN OF AUTOIMMUNE DISEASES
Bioinformatics has become a cornerstone in the investigation of genetic mutations associated with
autoimmune diseases by enabling the integration of high-throughput sequencing data with computational analysis. Approaches such as whole-genome sequencing (WGS), whole-exome sequencing
(WES), and genome-wide association studies (GWAS) generate massive datasets requiring robust
bioinformatics tools for variant detection, annotation, and interpretation. These computational
pipelines allow researchers to identify disease-associated variants, predict their functional consequences, and elucidate their roles in immune dysregulation (McKenna etal., 2010).
Central to this effort are publicly available databases that catalog genetic variation and genedisease associations. Databases like db SNP, dbVar, and ClinVar (NCBI databases) provide curated
information on small-scale and structural variants, while DisGeNET and OMIM focus on genedisease links. Tools like Ensembl, GeneCards, and the ImmPort platform offer gene annotations,
variant effect predictions, and immunological datasets crucial for autoimmune research.
Among the most studied mutations in autoimmune conditions are single nucleotide polymorphisms (SNPs), which may affect protein function, gene regulation, or immune signaling pathways.
Tools such as the Genome Analysis Toolkit (GATK), ANNOVAR, and Variant Effect Predictor
(VEP) facilitate the detection and interpretation of SNPs (McKenna etal., 2010).
Insertion and deletion mutations (Indels) are also implicated in autoimmune pathogenesis. These
mutations, particularly those causing frameshifts, can have deleterious effects on protein synthesis.
Software like Pindel and VarScan is used to detect indels in sequencing data.
Larger-scale genomic changes, including copy number variations (CNVs) and structural variations (SVs), affect gene dosage and regulation. Tools such as CNVkit and ExomeDepth detect
CNVs, while more comprehensive structural changes are analyzed using long-read sequencing data
and algorithms capable of identifying rearrangements.
Splicing mutations, which disrupt the accurate removal of introns and joining of exons, are another
important category. They are detected using splicing prediction tools such as SpliceAI, MaxEntScan,
and rMATS, which help assess the functional impact of mutations on transcript formation.
Beyond sequence mutations, epigenetic modications are increasingly recognized for their role
in autoimmune disease. Techniques such as ATAC-Seq and ChIP-Seq provide insights into chromatin accessibility and transcription factor binding. Bioinformatics tools like Bismark, MACS2,
ArchR, and DiffBind support the analysis of DNA methylation, histone modications, and chromatin state.
Microbiome research, particularly through metagenomics, has revealed that gut microbial communities can modulate immune responses. Tools like QIIME2, MetaPhlAn, and HUMAnN enable
microbial proling and functional annotation, linking dysbiosis to autoimmune pathophysiology.
RNA sequencing (RNA-Seq) has become indispensable for proling gene expression and alternative splicing in autoimmune diseases. Software such as STAR, HISAT2, DESeq2, and edgeR
allow for comprehensive transcriptomic analysis (Anders & Huber, 2010; Dobin etal., 2013).
Small RNA-Seq further enables investigation of microRNAs (miRNAs) and other non-coding
RNAs that regulate gene expression post-transcriptionally. Tools like miRDeep2 and TargetScan
identify and predict miRNA targets relevant to autoimmune disorders.
65 DO I: 10 .1201/ 97810 0 36 85 432-3

66 Bioinformatics of Autoimmune Diseases
Pathway analysis frameworks such as DAVID, KEGG, and Ingenuity Pathway Analysis (IPA)
help researchers understand how mutations perturb biological networks and identify potential therapeutic targets (Huang etal., 2009).
Overall, bioinformatics facilitates a multidimensional view of autoimmune disease genetics, from
mutations in coding sequences to alterations in regulatory and epigenomic elements. As machine
learning and multi-omics integration advance, bioinformatics will continue to play a pivotal role in
the identication of biomarkers and the design of personalized therapeutic interventions.
3.2 STUDY DESIGNS FOR SEQUENCE ANALYSIS IN AUTOIMMUNE DISEASES
Sequence-based analysis has revolutionized the study of autoimmune diseases by uncovering
genetic variants, epigenetic modications, and transcriptional changes associated with disease susceptibility and progression. The study designs used for sequence-based analysis vary depending on
the research question, data availability, and technological advancements. Below is a comprehensive
discussion of the most used study designs.
3.2.1 GENOME-WIDE ASSOCIATION STUDIES
GWAS are powerful tools in genetic epidemiology that identify associations between genetic variants and diseases. These studies have been instrumental in uncovering genetic risk factors for complex disorders, including autoimmune diseases. GWAS involves scanning the genomes of large
cohorts to nd SNPs that are more frequent in individuals with a disease than in healthy controls.
GWAS has signicantly advanced our understanding of autoimmune diseases by identifying
loci associated with immune system dysfunction. GWAS follows a structured methodology to identify genetic variants associated with autoimmune diseases. This process involves multiple stages,
including study design, genotyping, quality control (QC), statistical analysis, and post-GWAS interpretation. Each of these steps is crucial for ensuring the accuracy and reliability of the results.
3.2.1.1 Study Design
The initial and most pivotal phase of a GWAS is the meticulous selection of study participants. This
process involves assembling a well-characterized cohort of individuals diagnosed with a specic
autoimmune disease (cases) and a matched cohort of healthy individuals (controls). Accurate classication of both cases and controls is essential, as misclassication can lead to substantial bias and
compromise the study’s validity. To mitigate such risks, participant inclusion must rely on stringent
clinical criteria and standardized diagnostic guidelines.
Given that autoimmune diseases are typically polygenic, inuenced by numerous genetic loci
each exerting modest effects, large sample sizes are crucial for detecting statistically signicant
associations. Most robust GWAS require the inclusion of thousands to tens of thousands of individuals to ensure adequate statistical power and reduce the likelihood of false positives. The recruitment
of such large cohorts enables researchers to capture subtle genetic variations that may contribute to
disease susceptibility.
Equally important is the demographic matching of cases and controls based on variables such
as age, sex, and ancestry. Inadequate matching can introduce confounding variables that obscure
true genetic signals. Population stratication, in particular, presents a major challenge, as allele
frequency differences due to ancestral background rather than disease status can lead to spurious
associations. To address this, GWAS often incorporate statistical correction techniques such as
principal component analysis (PCA), which helps account for ancestral variation and ensures the
observed associations are truly disease-related (Price etal., 2006).
3.2.1.2 Genotyping and Sequencing
Once participants have been appropriately selected, the next critical phase in a GWAS involves
genomic DNA (gDNA) extraction from biological sources such as peripheral blood or saliva. The

67 Bioinformatics Design for Autoimmune Research
extracted DNA is then subjected to genotyping to identify genetic variants associated with disease
susceptibility. Traditionally, GWAS utilize microarray-based genotyping platforms that employ
high-density SNP arrays containing probes for millions of SNPs across the genome. These arrays
offer a cost-effective and efcient approach to capturing common genetic variants, enabling largescale population studies with high throughput.
However, despite their utility, SNP arrays have inherent limitations. They are primarily designed
to detect common variants (typically with a minor allele frequency >5%) and may miss rare variants
or structural changes that also contribute to disease phenotypes (Manolio etal., 2009). To address
this gap, many contemporary studies complement or replace SNP arrays with next-generation
sequencing (NGS) technologies, such as WGS and WES.
WGS provides a comprehensive assessment of the entire genome, capturing both coding and
non-coding regions, including rare single nucleotide variants (SNVs), structural variants, and CNVs
that may be implicated in autoimmune pathogenesis. In contrast, WES targets only the exonic
regions, which constitute about 1–2% of the genome but account for a signicant proportion of
known disease-causing mutations (Ng etal., 2009). While WES is more cost-effective than WGS,
it excludes regulatory and intergenic regions that may also harbor functionally relevant variants.
The selection of a genotyping strategy depends on the study’s objectives, available funding, and the
computational capacity for downstream data processing and analysis.
3.2.1.3 Quality Control
After genotyping is completed, a critical phase in GWAS involves applying stringent QC procedures to ensure the reliability and accuracy of the genetic data. Without rigorous QC, technical
artifacts or biological outliers can lead to false-positive associations and undermine the validity
of downstream analyses. At the sample level, standard QC steps include removing duplicate or
contaminated samples, verifying reported sex using X-chromosome heterozygosity, and identifying
cryptic relatedness or unexpected familial relationships. Individuals with high levels of missing
genotype data are also excluded, as they may indicate poor DNA quality or processing errors.
To control for confounding due to population stratication (systematic ancestry differences
between cases and controls), PCA or multidimensional scaling is often performed. These methods
enable researchers to detect and adjust for subtle population structure, thereby reducing the risk of
spurious associations arising from ancestry differences rather than true disease susceptibility loci
(Price etal., 2006).
At the variant level, additional QC lters are applied. SNPs with low genotyping call rates, typically below 95%, are discarded due to concerns about data completeness and accuracy. Furthermore,
SNPs that deviate signicantly from Hardy-Weinberg equilibrium (HWE) in the control group
are excluded, as such deviations may indicate genotyping errors, sample contamination, or hidden
population substructure. The HWE principle assumes a randomly mating population free from
selection, mutation, or migration, and is a foundational metric for validating genotypic distributions
in GWAS datasets. Variants with extremely low minor allele frequencies (MAFs), often below 1 or
5%, are also removed, since rare alleles may lack sufcient statistical power to demonstrate meaningful associations in most GWAS frameworks.
Together, these QC measures are essential for ensuring the integrity of GWAS ndings. They
minimize noise, reduce confounding, and increase the likelihood that detected associations reect
true biological relationships rather than technical anomalies or population biases.
3.2.1.4 Statistical Association Testing
Once high-quality genotype data is secured, the next phase of GWAS involves statistical association
testing to identify genetic variants that contribute to disease risk. The standard approach in GWAS
relies on single-marker tests, where each SNP is independently tested for its association with the
phenotype. In case-control designs, logistic regression is the preferred model to estimate the association between SNPs and binary outcomes, such as disease status. For continuous traits, linear

68 Bioinformatics of Autoimmune Diseases
regression models are used. These models typically adjust for relevant covariates such as age, sex,
and principal components derived from genotype data to control for confounding due to population
stratication (Price etal., 2006).
To account for the polygenic nature of complex diseases and subtle population structure, more
advanced statistical frameworks have been introduced. Mixed linear models (MLMs), as implemented in software tools like GEMMA and GCTA, incorporate a random effect for genetic relatedness, thereby improving control over confounding and enhancing statistical power. Bayesian
methods provide probabilistic estimates of variant effects and are particularly useful for integrating
prior knowledge or dealing with small effect sizes. In the realm of high-dimensional SNP datasets,
penalized regression models such as LASSO (Least Absolute Shrinkage and Selection Operator)
and elastic net are used to identify a sparse subset of informative variants. Machine learning techniques, including ensemble models and neural networks, have also emerged as promising tools for
capturing nonlinear interactions and improving phenotype prediction, particularly in the construction of polygenic risk scores (PRS).
Due to the enormous number of statistical tests conducted, often in the millions, correcting for multiple testing is essential to mitigate the risk of false positives. The Bonferroni correction, which sets
a genome-wide signicance threshold by dividing the desired alpha level (e.g., 0.05) by the number
of tests performed, is commonly applied but considered overly conservative. Alternative approaches
such as the false discovery rate (FDR), particularly the Benjamini-Hochberg procedure, offer a less
stringent yet robust method for identifying true associations while controlling for type I errors.
Genotype imputation is another key step in modern GWAS pipelines. Since SNP arrays do not
capture all genetic variation, imputation lls in unobserved genotypes using reference datasets
such as the 1000 Genomes Project or the Haplotype Reference Consortium. Tools like IMPUTE2
and Minimac3 apply statistical models to infer missing genotypes, thereby increasing variant density and enhancing discovery power. Imputation also enables cross-study comparisons and metaanalyses by harmonizing datasets genotyped on different platforms.
3.2.1.5 Post-GWAS Analysis
After identifying disease-associated SNPs through GWAS, the next crucial step is to determine
their functional signicance and biological relevance. This process begins with ne mapping, which
aims to pinpoint the true causal variants within statistically associated genomic regions. Because
GWAS often identify broad loci containing many correlated variants due to linkage disequilibrium, ne mapping helps distinguish the specic variants that directly inuence disease from those
that are merely correlated. This involves integrating high-resolution statistical models, functional
genomic annotations, and, when feasible, experimental validation.
Functional annotation is essential for interpreting the potential roles of identied variants,
especially those located in non-coding regions of the genome. Resources such as the ENCODE
(Encyclopedia of DNA Elements) project and the Genotype-Tissue Expression (GTEx) project provide comprehensive datasets on transcription factor binding, chromatin accessibility, histone modications, and gene expression patterns across tissues. These tools help researchers understand how
disease-associated SNPs may inuence gene regulation rather than protein coding. Notably, many
autoimmune-associated SNPs are enriched in regulatory elements active in immune cells, pointing
to dysregulated gene expression as a key mechanism of disease.
Expression quantitative trait loci (eQTL) mapping serves as a bridge between genetic variation
and gene expression by identifying SNPs that correlate with transcript abundance. eQTL analysis
allows researchers to link non-coding variants to specic gene regulatory effects in disease-relevant
tissues, particularly immune cell types in autoimmune diseases. Integrating GWAS signals with
eQTL data improves the interpretability of non-coding associations and facilitates the discovery of
novel gene-disease relationships.
To gain a systems-level understanding of genetic ndings, pathway and gene network analyses
are employed. These approaches identify clusters of genes and variants that participate in common

69 Bioinformatics Design for Autoimmune Research
biological functions, such as cytokine signaling, antigen processing, and T-cell differentiation,
core pathways implicated in autoimmunity. Such enrichment analyses help prioritize mechanistic
hypotheses and therapeutic targets by highlighting dysregulated pathways across multiple autoimmune conditions.
In addition, PRS are constructed using GWAS summary statistics to estimate an individual’s
cumulative genetic susceptibility to disease. PRS aggregate the effects of multiple independent risk
variants, each contributing a small effect size, into a single quantitative score. While still an emerging tool in clinical settings, PRS hold potential for early disease prediction, stratication of at-risk
individuals, and personalized intervention strategies in autoimmune disorders.
3.2.2 WHOLE-EXOME SEQUENCING
WES is a powerful high-throughput sequencing technique designed to target and sequence the
protein-coding regions of the genome, collectively referred to as the exome. Although exons represent
only about 1–2% of the human genome, they harbor approximately 85% of known disease-causing
mutations (Bamshad etal., 2011; Ng etal., 2009). This enrichment makes WES a cost-effective and
efcient alternative to WGS for identifying clinically relevant genetic variants. In the context of
autoimmune diseases (including rheumatoid arthritis (RA), systemic lupus erythematosus (SLE),
multiple sclerosis (MS), and type 1 diabetes (T1D)), WES has been instrumental in revealing both
rare inherited mutations and de novo variants that may contribute to immune dysregulation and
disease susceptibility.
Unlike GWAS, which primarily detect common variants with small effect sizes, WES enables
the identication of rare coding mutations that are more likely to have signicant functional consequences. This makes it especially valuable for investigating monogenic autoimmune syndromes,
as well as complex polygenic disorders where rare variants can contribute to disease heterogeneity. By focusing on exonic sequences, WES provides detailed insights into mutations affecting
key immune-related genes, including those encoding cytokines, transcription factors (e.g., FOXP3,
STAT3), and immune checkpoint molecules (e.g., CTLA4, PTPN22), many of which play critical
roles in maintaining immune tolerance and controlling inammation (de Lange etal., 2017).
The workow of WES encompasses several essential stages. These include careful study design
and phenotype denition, DNA extraction, library preparation, exome capture using hybridization
probes, high-throughput sequencing (often via platforms like Illumina), and a robust bioinformatics
pipeline involving QC, alignment to a reference genome, variant calling, and functional annotation
(McKenna etal., 2010). Each of these steps is vital to ensure data accuracy and to minimize false-
positive ndings. Sophisticated analytical tools such as GATK, ANNOVAR, and VEP are commonly used to identify potentially pathogenic variants and prioritize candidates for downstream
validation and functional studies.
The insights provided by WES are increasingly informing precision medicine approaches in
autoimmunity, enabling the identication of patient-specic genetic risk proles, guiding therapeutic decisions, and laying the groundwork for targeted interventions. As sequencing technologies continue to improve in depth, speed, and affordability, WES remains a cornerstone of genetic
research into the mechanisms of autoimmune diseases.
3.2.2.1 Study Design and Sample Selection
A well-designed study is fundamental to obtaining reliable and interpretable results in WES analysis, particularly when applied to autoimmune disease research. The rst step involves the careful
selection of an appropriate cohort, typically composed of individuals diagnosed with a specic autoimmune disorder (cases) and matched healthy individuals (controls). Alternatively, family-based
study designs are employed, particularly in rare or monogenic forms of autoimmunity, wherein
affected individuals are sequenced alongside their unaffected relatives to facilitate the identication
of inherited or de novo variants (Bamshad etal., 2011).

70 Bioinformatics of Autoimmune Diseases
Accurate and detailed phenotyping is essential to ensure that the cases genuinely reect the disease of interest. Autoimmune diseases are often clinically heterogeneous, with patients displaying
a wide spectrum of symptoms, disease courses, and treatment responses. As such, comprehensive
clinical data (including serological markers, imaging results, and age of disease onset) are critical
for rening the phenotype and enabling stratied analysis. Environmental exposures and ancestry
must also be accounted for, as these factors can signicantly impact both disease risk and variant
interpretation (de Lange etal., 2017). Ancestry-aware analysis is especially important in WES to
prevent spurious associations arising from population stratication.
Since WES targets rare and novel coding variants, the required sample size differs from that
of GWAS. Whereas GWAS often demands tens of thousands of participants to detect common
variants with small effect sizes, WES can yield informative results from much smaller cohorts,
particularly in studies of Mendelian autoimmune syndromes. However, in the context of complex
autoimmune diseases, larger sample sizes are still necessary to achieve sufcient statistical power
and to distinguish true disease-causing variants from background genetic noise. Ultimately, a robust
study design that integrates precise phenotyping, appropriate controls, and careful consideration of
genetic background is essential to maximize the utility of WES in autoimmune disease research.
3.2.2.2 DNA Extraction and Library Preparation
Once the study cohort is established, the extraction of high-quality gDNA from biological specimens (typically whole blood or saliva) is a critical step in ensuring the success of WES. The quality
and quantity of DNA are assessed using both spectrophotometry (e.g., NanoDrop) and uorometric
quantication (e.g., Qubit), which provide complementary measures of DNA purity and concentration. High-molecular-weight DNA is especially desirable, as it yields better coverage and sequencing delity.
Library preparation begins with the fragmentation of gDNA into smaller pieces, usually in the
range of 150–300 bp, suitable for high-throughput sequencing platforms. These DNA fragments
are then ligated to platform-specic adapters that allow for subsequent amplication and sequencing. Increasingly, unique molecular identiers (UMIs) are incorporated during this step to improve
downstream error correction and enhance the accuracy of variant detection, especially in lowfrequency or rare mutations.
3.2.2.3 Exome Capture and Target Enrichment
The dening feature of WES lies in its ability to selectively enrich coding regions of the genome.
This is achieved through a process called exome capture, where DNA fragments are hybridized
with biotinylated oligonucleotide probes designed to be complementary to exonic sequences. The
probes selectively bind to coding regions, while non-target sequences (including intronic and intergenic DNA) are removed through a series of wash steps.
Several commercial exome capture kits are commonly used in WES workows, including
Agilent SureSelect, Illumina Nextera Rapid Capture, and Roche NimbleGen SeqCap EZ. The
choice of kit depends on multiple factors such as desired target coverage, inclusion of untranslated regions (UTRs), cost, and compatibility with sequencing instrumentation (Clark et al., 2011).
Newer-generation capture kits often expand target regions to include exon-adjacent regulatory elements, which may harbor functionally signicant variants relevant to autoimmune diseases.
Following enrichment, the captured DNA is amplied via polymerase chain reaction (PCR)
and puried to eliminate excess reagents and contaminants. The nal sequencing library is then
quantied, typically using quantitative PCR (qPCR) or uorometric methods, and assessed for size
distribution and concentration before sequencing is initiated.
3.2.2.4 High-Throughput Sequencing
Sequencing of captured exonic regions is carried out using NGS platforms, such as Illumina
NovaSeq, HiSeq, and the more budget-conscious NextSeq series. These platforms are well-suited

71 Bioinformatics Design for Autoimmune Research
for high-throughput sequencing, producing millions to billions of short reads (typically ranging
from 100 to 150 bp in length) that represent the exome of each individual (Goodwin etal., 2016).
The widespread adoption of Illumina technology is largely due to its high accuracy, scalability, and
cost-effectiveness for WES applications.
A key metric in WES is sequencing depth, dened as the average number of times each base pair
in the target region is read during sequencing. To ensure condent variant detection and minimize
the inuence of random errors, a depth of at least 50–100× is generally recommended (Clark etal.,
2011). Higher coverage improves sensitivity, particularly for identifying rare and low-frequency
variants, and enhances the reliability of downstream variant calling.
The raw data generated from sequencing is output in FASTQ format, which contains both the
nucleotide sequence and associated base quality scores for each read. Before alignment and variant
detection, these raw reads undergo a series of preprocessing steps, including adapter trimming and
quality ltering. Tools such as Trimmomatic or FastQC are commonly used to eliminate low-quality
bases and remove residual adapter sequences that can interfere with mapping and variant calling.
This initial QC step is essential to reduce artifacts and ensure the integrity of downstream analysis.
3.2.2.5 Alignment and Variant Calling
Following sequencing, the rst critical step in the analysis pipeline is read alignment to a reference
human genome, typically GRCh38 or the earlier GRCh37/hg19 build. Alignment is performed using
highly efcient algorithms such as the Burrows–Wheeler Aligner (BWA) or Bowtie2, both of which
map short sequencing reads to the reference genome with high speed and accuracy. Proper alignment
is essential, as it forms the basis for all downstream variant detection. Key quality metrics such as
mapping rate, duplication rate, and coverage uniformity are evaluated to ensure alignment integrity.
Once alignment is complete, the pipeline proceeds to variant calling, which involves detecting
genomic alterations, including SNVs and indels. Several tools are used for this purpose:
GATK is the gold standard for germline variant calling, offering a robust pipeline with bestpractice workows for variant discovery (McKenna etal., 2010).
FreeBayes provides a Bayesian approach to identifying variants and is particularly useful for
multi-sample analysis.
DeepVariant, developed by Google, uses a deep learning framework to improve the precision of
variant calls by modeling sequencing data as images.
After calling variants, a variant ltering step is implemented to reduce false positives. Variants
are evaluated based on metrics such as read depth, genotype quality, and strand bias. Low-condence
variants are excluded to enhance specicity.
Finally, variant annotation is performed to interpret the biological relevance of each identied
mutation. Publicly available databases such as dbSNP, 1000 Genomes, and gnomAD are used to
annotate known variants, determine allele frequencies in global populations, and differentiate common polymorphisms from potentially disease-associated rare mutations.
3.2.2.6 Functional Annotation and Interpretation
Identifying genetic variants is only the beginning; the subsequent and more complex task involves
interpreting their biological and functional relevance. Variant annotation tools such as ANNOVAR,
VEP, and SnpEff are widely used to evaluate the genomic context of each variant and predict its
effect on gene function, coding potential, and transcript structure (McLaren et al., 2016). These
tools integrate information from gene models and known polymorphism databases to classify variants as synonymous, missense, nonsense, splice site, or frameshift, among others.
To assess the potential pathogenicity of identied mutations, several predictive algorithms are
employed. PolyPhen-2, SIFT, and CADD (Combined Annotation Dependent Depletion) are commonly used to prioritize variants that may be deleterious based on evolutionary conservation, protein structure modeling, and functional domain disruption (Adzhubei etal., 2010). Variants that

72 Bioinformatics of Autoimmune Diseases
result in premature stop codons, affect canonical splice sites, or lead to amino acid substitutions in
critical functional domains are agged for further analysis.
In the context of autoimmune diseases, annotated variants are cross-referenced with immunerelated databases such as ImmunoBase, the GWAS Catalog, and DisGeNET, which compile known
associations between immune genes and autoimmune conditions (Buniello etal., 2019). This inte-
grative approach enables researchers to determine whether a variant lies within an established
immune pathway or a gene known to inuence immune tolerance, cytokine signaling, or antigen
presentation.
To validate computational predictions and elucidate the functional consequences of candidate
variants, follow-up studies may be performed. These include gene expression analysis using RNASeq, CRISPR/Cas9-based genome editing to test variant effects in cell or animal models, and protein modeling to evaluate structural perturbations caused by amino acid changes. Such validation is
essential to bridge the gap between genotype and phenotype and to identify pathogenic mutations
with therapeutic potential.
3.2.3 WHOLE-GENOME SEQUENCING
WGS offers a comprehensive method for analyzing an individual’s entire genome, encompassing
both protein-coding (exonic) and non-coding (intronic and intergenic) regions. Unlike WES, which
restricts its scope to exonic regions, WGS provides an unbiased and complete assessment of genetic
variation across the genome. This broad scope is especially advantageous for studying complex diseases like autoimmune disorders, which involve a combination of polygenic inheritance, epigenetic
alterations, and environmental exposures.
Autoimmune diseases are known to be highly polygenic, meaning that numerous genetic variants, often with small effect sizes, contribute to disease susceptibility. Many of these variants reside
outside coding regions, where they inuence regulatory processes such as gene expression, chromatin accessibility, and transcription factor binding. WGS enables researchers to detect a broad
spectrum of genetic variation, including SNVs, indels, CNVs, SVs, and non-coding regulatory
mutations, which may collectively or individually disrupt immune system function.
By offering a complete catalog of genomic variation, WGS signicantly enhances our ability to discover novel disease-associated loci and rene PRS for improved genetic risk prediction.
Moreover, WGS allows for the identication of rare and population-specic variants that might
be missed by array-based genotyping or WES. These capabilities contribute to the advancement
of precision medicine in autoimmune disease by identifying personalized therapeutic targets and
stratifying patients based on their genomic proles.
Despite its potential, WGS presents several challenges. It is more expensive than targeted
approaches, generates vast quantities of data, and requires sophisticated computational infrastructure
for data processing and interpretation (Goodwin etal., 2016). However, with decreasing sequencing
costs and advances in cloud-based and high-performance computing, WGS is increasingly accessible and is expected to become integral to future autoimmune research and clinical diagnostics.
The WGS workow involves several essential steps: rigorous study design, high-quality DNA
extraction, NGS, raw data preprocessing, variant calling, annotation, and downstream functional
interpretation. Each stage must be carefully executed to ensure the reliability and reproducibility of
ndings in autoimmune genomics.
3.2.3.1 Study Design and Sample Selection
A well-structured study design is essential for generating reliable and biologically meaningful
results in WGS studies focused on autoimmune diseases. The foundational step involves the careful selection of an appropriate study cohort. Typically, this includes individuals diagnosed with a
well-characterized autoimmune disease (cases) and a group of healthy, ethnically matched individuals (controls). Accurate phenotyping and the use of standardized diagnostic criteria are critical to

73 Bioinformatics Design for Autoimmune Research
TotalNumberofBases Sequenced
G
Sequencin
Genomesize
NL×
De
minimize heterogeneity and misclassication, both of which can obscure true genetic associations
(Buniello etal., 2019).
In addition to case-control designs, family-based WGS studies have emerged as a powerful strategy to investigate inherited and de novo mutations. Sequencing affected individuals alongside their
unaffected relatives (such as parents or siblings) enables researchers to distinguish between inherited variants and novel mutations that may arise spontaneously and contribute to disease onset. This
approach is particularly useful for identifying rare, high-impact variants that may not be detectable
in population-level studies.
3.2.3.2 DNA Extraction and Library Preparation
Once participants are selected, the next critical step in WGS involves the extraction of high-quality
gDNA from biological specimens, typically peripheral blood or saliva. The integrity and concentration of the DNA are assessed using spectrophotometry (e.g., NanoDrop), uorometry (e.g., Qubit),
and gel electrophoresis to conrm the presence of high-molecular-weight DNA suitable for downstream applications.
Library preparation begins with the fragmentation of gDNA into smaller segments, usually ranging from 300 to 500 bp, using enzymatic digestion or mechanical shearing via sonication. These
fragments are then ligated to platform-specic sequencing adapters, which facilitate amplication
and allow the fragments to be read during sequencing. Increasingly, UMIs are incorporated during
adapter ligation to minimize amplication biases and sequencing artifacts, thereby improving variant calling accuracy, especially in low-frequency variants.
To promote uniform genome coverage, the prepared libraries undergo size selection and purication to remove undesired fragment sizes and contaminants. The nal libraries are quantied
using uorometric methods and quality-checked through automated electrophoresis or qPCR before
loading onto high-throughput sequencing platforms such as the Illumina NovaSeq or HiSeq series.
Given the polygenic nature of autoimmune diseases, achieving adequate statistical power requires
large sample sizes. Unlike monogenic conditions, where single variants have large effect sizes,
autoimmune diseases arise from the interplay of numerous genetic variants with modest effects.
Therefore, population-based WGS studies often include thousands of participants to detect statistically signicant associations. Additionally, rigorous matching of cases and controls based on factors
such as age, sex, ancestry, and relevant environmental exposures is necessary to reduce potential
confounders and population stratication.
3.2.3.3 High-Throughput Whole-Genome Sequencing
Sequencing is performed using NGS platforms such as Illumina NovaSeq, HiSeq, or PacBio. The
choice of sequencing platform depends on study requirements, such as read length, coverage depth,
and cost.
A critical parameter in WGS is sequencing depth, which refers to the average number of times
each base in the genome is read during sequencing (a measure of how many times a given nucleotide
in the genome has been sequenced). The basic formula of the sequencing depth is
gDepth (Coverage) =
Or more specically:
pth = (3.2)
where N is the number of reads, L is the length of each read (in base pairs), and G is the size of the
genome or target region (in base pairs).
(3.1)
Соседние файлы в папке Библиотека им академика М.И. Перельмана
