Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
94 Bioinformatics of Autoimmune Diseases
enable the identication of biochemical activities and metabolic networks associated with health or
disease.
Statistical analyses are used to assess microbial diversity and identify disease-associated patterns. Alpha diversity metrics, including Shannon and Simpson indices, measure the richness and
evenness of species within a single sample. Beta diversity metrics, such as Bray-Curtis dissimilarity
and UniFrac distances, quantify differences in microbial community structure between groups.
These diversity measures are often visualized through principal coordinate analysis (PCoA) or nonmetric multidimensional scaling (NMDS).
Emerging computational approaches, including machine learning models such as random forests
and deep neural networks, are increasingly used to identify microbial signatures that discriminate
between disease and healthy states. These methods enhance prediction accuracy and uncover complex microbial interactions and host associations.
3.2.7.5 Functional Insights and Integration with Autoimmune Disease Research
After identifying disease-associated microbial patterns, the next critical step is understanding their
functional role in the pathogenesis of autoimmune diseases. Microbial metabolites profoundly inuence host immune regulation. Notably, SCFAs (including butyrate, acetate, and propionate) are produced by gut commensals through the fermentation of dietary bers. These SCFAs play essential
roles in maintaining intestinal barrier integrity, modulating inammatory responses, and promoting
the differentiation and function of regulatory T cells (Tregs), which are critical for immune tolerance (Arpaia etal., 2012). In contrast, microbial-derived components such as lipopolysaccharides
(LPS) can activate toll-like receptor 4 (TLR4)-mediated signaling, leading to systemic inammation and the exacerbation of autoimmune processes.
To gain a deeper understanding of how microbial alterations inuence host biology, researchers
are increasingly adopting multi-omics integration strategies. By combining microbiome data with
host transcriptomics, proteomics, and metabolomics, it is possible to unravel how microbial-derived
signals modulate gene expression, protein activity, and metabolic networks involved in immune
regulation and disease progression. These integrative approaches offer a systems-level perspective
of host–microbe interactions and are particularly valuable for disentangling complex, multifactorial
autoimmune conditions.
The therapeutic potential of microbiome-based interventions is an emerging area of interest in
autoimmune disease management. FMT has demonstrated clinical benet in restoring microbial
diversity and reducing inammation, particularly in patients with ulcerative colitis. Moreover, the
use of probiotics (benecial live microbes) and prebiotics (compounds that promote the growth
of benecial microbes) is under investigation for their ability to modulate immune responses and
attenuate autoimmune pathology. In parallel, dietary interventions designed to enrich benecial
microbial metabolites or suppress inammatory taxa present a promising avenue for non-invasive,
long-term immunomodulation.
3.3 SUMMARY
This chapter offers a detailed and structured overview of the bioinformatics frameworks and
sequencing-based study designs essential for exploring the molecular and genetic basis of autoimmune diseases. The chapter begins by framing bioinformatics as a crucial discipline that
combines computational tools with high-throughput sequencing to identify and interpret genetic
variations associated with autoimmune disorders. It explains how technologies such as WGS,
WES, and GWAS are deployed to detect disease-associated mutations, including SNPs, indels,
SVs, and CNVs. These genetic changes can disrupt immune signaling pathways, affect protein
function, and contribute to immune dysregulation. Tools like GATK, A N NO VA R , SpliceAI, and
CNVkit are presented as instrumental in the detection, annotation, and functional prediction of
these variants.

95 Bioinformatics Design for Autoimmune Research
The chapter continues by integrating epigenomic and metagenomic insights into autoimmune
research. Epigenetic alterations, such as DNA methylation and histone modications, are described
as non-genetic regulators of gene expression, and their analysis is supported by tools like MACS2,
Bismark, and deepTools. Metagenomic approaches are highlighted for their ability to assess the role
of the microbiome in autoimmunity, particularly how microbial dysbiosis can inuence immune
function. Tools such as Kraken, MetaPhlAn, and HUMAnN are introduced as key resources for
microbial proling and functional analysis, showing how host–microbiome interactions may trigger
or exacerbate autoimmune responses.
A signicant portion of the chapter is devoted to describing the design and execution of GWAS.
The GWAS section outlines how large cohorts of cases and controls are genotyped to identify
SNPs associated with disease phenotypes. It elaborates on the necessity of QC procedures, statistical correction for population stratication, and the importance of multiple testing correction using
methods such as the Bonferroni adjustment or FDR. Post-GWAS analysis includes ne mapping,
functional annotation using databases like ENCODE and GTEx, and expression quantitative trait
loci (eQTL) analysis to link variants with gene expression. The chapter emphasizes that GWAS
ndings can be further distilled into PRS, offering predictive insights into genetic susceptibility.
The chapter then transitions to WES as a cost-effective strategy to uncover rare and potentially pathogenic mutations conned to the protein-coding regions of the genome. WES is detailed
step-by-step, from study design and sample selection to sequencing and functional annotation. It
describes how tools like BWA, GATK, and DeepVariant are used for alignment and variant calling,
while annotation tools like VEP and PolyPhen-2 assess variant pathogenicity. The power of WES
lies in its ability to uncover rare de novo or inherited coding variants in immune-related genes,
providing mechanistic insights into disease etiology.
WGS is described as a more comprehensive approach that captures all genomic elements,
including non-coding regulatory regions. It allows for the detection of structural variants, enhancer
mutations, and long-range chromatin interactions that may inuence gene regulation. The chapter
outlines the importance of sequencing depth, explaining the formulas and coverage levels required
for different research goals. Functional annotation and pathway analysis are emphasized for interpreting WGS data, particularly in the context of autoimmune diseases, where multiple weak-effect
variants interact to shape disease phenotypes.
RNA sequencing (RNA-Seq) is presented as an indispensable method for transcriptome analysis.
The chapter explains how RNA-Seq is used to quantify gene expression, detect alternative splicing,
and identify non-coding RNAs, all of which are critical to understanding autoimmune pathogenesis. It discusses sample handling, library preparation, and the impact of transcriptome size on
sequencing depth. Normalization techniques, including TPM, RPKM, and TMM, are introduced to
account for differences in library size and gene length. Statistical modeling of RNA-Seq data, particularly the use of negative binomial models as implemented in DESeq2 and edgeR, is explained in
detail. Limma-voom is introduced as an alternative for log-transformed data with empirical Bayes
variance shrinkage. Functional enrichment tools and pathway databases are used to interpret gene
expression changes, linking them to immune processes.
The chapter advances into the realm of single-cell RNA sequencing (scRNA-Seq), underscoring its transformative impact on resolving cellular heterogeneity in autoimmune diseases. It covers techniques for isolating individual cells, barcoding with UMIs, and preparing cDNA libraries
using droplet-based microuidics systems such as 10× Genomics Chromium. With read depth and
sequencing quality addressed, the chapter discusses computational methods like UMAP, PCA,
and clustering for analyzing single-cell data. It also introduces pseudotime trajectory analysis and
ligand–receptor interaction modeling as means of uncovering dynamic immune cell behavior and
intercellular communication in disease contexts.
EWAS are presented next, focusing on how epigenetic changes (such as DNA methylation and
histone modications) can be proled across the genome-using techniques like WGBS, RRBS, and
Innium methylation arrays. The chapter details the statistical modeling of ChIP-Seq data using

96 Bioinformatics of Autoimmune Diseases
negative binomial distributions to assess peak enrichment, differential binding, and downstream
functional interpretation through tools like HOMER, GREAT, and ChIPseeker. The integration
of chromatin accessibility data from ATAC-Seq is emphasized for identifying regulatory elements
involved in immune gene activation.
Finally, the chapter delves into metagenomic and microbiome sequencing, highlighting the
importance of microbial ecology in shaping autoimmune risk. It distinguishes between 16S rRNA
gene sequencing for taxonomic proling and WGSS for comprehensive microbial and functional
analysis. Tools such as QIIME2, Kraken2, and MetaPhlAn are introduced, alongside functional
annotation platforms like HUMAnN and KEGG. The chapter outlines methods to evaluate microbial diversity and explores how microbiome-derived metabolites, such as SCFAs, inuence immune
regulation.
In summary, Chapter 3 provides a richly detailed, methodologically grounded framework for
bioinformatics analysis in autoimmune disease research. It unies genetic, transcriptomic, epigenomic, and microbial dimensions through computational tools and statistical modeling, offering
an integrated systems biology approach to unravel the complexities of autoimmunity. Through
its comprehensive coverage, the chapter establishes a solid foundation for future studies aimed at
understanding pathogenesis, identifying biomarkers, and developing precision medicine strategies
for autoimmune disorders.
BIBLIOGRAPHY
Adzhubei, I. A., et al. (2010). A method and server for predicting damaging missense mutations. Nature
Methods, 7(4), 248–249. https://doi.org/10.1038/nmeth0410-248
Anders, S., & Huber, W. (2010). Differential expression analysis for sequence count data. Genome Biology, 11,
R106. https://doi.org/10.1186/gb -2010-11-10-r106
Andrews, S. (2010). FastQC: A quality control tool for high throughput sequence data. http://www.bioinfor-
matics.babraham.ac.uk/projects/fastqc
Arpaia, N., Green, J. A., Moltedo, B., Arvey, A., Hemmers, S., Yuan, S., Treuting, P. M., & Rudensky, A. Y.
(2015). A distinct function of regulatory T cells in tissue protection. Cell, 162(5), 1078–1089. https://doi.
org/10.1016/j.cell.2015.08.021
Bamshad, M. J., et al. (2011). Exome sequencing as a tool for Mendelian disease gene discovery. Nature
Reviews Genetics, 12(11), 745–755. https://doi.org/10.1038/nrg3031
Beghini, F., et al. (2021). Integrating taxonomic, functional, and strain-level proling of diverse microbial
communities with bioBakery 3. eLife, 10, e65088. https://doi.org/10.7554/eLife.65088
Buenrostro, J. D., et al. (2015). ATAC-seq: A method for assaying chromatin accessibility genome-wide.
Current Protocols in Molecular Biology, 109, 21–29. https://doi.org/10.1002/0471142727.mb2129s109
Buniello, A., et al. (2019). The NHGRI-EBI GWAS Catalog of published genome-wide association stud-
ies, targeted arrays and summary statistics. Nucleic Acids Research, 47(D1), D1005–D1012. https://doi.
org/10.1093/nar/gky1120 https://doi.org /10.1093/nar/gky1120
Clark, M. J., et al. (2011). Performance comparison of exome DNA sequencing technologies. Nature
Biotechnology, 29(10), 908–914. https://doi.org/10.1038/nbt.1975
Conesa, A., et al. (2016). A survey of best practices for RNA-seq data analysis. Genome Biology, 17, 13. https://
doi.org/10.1186/s13059-016-0881-8
de Lange, K. M., et al. (2017). Genome-wide association study implicates immune activation of multiple integrin
genes in inammatory bowel disease. Nature Genetics, 49(2), 256–261. https://doi.org/10.1038/ng.3760
Dobin, A., et al. (2013). STAR: Ultrafast universal RNA-seq aligner. Bioinformatics, 29(1), 15–21. https://doi.
org/10.1093/bioinformatics/ bts635
Feil, R., & Fraga, M. F. (2012). Epigenetics and the environment. Nature Reviews Genetics, 13(2), 97–109.
https://doi.org/10.1038/nrg3142
Goodwin, S., et al. (2016). Coming of age: ten years of next-generation sequencing technologies. Nature
Reviews Genetics, 17(6), 333–351. https://doi.org/10.1038/nrg.2016.49
Grosselin, K., et al. (2019). High-throughput single-cell ChIP-seq identies heterogeneity of chromatin states
in breast cancer. Nature Genetics, 51(6), 1060–1066. https://doi.org/10.1038/s41588- 019-042 4-9

97 Bioinformatics Design for Autoimmune Research
Haque, A., et al. (2017). A practical guide to single-cell RNA-sequencing. Genome Medicine, 9, 75. https://
doi.org/10.1186/s13073-017-0467-4
Huang, D. W., et al. (2009). DAVID bioinformatics resources. Nature Protocols, 4(1), 44–57. https://doi.
org/10.1038/nprot.2008.211
Love, M. I., et al. (2014). Moderated estimation of fold change and dispersion. Genome Biology, 15, 550.
ht tps://doi.org/10.1186/s130 59- 014- 0550-8
Manolio, T. A., et al. (2009). Finding the missing heritability of complex diseases. Nature, 461(7265), 747–753.
https://doi.org/10.1038/nature08494
Martin, M. (2011). Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.
Journal, 17(1), 10–12. htt ps://doi.org/10.14806/ej.17.1.200
McKenna, A., et al. (2010). The genome analysis toolkit. Genome Research, 20(9), 1297–1303. https://doi.
org/10.1101/gr.107524.110
M cL a re n, W. , et al. (2016). The ensembl variant effect predictor. Genome Biology, 17, 122. https://doi.
org/10.1186/s13059-016-0974-4
Michels, K. B., et al. (2013). Recommendations for EWAS design and analysis. Nature Methods, 10(10),
949–955. https://doi.org/10.1038/nmeth.2632
Ng, S. B., et al. (2009). Exome sequencing identies the cause of a Mendelian disorder. Nature Genetics, 42(1),
30–35. https://doi.org/10.1038/ng.499
Picelli, S., et al. (2014). Full-length RNA-seq from single cells using Smart-seq2. Nature Protocols, 9(1),
171–181. https://doi.org/10.1038/nprot.2014.006
Price, A. L., et al. (2006). Principal components analysis corrects for stratication. Nature Genetics, 38(8),
904–909. https://doi.org/10.1038/ng1847
Scher, J. U., et al. (2012). Periodontal disease and the oral microbiota in new-onset rheumatoid arthritis.
Arthritis & Rheumatism, 64(10), 3083–3094. https://doi.org/10.1002/art.34539
Wang, E. T., et al. (2009). Alternative isoform regulation in human tissue transcriptomes. Nature, 456(7221),
470–476. https://doi.org/10.1038/nature07509
Zheng, G. X. Y., et al. (2017). Massively parallel digital transcriptional proling of single cells. Nature
Communications, 8, 14049. https://doi.org/10.1038/ncomms14049

Bioinformatics Databases
4
4.1 BIOLOGICAL DATABASES FOR AUTOIMMUNE DISEASE RESEARCH
Bioinformatics databases have become indispensable tools in the study of autoimmune diseases,
offering extensive repositories of genetic, proteomic, transcriptomic, epigenetic, and functional
data. These resources facilitate the identication of disease-associated variants, analysis of gene
expression patterns, exploration of regulatory elements, and investigation of molecular pathways
that contribute to immune system dysfunction. The integration of high-throughput sequencing
technologies with bioinformatics has signicantly advanced our understanding of autoimmune
disorders and the development of precision medicine approaches tailored to individual genetic
proles.
One of the foundational databases in autoimmune disease research is the National Center for
Biotechnology Information (NCBI) GenBank. This resource provides access to an extensive collection of nucleotide sequences from a diverse array of organisms, allowing researchers to retrieve
genetic sequences linked to disease susceptibility. GenBank is particularly valuable for comparative genomic studies, aiding in the identication of genetic variations associated with autoimmune
diseases. The integration of GenBank with other NCBI resources, including database of single
nucleotide polymorphism (dbSNP) and ClinVar, enhances its utility by offering insights into single
nucleotide polymorphisms (SNPs) and their clinical implications. While GenBank is comprehensive, its vast dataset can sometimes be overwhelming, requiring signicant expertise to lter and
analyze relevant data.
The dbSNP is another critical resource, cataloging genetic variations across global populations.
Given that autoimmune diseases often have a strong genetic component, dbSNP allows researchers
to identify SNPs that differentiate affected individuals from healthy populations. This facilitates the
discovery of genetic risk factors, supporting personalized medicine initiatives that tailor treatments
based on an individual’s genetic predisposition. However, dbSNP’s reliance on population-wide data
means that ndings must be interpreted in conjunction with clinical validation to establish denitive
associations.
For more detailed genomic exploration, the Ensembl Genome Browser provides gene annotations, transcripts, and regulatory elements, allowing researchers to analyze functional consequences
of genetic variations. Ensembl is particularly useful for autoimmune disease studies, as it enables
the visualization of gene expression proles in immune-related tissues and integrates data from
genome-wide association studies (GWAS). Despite its advantages in data visualization and integration, Ensembl can present a steep learning curve for new users, necessitating prior bioinformatics
knowledge.
The Gene Expression Omnibus (GEO) serves as a critical repository for gene expression data,
encompassing microarray and RNA sequencing (RNA-Seq) datasets. By enabling comparisons of
gene expression between healthy and diseased individuals, GEO has been instrumental in identifying differentially expressed genes that may serve as biomarkers for autoimmune diseases. However,
given the complexity and variability of gene expression studies, careful experimental design and
validation are required to derive meaningful conclusions from GEO data.
Specialized RNA-Seq databases, such as the Sequence Read Archive (SRA) and Expression
Atlas, provide researchers with transcriptomic data crucial for understanding immune pathway
dysregulation in autoimmune diseases. These resources facilitate the identication of alternative
splicing events, long non-coding RNAs (lncRNAs), and regulatory RNA molecules implicated
in disease pathology. While RNA-Seq databases are invaluable for transcriptomic research, data
98 DO I: 10 .1201/9 7810 0 36 85432- 4

99 Bioinformatics Databases
analysis can be computationally intensive, necessitating expertise in bioinformatics tools and statistical interpretation.
The Online Mendelian Inheritance in Man (OMIM) database is an essential resource for
researchers studying genetic disorders, including autoimmune diseases. OMIM provides detailed
information on the genetic basis of diseases, integrating knowledge on gene–disease relationships,
mutations, and phenotypic variations. One of OMIM’s strengths is its curated and regularly updated
content, making it a trusted source for medical genetics. However, OMIM primarily focuses on
monogenic disorders, which may limit its applicability to complex autoimmune diseases that involve
multiple genetic and environmental factors.
The NCBI Database of Genotypes and Phenotypes (dbGaP) is another valuable resource that
provides access to large-scale genetic and clinical data from various studies, including autoimmune
disease research. dbGaP allows researchers to explore the relationship between genotypic variations and phenotypic manifestations in autoimmune conditions. The database’s controlled-access
structure ensures patient privacy while providing researchers with high-quality, well-documented
datasets. However, access to dbGaP requires an application process, which can be a barrier for some
researchers seeking immediate data retrieval.
UniProt is a fundamental proteomics database that offers comprehensive information on protein
sequences and functional annotations. In the context of autoimmune diseases, UniProt provides
insights into protein interactions, structural variations, and post-translational modications (PTMs)
that may inuence disease mechanisms. The database integrates experimental and computationally
predicted data, making it a valuable tool for understanding the molecular basis of immune system
dysfunction. However, because UniProt aggregates data from multiple sources, discrepancies and
uncertainties may arise, requiring careful validation before experimental application.
The Autoimmune Diseases Explorer (ADEX) is a unique resource integrating 82 curated transcriptomics and methylation studies, encompassing over 5609 samples from common autoimmune
diseases. It offers advanced analytical tools, including meta-analysis, differential expression, and
pathway analysis. This level of integration provides a comprehensive approach to studying autoimmune conditions. However, ADEX’s reliance on curated studies means that updates may lag behind
the latest research.
The Interactive Analysis and Atlas for Autoimmune Diseases (IAAA) compiles bulk RNASeq data from 929 samples spanning ten autoimmune diseases, along with single-cell RNA-Seq
data from over 783,203 cells covering six autoimmune diseases. With functionalities such as gene
expression analysis, correlation studies, and cell–cell interaction insights, IAAA provides a highly
detailed landscape of autoimmune disease biology. The challenge with IAAA is its complexity,
requiring advanced computational skills for meaningful data interpretation.
In the realm of proteomics, the Human Protein Atlas (HPA) provides insights into protein expression across different tissues, including immune-related organs. Given that autoimmune diseases
involve aberrant immune responses leading to tissue damage, studying protein expression patterns
through HPA can identify crucial molecular players. The main limitation of HPA is that it focuses
on protein-level data, necessitating integration with genomic and transcriptomic ndings for a holistic understanding of disease mechanisms.
In summary, bioinformatics databases continue to drive breakthroughs in autoimmune disease
research by offering extensive datasets and analytical tools. Resources such as OMIM, dbGaP, and
UniProt provide invaluable insights into the genetic and molecular mechanisms of autoimmune
disorders. While these databases have distinct advantages, their limitations necessitate careful
data integration and interpretation. The continued renement of bioinformatics methodologies will
enhance our understanding of autoimmune disorders, paving the way for more effective diagnostics
and therapeutics.
Table 4.1 provides a comprehensive overview of bioinformatics databases relevant to autoim-
mune disease research, listing their curated institutions, URLs, types of data, and primary applications. Each database serves a specic purpose, ranging from genetic sequence repositories such

100 Bioinformatics of Autoimmune Diseases
TABLE 4.1
Bioinformatics Databases Relevant to Autoimmune Disease Research
Database
Name
GenBank
dbSNP
Ensembl
GEO
SRA
Expression
Atlas
OMIM
dbGaP
UniProt
ADEX
IAAA
H PA
RABC
DisGeNET
ENCODE
KEGG
Reactome
LOVD
UCSC
Browser
IMGT/
HLA
dbMHC
Institution
NCBI
NCBI
EMBL-EBI
NCBI
NCBI
EMBL-EBI
Johns
Hopkins
University
NCBI
UniProt
GENyO
Autoimmune
Atlas
SciLifeLab
Rheumatoid
Arthritis
Cons.
Uni. of
Pompeu
ENCODE
KEGG
Consortium
Reactome
Leiden
University
UCSC
IMGT/HLA
NCBI
URL
https://www.ncbi.nlm.nih.gov/
genbank/
https://www.ncbi.nlm.nih.gov/snp/
https://www.ensembl.org/
https://www.ncbi.nlm.nih.gov/geo/
https://www.ncbi.nlm.nih.gov/sra/
https://www.ebi.ac.uk/gxa/home
https://www.omim.org/
https://www.ncbi.nlm.nih.gov/gap/
https://www.uniprot.org/
https://adex.genyo.es/
https://www.autoimmuneatlas.org/
https://www.proteinatlas.org/
http://www.onethird-lab.com/RABC/
https://www.disgenet.org/
https://www.encodeproject.org/
https://www.kegg.jp/
https://reactome.org/
https://www.lovd.nl/
https://genome.ucsc.edu/
https://www.ebi.ac.uk/ipd/imgt/hla/
https://www.ncbi.nlm.nih.gov/
projects/mhc/
Type of Data
Sequence data
Genetic
variations
Genomic
annotations
Gene expression
RNA-Seq data
Transcriptomics
Genetic disorders
Genotype–
phenotype
Protein sequences
Transcriptomics
RNA-Seq (bulk
and S-cell)
Protein
expression
Multi-omics
Gene–disease
associations
Epigenetic
modications
Pathway analysis
Pathway
interactions
Genetic
variations
Genomic browser
and annotations
HLA sequence
data
HLA sequences
Use
Genomics, genetic sequence
retrieval
Identifying disease-associated
SNPs
Gene structure and function
analysis
Differential gene expression
analysis
RNA sequencing data repository
Functional genomics and
expression
Inherited diseases and gene
mutations
Genotypic variations with
phenotypic traits
Protein function, interactions, and
structure
Transcriptomics and methylation
analysis
Gene expression and immune
interactions
Mapping protein expression
Studying rheumatoid arthritis
Genetic links to autoimmune
diseases
Epigenetic inuences on gene
expression
Pathway mapping for disease
mechanisms
Signaling pathways and disease
interactions
Tracking and curating sequence
variations
Genome-wide data visualization
Studying immune gene
polymorphisms
Genetic testing for autoimmune
conditions

101 Bioinformatics Databases
as GenBank to specialized resources like ADEX and IAAA, which focus on transcriptomics and
immune-related data. By compiling this information, the table highlights the diverse range of bioinformatics resources available to researchers, aiding in the selection of appropriate databases for
specic investigative needs in autoimmune disease studies.
4.2 MAJOR BIOINFORMATICS DATABASES
In the following sections, we explore major bioinformatics databases that store vast amounts of data
on autoimmune diseases. We will discuss the structure of their records, including key data elds, and
examine the application programming interfaces (APIs) available for retrieving this information.
4.2.1 NCBI ENTREZ DATABASE
The NCBI Entrez database is a powerful search and retrieval system that provides access to a vast
collection of biological data, covering various disciplines such as genomics, proteomics, medicine,
and evolutionary biology. It serves as a central hub for accessing a wide range of interconnected
databases maintained by the NCBI. Entrez is designed to facilitate seamless navigation between
different types of biological information, linking data across genetic sequences, protein structures,
scientic literature, and clinical studies. The system enables users to query multiple databases
simultaneously, retrieve relevant records efciently, and explore complex relationships between biological entities.
GenBank is one of the most fundamental databases within Entrez, serving as a comprehensive
public repository of nucleotide sequences, including those relevant to autoimmune diseases. It provides access to a vast collection of DNA and RNA sequences submitted by researchers worldwide,
facilitating the study of genetic variations and mutations associated with autoimmune disorders.
The SRA is another crucial database, containing high-throughput sequencing data, including raw
sequencing reads from studies investigating the genetic basis of autoimmune diseases. SRA data
can be analyzed to identify novel genetic variants, gene expression patterns, and microbiome associations that may contribute to autoimmune conditions.
The Gene database complements GenBank and SRA by offering detailed annotations on genes
implicated in autoimmune diseases, providing information on gene function, expression, and known
variants. The dbSNP database is essential for investigating SNPs linked to autoimmune conditions,
helping researchers understand genetic predispositions. The GEO database offers gene expression
proles from various studies, allowing the analysis of differentially expressed genes in autoimmune
diseases. The ClinVar database is an invaluable resource for exploring genetic variants associated
with clinical conditions, providing expert-curated information on their pathogenicity. The OMIM
database contains detailed descriptions of genetic disorders, including autoimmune diseases, linking phenotypic traits with underlying genetic causes.
Protein-level data relevant to autoimmune diseases can be found in the Protein database, which
includes information on protein sequences, structures, and functions. The Molecular Modeling
Database (MMDB) database provides 3D structural insights that can help in understanding how
mutations affect protein function and immune responses. The UniProtKB/Swiss-Prot database,
accessible through Entrez, contains manually curated protein information that includes disease
associations and functional annotations. The BioSystems database is valuable for exploring biochemical pathways and networks that are dysregulated in autoimmune diseases.
For researchers focusing on medical and clinical aspects, the PubMed database serves as a comprehensive repository of scientic literature, offering access to millions of research articles, case
studies, and reviews on autoimmune diseases. The MedGen database consolidates information on
genetic disorders and their clinical manifestations, making it useful for clinicians and researchers
alike. The dbGaP stores data from GWAS, which help identify genetic factors contributing to autoimmune diseases.

102 Bioinformatics of Autoimmune Diseases
Pathogen-related aspects of autoimmune diseases can be explored using the Taxonomy database,
which provides classication and relationships of microbial species that may trigger or inuence
autoimmune responses. The RefSeq database offers curated sequences of genes, transcripts, and
proteins, serving as a reliable reference for comparative studies.
By integrating data from these diverse sources, Entrez enables researchers to conduct comprehensive investigations into autoimmune diseases, facilitating the discovery of genetic risk factors,
molecular mechanisms, and potential therapeutic targets.
NCBI’s E-Utils and RESTful API provide programmatic access to data from Entrez databases,
enabling users to fetch records in various structured formats depending on the type of data and the
specic database being queried. These formats are designed to support bioinformatics analysis,
visualization, and data integration across different platforms.
To use NCBI E-utilities with Biopython for retrieving information on autoimmune diseases, we
will explore examples of how to fetch data from commonly used Entrez databases. In the following
sections, we will discuss the various data formats that we may encounter when using an API for
data searching and retrieval, and then we will demonstrate how to use the E-Utils API to retrieve
data from some NCBI databases.
4.2.1.1 Common File Formats in NCBI Entrez Databases
4.2.1.1.1 XML Fo r mat
XML (eXtensible Markup Language) is a hierarchical, structured format widely used in bioinformatics to store and exchange complex biological data. It is particularly useful for applications
that require structured metadata, such as linking genes, proteins, and diseases. XML allows for
easy parsing using programming languages like Python and Java, making it a preferred format for
large-scale data retrieval from Entrez databases. Commonly indexed elds in XML include unique
identiers (IDs) (e.g., GeneID, PMID), sequence data, annotations, taxonomic classication, and
references to scientic literature. It is frequently used for data retrieval in GenBank, PubMed, Gene,
and ClinVar, where relationships between entities must be preserved and navigable.
The le N M _ 002116.8.x m l is provided in the supplementary materials as an example. It con-
tains the XML representation of the RefSeq HLA-A transcript (accession: NM_002116.8). Open the
le and examine its content.
The XML le format of the GenBank database follows a structured format designed to store
and annotate biological sequence data. The <Bioseq-set> structure encapsulates sequences,
metadata, and annotations. The sequence data itself is typically stored within <Seq-data> or
<S eq -i n st> , while the coding sequences (CDS) and associated annotations are found within
<S e q-feat>. The taxonomy information is stored in the <BioSource> section, which provides
details about the organism, lineage, and classication.
The sequences in GenBank XML les are often encoded in numbers and letters instead of the
traditional Adenine (A), Thymine (T), Cytosine (C), and Guanine (G) nucleotide representation for
efciency and data standardization. This encoding allows for streamlined processing, compression,
and error detection when handling large-scale genomic data. The use of numerical IDs can also help
with indexing sequences across multiple databases and computational tools.
Key elds in the GenBank XML format include <Org-ref _ taxname>, which species
the species; <Pubdesc _ pub>, which provides literature references; <Seq-inst _ length>,
which indicates the sequence length; < S e q-i ns t _ mo l> , describing the type of molecule (e.g.,
DNA, RNA, protein); and <Seq-inst _ seq-data>, which contains the sequence data itself.
The <Seq-fe at> section holds annotations such as exons, introns, and regulatory elements, while
<GBFeature> elements contain information on specic genes CDS, including their start and stop
positions, strand orientation, and potential translation products. The <CDS> element specically
identies protein-coding regions, linking genomic DNA to its corresponding amino acid sequence.
This structured format ensures interoperability between different bioinformatics tools while
maintaining a rich set of metadata for analysis and research.

For example, assume that you have an “ex a mple.x ml” le, structure of which is as follows:
<?xml version="1.0" encoding="UTF-8"?>
<Bioseq-set>
<Organism>Homo sapiens</Organism>
<CommonName>human</CommonName>
<Chromosome>6</Chromosome>
<Location>6p22.1</Location>
<Sequence>ATGCGTACGTTAGCGT...</Sequence>
<Publications>
<Publication>
<PMID>38946372</PMID>
<Title>Genetic study of a rare Chinese pedigree.</Title>
</Publication>
<Publication>
<PMID>38809622</PMID>
<Title>HLA-A, HLA-B, and HLA-DRB1.</Title>
</Publication>
</Publications>
</Bioseq-set>
To parse “ex ample.x ml” le using Python, save the following codes in a Python le “parse _
x ml.p y” and run it:
103 Bioinformatics Databases
import xml.etree.ElementTree as ET
# Load and parse the XML file
file_path = "example.xml"
tree = ET.parse(file_path)
root = tree.getroot()
# Extract essential data
data = {
"Organism": root.find("Organism").text,
"Common Name": root.find("CommonName").text,
"Chromosome": root.find("Chromosome").text,
"Location": root.find("Location").text,
"Sequence": root.find("Sequence").text
}
# Extract publication data
pubs = {}
for pub in root.findall("Publications/Publication"):
pmid = pub.find("PMID").text
title = pub.find("Title").text
pubs[pmid] = title
for key in data:
print(f"{key}: {data[key]}")
for key in pubs:
print(f"{key} : {pubs[key]}")
The output will be as follows:
Chromosome: 6
Location: 6p22.1
Sequence: ATGCGTACGTTAGCGT...
38946372 : Genetic study of a rare Chinese pedigree.
38809622 : HLA-A, HLA-B, and HLA-DRB1.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
