Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
26 Мб
Скачать
94 Bioinformatics of Autoimmune Diseases
enable the identication of biochemical activities and metabolic networks associated with health or disease.
Statistical analyses are used to assess microbial diversity and identify disease-associated pat­terns. Alpha diversity metrics, including Shannon and Simpson indices, measure the richness and evenness of species within a single sample. Beta diversity metrics, such as Bray-Curtis dissimilarity and UniFrac distances, quantify differences in microbial community structure between groups. These diversity measures are often visualized through principal coordinate analysis (PCoA) or non­metric multidimensional scaling (NMDS).
Emerging computational approaches, including machine learning models such as random forests and deep neural networks, are increasingly used to identify microbial signatures that discriminate between disease and healthy states. These methods enhance prediction accuracy and uncover com­plex microbial interactions and host associations.
3.2.7.5 Functional Insights and Integration with Autoimmune Disease Research
After identifying disease-associated microbial patterns, the next critical step is understanding their functional role in the pathogenesis of autoimmune diseases. Microbial metabolites profoundly inu­ence host immune regulation. Notably, SCFAs (including butyrate, acetate, and propionate) are pro­duced by gut commensals through the fermentation of dietary bers. These SCFAs play essential roles in maintaining intestinal barrier integrity, modulating inammatory responses, and promoting the differentiation and function of regulatory T cells (Tregs), which are critical for immune toler­ance (Arpaia etal., 2012). In contrast, microbial-derived components such as lipopolysaccharides (LPS) can activate toll-like receptor 4 (TLR4)-mediated signaling, leading to systemic inamma­tion and the exacerbation of autoimmune processes.
To gain a deeper understanding of how microbial alterations inuence host biology, researchers are increasingly adopting multi-omics integration strategies. By combining microbiome data with host transcriptomics, proteomics, and metabolomics, it is possible to unravel how microbial-derived signals modulate gene expression, protein activity, and metabolic networks involved in immune regulation and disease progression. These integrative approaches offer a systems-level perspective of host–microbe interactions and are particularly valuable for disentangling complex, multifactorial autoimmune conditions.
The therapeutic potential of microbiome-based interventions is an emerging area of interest in autoimmune disease management. FMT has demonstrated clinical benet in restoring microbial diversity and reducing inammation, particularly in patients with ulcerative colitis. Moreover, the use of probiotics (benecial live microbes) and prebiotics (compounds that promote the growth of benecial microbes) is under investigation for their ability to modulate immune responses and attenuate autoimmune pathology. In parallel, dietary interventions designed to enrich benecial microbial metabolites or suppress inammatory taxa present a promising avenue for non-invasive, long-term immunomodulation.
3.3 SUMMARY
This chapter offers a detailed and structured overview of the bioinformatics frameworks and sequencing-based study designs essential for exploring the molecular and genetic basis of auto­immune diseases. The chapter begins by framing bioinformatics as a crucial discipline that combines computational tools with high-throughput sequencing to identify and interpret genetic variations associated with autoimmune disorders. It explains how technologies such as WGS, WES, and GWAS are deployed to detect disease-associated mutations, including SNPs, indels, SVs, and CNVs. These genetic changes can disrupt immune signaling pathways, affect protein function, and contribute to immune dysregulation. Tools like GATK, A N NO VA R , SpliceAI, and CNVkit are presented as instrumental in the detection, annotation, and functional prediction of these variants.
95 Bioinformatics Design for Autoimmune Research
The chapter continues by integrating epigenomic and metagenomic insights into autoimmune research. Epigenetic alterations, such as DNA methylation and histone modications, are described as non-genetic regulators of gene expression, and their analysis is supported by tools like MACS2, Bismark, and deepTools. Metagenomic approaches are highlighted for their ability to assess the role of the microbiome in autoimmunity, particularly how microbial dysbiosis can inuence immune function. Tools such as Kraken, MetaPhlAn, and HUMAnN are introduced as key resources for microbial proling and functional analysis, showing how host–microbiome interactions may trigger or exacerbate autoimmune responses.
A signicant portion of the chapter is devoted to describing the design and execution of GWAS. The GWAS section outlines how large cohorts of cases and controls are genotyped to identify SNPs associated with disease phenotypes. It elaborates on the necessity of QC procedures, statisti­cal correction for population stratication, and the importance of multiple testing correction using methods such as the Bonferroni adjustment or FDR. Post-GWAS analysis includes ne mapping, functional annotation using databases like ENCODE and GTEx, and expression quantitative trait loci (eQTL) analysis to link variants with gene expression. The chapter emphasizes that GWAS ndings can be further distilled into PRS, offering predictive insights into genetic susceptibility.
The chapter then transitions to WES as a cost-effective strategy to uncover rare and poten­tially pathogenic mutations conned to the protein-coding regions of the genome. WES is detailed step-by-step, from study design and sample selection to sequencing and functional annotation. It describes how tools like BWA, GATK, and DeepVariant are used for alignment and variant calling, while annotation tools like VEP and PolyPhen-2 assess variant pathogenicity. The power of WES lies in its ability to uncover rare de novo or inherited coding variants in immune-related genes, providing mechanistic insights into disease etiology.
WGS is described as a more comprehensive approach that captures all genomic elements, including non-coding regulatory regions. It allows for the detection of structural variants, enhancer mutations, and long-range chromatin interactions that may inuence gene regulation. The chapter outlines the importance of sequencing depth, explaining the formulas and coverage levels required for different research goals. Functional annotation and pathway analysis are emphasized for inter­preting WGS data, particularly in the context of autoimmune diseases, where multiple weak-effect variants interact to shape disease phenotypes.
RNA sequencing (RNA-Seq) is presented as an indispensable method for transcriptome analysis. The chapter explains how RNA-Seq is used to quantify gene expression, detect alternative splicing, and identify non-coding RNAs, all of which are critical to understanding autoimmune pathogen­esis. It discusses sample handling, library preparation, and the impact of transcriptome size on sequencing depth. Normalization techniques, including TPM, RPKM, and TMM, are introduced to account for differences in library size and gene length. Statistical modeling of RNA-Seq data, par­ticularly the use of negative binomial models as implemented in DESeq2 and edgeR, is explained in detail. Limma-voom is introduced as an alternative for log-transformed data with empirical Bayes variance shrinkage. Functional enrichment tools and pathway databases are used to interpret gene expression changes, linking them to immune processes.
The chapter advances into the realm of single-cell RNA sequencing (scRNA-Seq), underscor­ing its transformative impact on resolving cellular heterogeneity in autoimmune diseases. It cov­ers techniques for isolating individual cells, barcoding with UMIs, and preparing cDNA libraries using droplet-based microuidics systems such as 10× Genomics Chromium. With read depth and sequencing quality addressed, the chapter discusses computational methods like UMAP, PCA, and clustering for analyzing single-cell data. It also introduces pseudotime trajectory analysis and ligand–receptor interaction modeling as means of uncovering dynamic immune cell behavior and intercellular communication in disease contexts.
EWAS are presented next, focusing on how epigenetic changes (such as DNA methylation and histone modications) can be proled across the genome-using techniques like WGBS, RRBS, and Innium methylation arrays. The chapter details the statistical modeling of ChIP-Seq data using
96 Bioinformatics of Autoimmune Diseases
negative binomial distributions to assess peak enrichment, differential binding, and downstream functional interpretation through tools like HOMER, GREAT, and ChIPseeker. The integration of chromatin accessibility data from ATAC-Seq is emphasized for identifying regulatory elements involved in immune gene activation.
Finally, the chapter delves into metagenomic and microbiome sequencing, highlighting the importance of microbial ecology in shaping autoimmune risk. It distinguishes between 16S rRNA gene sequencing for taxonomic proling and WGSS for comprehensive microbial and functional analysis. Tools such as QIIME2, Kraken2, and MetaPhlAn are introduced, alongside functional annotation platforms like HUMAnN and KEGG. The chapter outlines methods to evaluate micro­bial diversity and explores how microbiome-derived metabolites, such as SCFAs, inuence immune regulation.
In summary, Chapter 3 provides a richly detailed, methodologically grounded framework for bioinformatics analysis in autoimmune disease research. It unies genetic, transcriptomic, epig­enomic, and microbial dimensions through computational tools and statistical modeling, offering an integrated systems biology approach to unravel the complexities of autoimmunity. Through its comprehensive coverage, the chapter establishes a solid foundation for future studies aimed at understanding pathogenesis, identifying biomarkers, and developing precision medicine strategies for autoimmune disorders.
BIBLIOGRAPHY
Adzhubei, I. A., et al. (2010). A method and server for predicting damaging missense mutations. Nature
Methods, 7(4), 248–249. https://doi.org/10.1038/nmeth0410-248
Anders, S., & Huber, W. (2010). Differential expression analysis for sequence count data. Genome Biology, 11,
R106. https://doi.org/10.1186/gb -2010-11-10-r106
Andrews, S. (2010). FastQC: A quality control tool for high throughput sequence data. http://www.bioinfor-
matics.babraham.ac.uk/projects/fastqc
Arpaia, N., Green, J. A., Moltedo, B., Arvey, A., Hemmers, S., Yuan, S., Treuting, P. M., & Rudensky, A. Y.
(2015). A distinct function of regulatory T cells in tissue protection. Cell, 162(5), 1078–1089. https://doi.
org/10.1016/j.cell.2015.08.021
Bamshad, M. J., et al. (2011). Exome sequencing as a tool for Mendelian disease gene discovery. Nature
Reviews Genetics, 12(11), 745–755. https://doi.org/10.1038/nrg3031
Beghini, F., et al. (2021). Integrating taxonomic, functional, and strain-level proling of diverse microbial
communities with bioBakery 3. eLife, 10, e65088. https://doi.org/10.7554/eLife.65088
Buenrostro, J. D., et al. (2015). ATAC-seq: A method for assaying chromatin accessibility genome-wide.
Current Protocols in Molecular Biology, 109, 21–29. https://doi.org/10.1002/0471142727.mb2129s109
Buniello, A., et al. (2019). The NHGRI-EBI GWAS Catalog of published genome-wide association stud-
ies, targeted arrays and summary statistics. Nucleic Acids Research, 47(D1), D1005–D1012. https://doi.
org/10.1093/nar/gky1120 https://doi.org /10.1093/nar/gky1120
Clark, M. J., et al. (2011). Performance comparison of exome DNA sequencing technologies. Nature
Biotechnology, 29(10), 908–914. https://doi.org/10.1038/nbt.1975
Conesa, A., et al. (2016). A survey of best practices for RNA-seq data analysis. Genome Biology, 17, 13. https://
doi.org/10.1186/s13059-016-0881-8
de Lange, K. M., et al. (2017). Genome-wide association study implicates immune activation of multiple integrin
genes in inammatory bowel disease. Nature Genetics, 49(2), 256–261. https://doi.org/10.1038/ng.3760
Dobin, A., et al. (2013). STAR: Ultrafast universal RNA-seq aligner. Bioinformatics, 29(1), 15–21. https://doi.
org/10.1093/bioinformatics/ bts635
Feil, R., & Fraga, M. F. (2012). Epigenetics and the environment. Nature Reviews Genetics, 13(2), 97–109.
https://doi.org/10.1038/nrg3142
Goodwin, S., et al. (2016). Coming of age: ten years of next-generation sequencing technologies. Nature
Reviews Genetics, 17(6), 333–351. https://doi.org/10.1038/nrg.2016.49
Grosselin, K., et al. (2019). High-throughput single-cell ChIP-seq identies heterogeneity of chromatin states
in breast cancer. Nature Genetics, 51(6), 1060–1066. https://doi.org/10.1038/s41588- 019-042 4-9
97 Bioinformatics Design for Autoimmune Research
Haque, A., et al. (2017). A practical guide to single-cell RNA-sequencing. Genome Medicine, 9, 75. https://
doi.org/10.1186/s13073-017-0467-4
Huang, D. W., et al. (2009). DAVID bioinformatics resources. Nature Protocols, 4(1), 44–57. https://doi.
org/10.1038/nprot.2008.211
Love, M. I., et al. (2014). Moderated estimation of fold change and dispersion. Genome Biology, 15, 550.
ht tps://doi.org/10.1186/s130 59- 014- 0550-8
Manolio, T. A., et al. (2009). Finding the missing heritability of complex diseases. Nature, 461(7265), 747–753.
https://doi.org/10.1038/nature08494
Martin, M. (2011). Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.
Journal, 17(1), 10–12. htt ps://doi.org/10.14806/ej.17.1.200
McKenna, A., et al. (2010). The genome analysis toolkit. Genome Research, 20(9), 1297–1303. https://doi.
org/10.1101/gr.107524.110
M cL a re n, W. , et al. (2016). The ensembl variant effect predictor. Genome Biology, 17, 122. https://doi.
org/10.1186/s13059-016-0974-4
Michels, K. B., et al. (2013). Recommendations for EWAS design and analysis. Nature Methods, 10(10),
949–955. https://doi.org/10.1038/nmeth.2632
Ng, S. B., et al. (2009). Exome sequencing identies the cause of a Mendelian disorder. Nature Genetics, 42(1),
30–35. https://doi.org/10.1038/ng.499
Picelli, S., et al. (2014). Full-length RNA-seq from single cells using Smart-seq2. Nature Protocols, 9(1),
171–181. https://doi.org/10.1038/nprot.2014.006
Price, A. L., et al. (2006). Principal components analysis corrects for stratication. Nature Genetics, 38(8),
904–909. https://doi.org/10.1038/ng1847
Scher, J. U., et al. (2012). Periodontal disease and the oral microbiota in new-onset rheumatoid arthritis.
Arthritis & Rheumatism, 64(10), 3083–3094. https://doi.org/10.1002/art.34539
Wang, E. T., et al. (2009). Alternative isoform regulation in human tissue transcriptomes. Nature, 456(7221),
470–476. https://doi.org/10.1038/nature07509
Zheng, G. X. Y., et al. (2017). Massively parallel digital transcriptional proling of single cells. Nature
Communications, 8, 14049. https://doi.org/10.1038/ncomms14049
Bioinformatics Databases
4
4.1 BIOLOGICAL DATABASES FOR AUTOIMMUNE DISEASE RESEARCH
Bioinformatics databases have become indispensable tools in the study of autoimmune diseases, offering extensive repositories of genetic, proteomic, transcriptomic, epigenetic, and functional data. These resources facilitate the identication of disease-associated variants, analysis of gene expression patterns, exploration of regulatory elements, and investigation of molecular pathways that contribute to immune system dysfunction. The integration of high-throughput sequencing technologies with bioinformatics has signicantly advanced our understanding of autoimmune disorders and the development of precision medicine approaches tailored to individual genetic proles.
One of the foundational databases in autoimmune disease research is the National Center for Biotechnology Information (NCBI) GenBank. This resource provides access to an extensive col­lection of nucleotide sequences from a diverse array of organisms, allowing researchers to retrieve genetic sequences linked to disease susceptibility. GenBank is particularly valuable for compara­tive genomic studies, aiding in the identication of genetic variations associated with autoimmune diseases. The integration of GenBank with other NCBI resources, including database of single nucleotide polymorphism (dbSNP) and ClinVar, enhances its utility by offering insights into single nucleotide polymorphisms (SNPs) and their clinical implications. While GenBank is comprehen­sive, its vast dataset can sometimes be overwhelming, requiring signicant expertise to lter and analyze relevant data.
The dbSNP is another critical resource, cataloging genetic variations across global populations. Given that autoimmune diseases often have a strong genetic component, dbSNP allows researchers to identify SNPs that differentiate affected individuals from healthy populations. This facilitates the discovery of genetic risk factors, supporting personalized medicine initiatives that tailor treatments based on an individual’s genetic predisposition. However, dbSNP’s reliance on population-wide data means that ndings must be interpreted in conjunction with clinical validation to establish denitive associations.
For more detailed genomic exploration, the Ensembl Genome Browser provides gene annota­tions, transcripts, and regulatory elements, allowing researchers to analyze functional consequences of genetic variations. Ensembl is particularly useful for autoimmune disease studies, as it enables the visualization of gene expression proles in immune-related tissues and integrates data from genome-wide association studies (GWAS). Despite its advantages in data visualization and integra­tion, Ensembl can present a steep learning curve for new users, necessitating prior bioinformatics knowledge.
The Gene Expression Omnibus (GEO) serves as a critical repository for gene expression data, encompassing microarray and RNA sequencing (RNA-Seq) datasets. By enabling comparisons of gene expression between healthy and diseased individuals, GEO has been instrumental in identify­ing differentially expressed genes that may serve as biomarkers for autoimmune diseases. However, given the complexity and variability of gene expression studies, careful experimental design and validation are required to derive meaningful conclusions from GEO data.
Specialized RNA-Seq databases, such as the Sequence Read Archive (SRA) and Expression Atlas, provide researchers with transcriptomic data crucial for understanding immune pathway dysregulation in autoimmune diseases. These resources facilitate the identication of alternative splicing events, long non-coding RNAs (lncRNAs), and regulatory RNA molecules implicated in disease pathology. While RNA-Seq databases are invaluable for transcriptomic research, data
98 DO I: 10 .1201/9 7810 0 36 85432- 4
99 Bioinformatics Databases
analysis can be computationally intensive, necessitating expertise in bioinformatics tools and sta­tistical interpretation.
The Online Mendelian Inheritance in Man (OMIM) database is an essential resource for researchers studying genetic disorders, including autoimmune diseases. OMIM provides detailed information on the genetic basis of diseases, integrating knowledge on gene–disease relationships, mutations, and phenotypic variations. One of OMIM’s strengths is its curated and regularly updated content, making it a trusted source for medical genetics. However, OMIM primarily focuses on monogenic disorders, which may limit its applicability to complex autoimmune diseases that involve multiple genetic and environmental factors.
The NCBI Database of Genotypes and Phenotypes (dbGaP) is another valuable resource that provides access to large-scale genetic and clinical data from various studies, including autoimmune disease research. dbGaP allows researchers to explore the relationship between genotypic varia­tions and phenotypic manifestations in autoimmune conditions. The database’s controlled-access structure ensures patient privacy while providing researchers with high-quality, well-documented datasets. However, access to dbGaP requires an application process, which can be a barrier for some researchers seeking immediate data retrieval.
UniProt is a fundamental proteomics database that offers comprehensive information on protein sequences and functional annotations. In the context of autoimmune diseases, UniProt provides insights into protein interactions, structural variations, and post-translational modications (PTMs) that may inuence disease mechanisms. The database integrates experimental and computationally predicted data, making it a valuable tool for understanding the molecular basis of immune system dysfunction. However, because UniProt aggregates data from multiple sources, discrepancies and uncertainties may arise, requiring careful validation before experimental application.
The Autoimmune Diseases Explorer (ADEX) is a unique resource integrating 82 curated tran­scriptomics and methylation studies, encompassing over 5609 samples from common autoimmune diseases. It offers advanced analytical tools, including meta-analysis, differential expression, and pathway analysis. This level of integration provides a comprehensive approach to studying autoim­mune conditions. However, ADEX’s reliance on curated studies means that updates may lag behind the latest research.
The Interactive Analysis and Atlas for Autoimmune Diseases (IAAA) compiles bulk RNA­Seq data from 929 samples spanning ten autoimmune diseases, along with single-cell RNA-Seq data from over 783,203 cells covering six autoimmune diseases. With functionalities such as gene expression analysis, correlation studies, and cell–cell interaction insights, IAAA provides a highly detailed landscape of autoimmune disease biology. The challenge with IAAA is its complexity, requiring advanced computational skills for meaningful data interpretation.
In the realm of proteomics, the Human Protein Atlas (HPA) provides insights into protein expres­sion across different tissues, including immune-related organs. Given that autoimmune diseases involve aberrant immune responses leading to tissue damage, studying protein expression patterns through HPA can identify crucial molecular players. The main limitation of HPA is that it focuses on protein-level data, necessitating integration with genomic and transcriptomic ndings for a holis­tic understanding of disease mechanisms.
In summary, bioinformatics databases continue to drive breakthroughs in autoimmune disease research by offering extensive datasets and analytical tools. Resources such as OMIM, dbGaP, and UniProt provide invaluable insights into the genetic and molecular mechanisms of autoimmune disorders. While these databases have distinct advantages, their limitations necessitate careful data integration and interpretation. The continued renement of bioinformatics methodologies will enhance our understanding of autoimmune disorders, paving the way for more effective diagnostics and therapeutics.
Table 4.1 provides a comprehensive overview of bioinformatics databases relevant to autoim-
mune disease research, listing their curated institutions, URLs, types of data, and primary applica­tions. Each database serves a specic purpose, ranging from genetic sequence repositories such
100 Bioinformatics of Autoimmune Diseases
TABLE 4.1 Bioinformatics Databases Relevant to Autoimmune Disease Research
Database Name
GenBank
dbSNP
Ensembl
GEO
SRA Expression
Atlas
OMIM
dbGaP
UniProt
ADEX
IAAA
H PA
RABC
DisGeNET
ENCODE
KEGG
Reactome
LOVD
UCSC
Browser
IMGT/
HLA
dbMHC
Institution
NCBI
NCBI
EMBL-EBI
NCBI
NCBI EMBL-EBI
Johns
Hopkins University
NCBI
UniProt
GENyO
Autoimmune
Atlas
SciLifeLab
Rheumatoid
Arthritis Cons.
Uni. of
Pompeu
ENCODE
KEGG
Consortium
Reactome
Leiden
University
UCSC
IMGT/HLA
NCBI
URL
https://www.ncbi.nlm.nih.gov/
genbank/
https://www.ncbi.nlm.nih.gov/snp/
https://www.ensembl.org/
https://www.ncbi.nlm.nih.gov/geo/
https://www.ncbi.nlm.nih.gov/sra/ https://www.ebi.ac.uk/gxa/home
https://www.omim.org/
https://www.ncbi.nlm.nih.gov/gap/
https://www.uniprot.org/
https://adex.genyo.es/
https://www.autoimmuneatlas.org/
https://www.proteinatlas.org/
http://www.onethird-lab.com/RABC/
https://www.disgenet.org/
https://www.encodeproject.org/
https://www.kegg.jp/
https://reactome.org/
https://www.lovd.nl/
https://genome.ucsc.edu/
https://www.ebi.ac.uk/ipd/imgt/hla/
https://www.ncbi.nlm.nih.gov/
projects/mhc/
Type of Data
Sequence data
Genetic
variations
Genomic
annotations
Gene expression
RNA-Seq data Transcriptomics
Genetic disorders
Genotype–
phenotype
Protein sequences
Transcriptomics
RNA-Seq (bulk
and S-cell)
Protein
expression
Multi-omics
Gene–disease
associations
Epigenetic
modications
Pathway analysis
Pathway
interactions
Genetic
variations
Genomic browser
and annotations
HLA sequence
data
HLA sequences
Use
Genomics, genetic sequence
retrieval
Identifying disease-associated
SNPs
Gene structure and function
analysis
Differential gene expression
analysis RNA sequencing data repository Functional genomics and
expression Inherited diseases and gene
mutations
Genotypic variations with
phenotypic traits Protein function, interactions, and
structure Transcriptomics and methylation
analysis Gene expression and immune
interactions Mapping protein expression
Studying rheumatoid arthritis
Genetic links to autoimmune
diseases Epigenetic inuences on gene
expression Pathway mapping for disease
mechanisms Signaling pathways and disease
interactions Tracking and curating sequence
variations Genome-wide data visualization
Studying immune gene
polymorphisms Genetic testing for autoimmune
conditions
101 Bioinformatics Databases
as GenBank to specialized resources like ADEX and IAAA, which focus on transcriptomics and immune-related data. By compiling this information, the table highlights the diverse range of bio­informatics resources available to researchers, aiding in the selection of appropriate databases for specic investigative needs in autoimmune disease studies.
4.2 MAJOR BIOINFORMATICS DATABASES
In the following sections, we explore major bioinformatics databases that store vast amounts of data on autoimmune diseases. We will discuss the structure of their records, including key data elds, and examine the application programming interfaces (APIs) available for retrieving this information.
4.2.1 NCBI ENTREZ DATABASE
The NCBI Entrez database is a powerful search and retrieval system that provides access to a vast collection of biological data, covering various disciplines such as genomics, proteomics, medicine, and evolutionary biology. It serves as a central hub for accessing a wide range of interconnected databases maintained by the NCBI. Entrez is designed to facilitate seamless navigation between different types of biological information, linking data across genetic sequences, protein structures, scientic literature, and clinical studies. The system enables users to query multiple databases simultaneously, retrieve relevant records efciently, and explore complex relationships between bio­logical entities.
GenBank is one of the most fundamental databases within Entrez, serving as a comprehensive public repository of nucleotide sequences, including those relevant to autoimmune diseases. It pro­vides access to a vast collection of DNA and RNA sequences submitted by researchers worldwide, facilitating the study of genetic variations and mutations associated with autoimmune disorders. The SRA is another crucial database, containing high-throughput sequencing data, including raw sequencing reads from studies investigating the genetic basis of autoimmune diseases. SRA data can be analyzed to identify novel genetic variants, gene expression patterns, and microbiome asso­ciations that may contribute to autoimmune conditions.
The Gene database complements GenBank and SRA by offering detailed annotations on genes implicated in autoimmune diseases, providing information on gene function, expression, and known variants. The dbSNP database is essential for investigating SNPs linked to autoimmune conditions, helping researchers understand genetic predispositions. The GEO database offers gene expression proles from various studies, allowing the analysis of differentially expressed genes in autoimmune diseases. The ClinVar database is an invaluable resource for exploring genetic variants associated with clinical conditions, providing expert-curated information on their pathogenicity. The OMIM database contains detailed descriptions of genetic disorders, including autoimmune diseases, link­ing phenotypic traits with underlying genetic causes.
Protein-level data relevant to autoimmune diseases can be found in the Protein database, which includes information on protein sequences, structures, and functions. The Molecular Modeling Database (MMDB) database provides 3D structural insights that can help in understanding how mutations affect protein function and immune responses. The UniProtKB/Swiss-Prot database, accessible through Entrez, contains manually curated protein information that includes disease associations and functional annotations. The BioSystems database is valuable for exploring bio­chemical pathways and networks that are dysregulated in autoimmune diseases.
For researchers focusing on medical and clinical aspects, the PubMed database serves as a com­prehensive repository of scientic literature, offering access to millions of research articles, case studies, and reviews on autoimmune diseases. The MedGen database consolidates information on genetic disorders and their clinical manifestations, making it useful for clinicians and researchers alike. The dbGaP stores data from GWAS, which help identify genetic factors contributing to auto­immune diseases.
102 Bioinformatics of Autoimmune Diseases
Pathogen-related aspects of autoimmune diseases can be explored using the Taxonomy database, which provides classication and relationships of microbial species that may trigger or inuence autoimmune responses. The RefSeq database offers curated sequences of genes, transcripts, and proteins, serving as a reliable reference for comparative studies.
By integrating data from these diverse sources, Entrez enables researchers to conduct compre­hensive investigations into autoimmune diseases, facilitating the discovery of genetic risk factors, molecular mechanisms, and potential therapeutic targets.
NCBI’s E-Utils and RESTful API provide programmatic access to data from Entrez databases, enabling users to fetch records in various structured formats depending on the type of data and the specic database being queried. These formats are designed to support bioinformatics analysis, visualization, and data integration across different platforms.
To use NCBI E-utilities with Biopython for retrieving information on autoimmune diseases, we will explore examples of how to fetch data from commonly used Entrez databases. In the following sections, we will discuss the various data formats that we may encounter when using an API for data searching and retrieval, and then we will demonstrate how to use the E-Utils API to retrieve data from some NCBI databases.
4.2.1.1 Common File Formats in NCBI Entrez Databases
4.2.1.1.1 XML Fo r mat
XML (eXtensible Markup Language) is a hierarchical, structured format widely used in bioin­formatics to store and exchange complex biological data. It is particularly useful for applications that require structured metadata, such as linking genes, proteins, and diseases. XML allows for easy parsing using programming languages like Python and Java, making it a preferred format for large-scale data retrieval from Entrez databases. Commonly indexed elds in XML include unique identiers (IDs) (e.g., GeneID, PMID), sequence data, annotations, taxonomic classication, and references to scientic literature. It is frequently used for data retrieval in GenBank, PubMed, Gene, and ClinVar, where relationships between entities must be preserved and navigable.
The le N M _ 002116.8.x m l is provided in the supplementary materials as an example. It con- tains the XML representation of the RefSeq HLA-A transcript (accession: NM_002116.8). Open the le and examine its content.
The XML le format of the GenBank database follows a structured format designed to store and annotate biological sequence data. The <Bioseq-set> structure encapsulates sequences, metadata, and annotations. The sequence data itself is typically stored within <Seq-data> or
<S eq -i n st> , while the coding sequences (CDS) and associated annotations are found within <S e q-feat>. The taxonomy information is stored in the <BioSource> section, which provides
details about the organism, lineage, and classication.
The sequences in GenBank XML les are often encoded in numbers and letters instead of the traditional Adenine (A), Thymine (T), Cytosine (C), and Guanine (G) nucleotide representation for efciency and data standardization. This encoding allows for streamlined processing, compression, and error detection when handling large-scale genomic data. The use of numerical IDs can also help with indexing sequences across multiple databases and computational tools.
Key elds in the GenBank XML format include <Org-ref _ taxname>, which species the species; <Pubdesc _ pub>, which provides literature references; <Seq-inst _ length>, which indicates the sequence length; < S e q-i ns t _ mo l> , describing the type of molecule (e.g., DNA, RNA, protein); and <Seq-inst _ seq-data>, which contains the sequence data itself. The <Seq-fe at> section holds annotations such as exons, introns, and regulatory elements, while <GBFeature> elements contain information on specic genes CDS, including their start and stop positions, strand orientation, and potential translation products. The <CDS> element specically identies protein-coding regions, linking genomic DNA to its corresponding amino acid sequence.
This structured format ensures interoperability between different bioinformatics tools while maintaining a rich set of metadata for analysis and research.
For example, assume that you have an “ex a mple.x ml” le, structure of which is as follows:
<?xml version="1.0" encoding="UTF-8"?> <Bioseq-set>
<Organism>Homo sapiens</Organism> <CommonName>human</CommonName> <Chromosome>6</Chromosome> <Location>6p22.1</Location> <Sequence>ATGCGTACGTTAGCGT...</Sequence> <Publications>
<Publication>
<PMID>38946372</PMID>
<Title>Genetic study of a rare Chinese pedigree.</Title> </Publication> <Publication>
<PMID>38809622</PMID>
<Title>HLA-A, HLA-B, and HLA-DRB1.</Title> </Publication>
</Publications>
</Bioseq-set>
To parse “ex ample.x ml” le using Python, save the following codes in a Python le “parse _
x ml.p y” and run it:
103 Bioinformatics Databases
import xml.etree.ElementTree as ET # Load and parse the XML file file_path = "example.xml" tree = ET.parse(file_path) root = tree.getroot()
# Extract essential data data = {
"Organism": root.find("Organism").text, "Common Name": root.find("CommonName").text, "Chromosome": root.find("Chromosome").text, "Location": root.find("Location").text, "Sequence": root.find("Sequence").text
} # Extract publication data pubs = {} for pub in root.findall("Publications/Publication"):
pmid = pub.find("PMID").text title = pub.find("Title").text pubs[pmid] = title
for key in data:
print(f"{key}: {data[key]}")
for key in pubs:
print(f"{key} : {pubs[key]}")
The output will be as follows:
Chromosome: 6 Location: 6p22.1 Sequence: ATGCGTACGTTAGCGT... 38946372 : Genetic study of a rare Chinese pedigree. 38809622 : HLA-A, HLA-B, and HLA-DRB1.