Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
26 Мб
Скачать
RNA Sequencing
5
5.1 INTRODUCTION
Gene expression is the intricate biological process through which genetic information encoded within DNA is transcribed and translated into functional molecules such as proteins and non­coding RNAs. This process is fundamental to cellular function, development, and adaptability, ensuring that each cell type executes its specialized role. The regulation of gene expression is highly dynamic, involving multiple layers of control, including transcriptional, post-transcriptional, trans­lational, and post-translational mechanisms. These regulatory networks allow cells to respond to intrinsic and extrinsic stimuli, including environmental changes, developmental cues, and immune challenges.
The immune system is particularly dependent on precise gene expression control, as it must rapidly adapt to pathogens while maintaining self-tolerance. A breakdown in the regulation of gene expression can lead to immune dysfunction, contributing to conditions such as autoimmune dis­eases. Autoimmune disorders arise when immune cells erroneously target the body’s own tissues, causing chronic inammation and tissue damage. The pathogenesis of these diseases is inuenced by a complex interplay between genetic predisposition and environmental factors, with gene expres­sion playing a crucial role in disease initiation, progression, and severity. Dysregulated gene expres­sion can result in the aberrant activation of immune cells, an imbalance between pro-inammatory and anti-inammatory cytokines, and a failure to establish immune tolerance.
The advent of high-throughput sequencing technologies, particularly RNA sequencing (RNA­Seq), has revolutionized the study of gene expression in autoimmune diseases. These approaches enable researchers to systematically prole transcriptomic changes in patients, uncovering gene expression signatures associated with different disease states. For example, systemic lupus ery­thematosus (SLE) is characterized by an elevated expression of interferon-stimulated genes, high­lighting the role of type I interferons in disease pathogenesis. Similarly, impaired expression of genes involved in regulatory T-cell (Treg) function has been implicated in the loss of immune tolerance, a hallmark of autoimmune diseases such as type 1 diabetes and multiple sclerosis (MS). By analyzing these expression patterns, researchers gain insights into the molecular mecha­nisms driving autoimmunity and identify potential biomarkers for early diagnosis and disease stratication.
Gene expression analysis also facilitates the development of targeted therapies, allowing for the design of interventions that modulate specic disease-related pathways. Biologic drugs such as monoclonal antibodies targeting tumor necrosis factor-alpha (TNF-α) or interleukin-6 (IL-6) have been developed based on transcriptomic studies and have transformed the management of diseases like rheumatoid arthritis and inammatory bowel disease. Moreover, the emergence of personalized medicine underscores the importance of tailoring treatments based on individual gene expression proles. Since patients exhibit varying transcriptomic responses to immunosuppressive drugs, pharmacogenomic studies help optimize therapeutic strategies, reducing adverse effects and enhancing efcacy.
Beyond genetic inuences, epigenetic modications signicantly impact gene expression in autoimmune diseases. DNA methylation, histone modications, and non-coding RNAs serve as regulatory elements that ne-tune gene activity in response to environmental stimuli. External factors such as viral infections, diet, and stress can induce epigenetic changes that alter immune function and contribute to disease susceptibility. For instance, DNA hypomethylation of immune­related genes has been observed in SLE patients, leading to the hyperactivation of immune cells and
164 DO I: 10 .1201/9 7810 0 36 85432-5
165 RNA Sequencing
increased autoantibody production. These ndings emphasize the need to integrate epigenetic data with transcriptomic analyses to obtain a comprehensive understanding of autoimmunity.
Experimental studies of gene expression in autoimmune diseases rely on meticulously designed RNA-Seq workows to ensure robust and biologically meaningful results. A well-structured study begins with the clear denition of research objectives, such as comparing gene expression pro­les between diseased and healthy individuals or assessing transcriptional changes in response to treatment. Patient cohort selection is critical, requiring stringent inclusion and exclusion criteria to minimize confounding variables. Factors such as disease severity, medication history, genetic background, and demographic characteristics must be carefully controlled to enhance the reliability of ndings.
Once the study design is established, biological samples are collected from relevant tissues, including peripheral blood mononuclear cells (PBMCs), synovial tissue in rheumatoid arthritis, or intestinal biopsies in Crohn’s disease. The choice of sample type is crucial, as autoimmune dis­eases exhibit cell-type-specic gene expression changes. To preserve RNA integrity, samples are processed under standardized conditions, and RNA extraction protocols are optimized to mini­mize degradation. The quality of extracted RNA is assessed using bioanalyzers, ensuring high RNA integrity numbers (RIN) before sequencing. Depending on the study’s goals, either poly(A) enrichment or ribosomal RNA depletion is performed to focus on coding or total RNA populations, respe ctive ly.
Library preparation involves converting RNA into cDNA, fragmenting it, and ligating sequenc­ing adapters. To reduce amplication bias, unique molecular identiers (UMIs) may be incorpo­rated. Sequencing is then carried out on high-throughput platforms such as Illumina NovaSeq, with sequencing depth tailored to the experimental objectives. For differential gene expression analysis, read depths of 30–50 million reads per sample are typically sufcient, whereas studies investigating alternative splicing or rare transcripts may require deeper sequencing.
Post-sequencing, bioinformatics pipelines process the raw sequencing data. Reads are rst sub­jected to quality control using FastQC to assess sequencing accuracy, remove adapter sequences, and lter out low-quality bases. High-quality reads are then aligned to the reference genome using alignment tools like STAR or HISAT2, ensuring accurate mapping. Gene expression levels are quantied using featureCounts or Salmon, and normalization techniques such as transcripts per million (TPM) or variance-stabilizing transformation are applied to account for sequencing biases.
Differential expression analysis is conducted using statistical frameworks such as DESeq2 or edgeR, identifying genes that exhibit signicant expression changes between disease and control groups. Functional enrichment analyses follow, employing tools like gene set enrichment analy­sis (GSEA) or DAVID to determine biological pathways associated with differentially expressed genes. By integrating transcriptomic data with proteomic and epigenomic datasets, researchers can construct comprehensive disease models that capture the interplay between gene regulation and immune dysfunction.
Validation of key ndings is a critical step, ensuring the reproducibility of results. Quantitative real-time PCR (qRT-PCR) is commonly used to conrm differentially expressed genes, while func­tional assays, such as cytokine secretion measurements or CRISPR-based gene perturbation, elu­cidate the biological relevance of identied genes. These experimental approaches bridge the gap between gene expression studies and mechanistic insights, paving the way for novel therapeutic strategies.
RNA-Seq data visualization techniques play an essential role in interpreting complex transcrip­tomic datasets. Heatmaps are used to display hierarchical clustering of gene expression patterns across samples, while volcano plots highlight signicantly dysregulated genes. Principal component analysis (PCA) enables the identication of sample clustering and outliers, providing a global over­view of gene expression variations. Such visualizations facilitate hypothesis generation and enhance the interpretability of large-scale transcriptomic studies.
166 Bioinformatics of Autoimmune Diseases
The integration of multi-omics approaches with gene expression analysis represents the future of autoimmune disease research. By combining RNA-Seq with single-cell sequencing, chromatin accessibility proling (ATAC-Seq ), and proteomics, researchers can gain unprecedented insights into cellular heterogeneity and molecular drivers of autoimmunity. Ultimately, these advances will contribute to more precise diagnostic tools, improved therapeutic interventions, and personalized treatment strategies, revolutionizing the management of autoimmune disorders.
RNA-Seq has revolutionized the study of autoimmune diseases by providing a powerful tool for uncovering molecular mechanisms, identifying diagnostic biomarkers, and guiding personalized treatments. Unlike traditional methods, RNA-Seq enables comprehensive analysis of gene expres­sion across different immune cell populations, tissues, and disease states, leading to breakthroughs in understanding the underlying pathology of autoimmune disorders.
One of the most signicant successes of RNA-Seq in autoimmune disease research has been the identication of disease-specic gene expression signatures. For instance, in SLE, a chronic auto­immune disease affecting multiple organs, RNA-Seq has revealed distinct interferon-related gene expression patterns in patients’ blood cells. These ndings have not only conrmed the central role of type I interferons in SLE but have also led to the development of new targeted therapies, such as anifrolumab, which specically blocks the type I interferon receptor and has shown efcacy in clinical trials.
Similarly, in rheumatoid arthritis (RA), RNA-Seq has been instrumental in identifying inam­matory pathways that drive disease progression. By analyzing synovial tissue from RA patients, researchers have uncovered key differences in gene expression between those who respond to stan­dard treatments and those who do not. This has paved the way for precision medicine approaches, where RNA-Seq data can predict which patients will benet from biologic therapies such as TNF inhibitors or JAK inhibitors. Such stratication not only improves treatment efcacy but also mini­mizes unnecessary exposure to ineffective drugs and their potential side effects.
Beyond diagnosis and treatment selection, RNA-Seq has contributed to the discovery of novel therapeutic targets. In MS, a disease characterized by immune-mediated damage to the nervous system, RNA-Seq has helped identify previously unrecognized immune cell subtypes involved in disease progression. This has led to the exploration of new drug candidates that modulate these cell populations, potentially offering more effective treatment options for MS patients who do not respond to conventional therapies.
RNA-Seq has also facilitated the understanding of complex autoimmune conditions such as inammatory bowel disease (IBD), which includes Crohn’s disease and ulcerative colitis. By prol­ing intestinal tissue samples, researchers have identied molecular subtypes of IBD, some of which respond differently to available treatments. This knowledge is being used to develop personalized therapeutic strategies, ensuring that patients receive the most appropriate interventions based on their specic gene expression proles.
Perhaps one of the most exciting applications of RNA-Seq in autoimmune disease research is the ability to track disease progression and treatment response over time. By sequencing RNA from patient samples at multiple time points, scientists can monitor how immune activity changes in response to therapy, providing insights into disease remission and potential relapses. This real-time molecular monitoring has the potential to transform disease management, allowing clinicians to adjust treatments proactively rather than reactively.
The success of RNA-Seq in autoimmune disease research highlights its potential as a cornerstone of future diagnostic and therapeutic approaches. As sequencing technologies continue to advance, costs decrease, and computational methods improve, RNA-Seq is poised to become an integral part of routine clinical practice. Its ability to provide detailed molecular insights into immune dys­function not only enhances our understanding of autoimmune diseases but also brings us closer to achieving truly personalized medicine for patients suffering from these complex conditions.
In the following sections, we demonstrate how to retrieve both raw data and metadata from the NCBI SRA database. The raw data typically consists of sequencing reads in FASTQ format, while
167 RNA Sequencing
the metadata includes study design information, which is essential for meaningful RNA-Seq data analysis. If you are working with your own data, you can substitute your own study design and make minimal adjustments to the provided code to suit your specic objectives. To keep the focus on the workow rather than the biological interpretation, gene names in the results, such as tables and plots, are anonymized using labels like gene1, gene2, and so on. All the scripts used in this chapter are available in the GitHub repository referenced in the introduction of this book.
5.2 EXTRACTING BIOPROJECT STUDY DESIGN FROM SRA DATABASE
We discussed the principles of study design in Chapter 3. A well-dened study design is essential prior to performing RNA-Seq experiments, as it determines the structure of the experiment, the number and nature of sample groups, and the specic biological questions to be addressed. By the time we enter the RNA-Seq data analysis phase, critical decisions (such as sample groupings, experimental conditions, and the hypotheses to be tested) should already be in place.
Researchers commonly deposit raw RNA-Seq data into public repositories such as the NCBI Sequence Read Archive (SRA), along with associated metadata that describes the study, sample characteristics, and sequencing details. This standardized archiving ensures data transparency, reproducibility, and reusability by the broader scientic community.
In this chapter, we demonstrate how to access and retrieve RNA-Seq raw data from the NCBI SRA database. Specically, we focus on the NCBI BioProject PRJEB52174, which contains paired-end FASTQ les representing 184 RNA-Seq samples from human subjects diagnosed with rheumatoid arthritis (RA). To simplify our analysis and reduce computational demands, we will work with a representative subset of 20 samples (National Center for Biotechnology Information,
2022).
The associated study metadata (such as SRA run accession numbers, experimental conditions, patient demographics, and disease states) can be retrieved either manually through the SRA Run Selector interface on the NCBI website or programmatically using scripting tools such as Bash or Python. For reproducibility and automation, we will use Python to download and parse the meta­data, after which we will select the 20 samples to be included in our downstream analyses.
The Python script “sr a _ ru n s el ec to r.p y ” is designed to extract detailed study design metadata from the SRA (Sequence Read Archive) database at NCBI, specically for a given BioProject, and save that metadata into a CSV le. The goal is to capture not just the basic run and sample information but also the custom annotations found in the BioSample, such as tissue type, time point, treatment conditions, or developmental stage. These annotations are commonly used in the SRA Run Selector interface on NCBI’s website and are critical for understanding the experi­mental context of the sequencing data.
To begin with, the script uses the BioPython Entrez module to search the SRA database for all entries associated with the specied BioProject ID. The esearch function retrieves a list of unique IDs corresponding to SRA records. These IDs are then individually processed in a loop, where each one is fetched in XML format using efetch. However, instead of using Entrez.
re ad() to parse the XML, as was done initially, the script now uses Python’s built-in xm l.etr ee. ElementTree module. This change is necessary because many SRA XML records do not include
a formal Document Type Denition (DTD) or XML Schema, which makes them incompatible with BioPython’s strict parser.
For each XML record retrieved, the script navigates the tree structure to locate the relevant sections: EXPERIMENT_PACKAGE, RUN, and SAMPLE. It extracts the run accession number, the BioSample ID, and the scientic name of the sample. The most important section is SAMPLE_ ATTRIBUTES, where the submitter-dened metadata resides. This part of the XML contains repeated elements with a TAG and VALUE, which represent various custom attributes describing the sample and study design. The script iterates through these attributes and dynamically adds them as new columns in a metadata dictionary. Each dictionary corresponds to a row in the nal dataset.
168 Bioinformatics of Autoimmune Diseases
FIGURE 5.1 Partial view of the metadata contained in the CSV le.
All rows are collected into a list and converted into a pandas DataFrame, which is then saved as a CSV le (Figure 5.1). The output le contains standard identiers such as run and sample accession numbers, along with any additional study-specic elds submitted by the researchers. By using ElementTree, the script ensures compatibility with the XML format of SRA records while capturing rich contextual metadata that can be used for downstream analysis or curation. It also includes error handling and respects NCBI’s rate-limiting guidelines by inserting a delay between fetch requests. This approach provides a reliable and scalable way to programmatically extract structured metadata from the SRA for any BioProject of interest.
5.3 STUDY VARIABLES
After downloading the study metadata, we selected a set of variables, listed in Table 5.1, that will be used in the RNA-Seq data analysis. The clinical variables and their relevance to the study are described in detail below.
TABLE 5.1 Selected Variables of the Rheumatoid Arthritis Study
runID Age Anti-CCP ADAS cellType CRP ESR Ethnic Pathotype RF Sex SJC TJC
ERR9539313 40 Positive 75 Bpoor 39 67 African Lymphoid Positive Female 5 11 ERR9539319 56 Negative 98 Brich 11 30 African Lymphoid Positive Female 2 8 ERR9539328 46 Positive 95 Brich 29 26 African Lymphoid Negative Female 11 15 ERR9539343 45 Positive 4 Brich 14 19 African Lymphoid Positive Male 6 2 ERR9539358 24 Positive 71 Brich 36 52 African Lymphoid Positive Female 4 16 ERR9539216 67 Negative 17 Bpoor 0 5 African Fibroid Negative Female 10 12 ERR9539393 38 Positive 40 Brich 9 5 Asian Lymphoid Positive Male 7 5 ERR9539217 41 Positive 82 Bpoor 16 42 Asian Fibroid Positive Female 18 19 ERR9539241 39 Negative 21 Brich 6 52 Asian Lymphoid Positive Male 0 10 ERR9539243 43 Positive 89 Brich 89 131 Asian Lymphoid Positive Female 3 5 ERR9539246 59 Positive 85 Bpoor 0 35 Asian Fibroid Positive Female 3 11 ERR9539251 42 Negative 70 Bpoor 13 45 Asian Lymphoid Negative Female 11 28 ERR9539213 31 Positive 93 Bpoor 6 12 Caucasian Fibroid Negative Female 14 18 ERR9539320 62 Positive 55 Bpoor 0 20 Caucasian Fibroid Positive Female 5 6 ERR9539331 60 Negative 41 Brich 4 12 Caucasian Lymphoid Positive Female 3 2 ERR9539353 69 Negative 56 Brich 1 5 Caucasian Lymphoid Negative Male 1 2 ERR9539364 70 Positive 49 Bpoor 6 25 Caucasian Fibroid Positive Female 16 13 ERR9539375 59 Positive 57 Bpoor 0 8 Caucasian Myeloid Negative Female 7 13
i. Anti-cyclic citrullinated peptide antibody measurement
Anti-cyclic citrullinated peptide (anti-CCP) antibody measurement is a blood test used primarily to help diagnose rheumatoid arthritis (RA). This test detects the presence of spe­cic antibodies that the immune system produces against peptides or proteins that contain the amino acid citrulline (a modied form of the amino acid arginine). In people with RA, the immune system mistakenly targets these citrullinated proteins as if they were foreign invaders.
The process of citrullination happens normally in the body, but in autoimmune condi­tions like RA, it becomes excessive or dysregulated, leading to an immune response. The anti-CCP test looks for antibodies directed against cyclic citrullinated peptides, which are synthetic versions of citrullinated proteins. A high level of anti-CCP antibodies in the blood is strongly associated with RA and can even appear in the blood before symptoms start, making the test useful for early diagnosis.
ii. B Cells
In RA, B cells play a signicant and multifaceted role in driving the autoimmune response that leads to chronic joint inammation and tissue damage. B cells are a type of white blood cell that are part of the adaptive immune system, normally responsible for producing antibodies to ght infections. However, in autoimmune diseases like RA, their functions become dysregulated, contributing to the development and persistence of the disease.
One of the hallmark contributions of B cells in RA is their production of autoantibod­ies, such as rheumatoid factor (RF) and anti-CCP antibodies. These autoantibodies target the body’s own proteins, forming immune complexes that deposit in the joints and trigger inammatory cascades. The presence of these antibodies is strongly associated with dis­ease severity and progression, and anti-CCP antibodies in particular are highly specic to RA. B cells also function as antigen-presenting cells, meaning they process and present antigens to T cells, which further amplies the autoimmune response.
Beyond antibody production and antigen presentation, B cells secrete pro-inammatory cytokines such as IL-6, which contribute to the recruitment and activation of other immune cells in the synovial membrane. This leads to the formation of pannus tissue, an abnormal thickening of the synovial lining that invades cartilage and bone. B cells are found in increased numbers within the inamed synovial tissue of RA patients, often organized into structures that resemble germinal centers (GCs), similar to those seen in lymph nodes.
The central role of B cells in RA pathogenesis has made them a therapeutic target. Treatments such as rituximab, a monoclonal antibody that depletes B cells by targeting CD20, have been effective in reducing disease activity in RA patients, particularly those who are seropositive for RF or anti-CCP antibodies. This underscores the importance of B cells not just in initiating the autoimmune process but also in sustaining the chronic inammation characteristic of rheumatoid arthritis.
In the context of RA, the terms B-cell rich, B-cell poor, and GC refer to distinct pat­terns of immune cell organization and activity observed in the synovial tissue, which lines the joints. These patterns are identied through histological and immunological analysis, often during research studies or biopsy-based evaluations, and they help characterize the immune microenvironment of the affected joints. Understanding these categories can pro­vide insights into the disease mechanism and may even guide therapeutic strategies.
B-cell-rich synovitis describes synovial tissue that contains a high number of B lym­phocytes. In this pattern, B cells are often found clustered together, sometimes forming organized structures resembling GCs, which are areas where B cells proliferate, mutate their antibody genes, and undergo selection. These structures are similar to those found in lymphoid tissues like lymph nodes and suggest a highly active and organized immune response within the joint. B-cell-rich synovitis is typically associated with the presence of autoantibodies such as rheumatoid factor and anti-CCP antibodies, as well as more severe
169 RNA Sequencing
170 Bioinformatics of Autoimmune Diseases
inammation. Patients with this pattern may respond well to therapies that target B cells, such as rituximab.
iii. C-reactive protein
C-reactive protein (CRP) is a key biomarker of inammation that is widely used to assess disease activity and monitor response to treatment. CRP is a protein produced by the liver in response to inammation, particularly when stimulated by cytokines like IL-6. Its lev­els in the blood rise rapidly during episodes of acute inammation and fall just as quickly when the inammation subsides, making it a sensitive and dynamic indicator.
In RA, which is a chronic inammatory autoimmune disease, elevated CRP levels often reect active joint inammation and tissue damage. Clinicians use CRP values alongside other measures, such as joint counts, patient symptoms, and imaging, to determine how active the disease is. A high CRP suggests ongoing inammation and possibly uncon­trolled disease, while a low or normal CRP can indicate effective disease control or remis­sion. However, CRP is a general marker of inammation and is not specic to RA; it can be elevated in infections, injuries, or other autoimmune conditions as well.
CRP is also a component of several composite disease activity scores used in RA, such as the Disease Activity Score 28-CRP (DAS28-CRP) and Simplied Disease Activity Index (SDAI). In these indices, CRP contributes to the overall evaluation of disease sever­ity. Tracking CRP over time helps rheumatologists adjust medications, evaluate ares, and assess whether a patient is responding well to therapy. Overall, CRP is a practical, acces­sible, and clinically valuable tool in the diagnosis and management of rheumatoid arthritis.
iv. Erythrocyte sedimentation rate
Erythrocyte sedimentation rate (ESR) is a common blood test used to measure the level of inammation in the body. It reects how quickly red blood cells (erythrocytes) settle to the bottom of a test tube over a specied period, typically 1 hour. In healthy individuals, red blood cells settle slowly, but in the presence of inammation, certain proteins such as brinogen cause the cells to clump together and settle more rapidly. A higher ESR indi­cates a higher level of systemic inammation.
In RA, which is a chronic autoimmune condition characterized by persistent joint inammation, ESR is used as a marker of disease activity. It helps clinicians assess how active the disease is, monitor the patient’s response to therapy, and detect disease ares. Elevated ESR is often found in patients with active RA and may correlate with symptoms such as joint pain, swelling, stiffness, and fatigue. However, like CRP, ESR is a nonspecic test, meaning it can also be elevated in other conditions like infections, other autoimmune diseases, or even some cancers.
ESR is included in several composite scores used to measure disease activity in RA, such as the Disease Activity Score 28 (DAS28-ESR). In this context, ESR is combined with joint assessments and patient-reported symptoms to generate a numerical score that helps guide treatment decisions. Although it is a useful tool, ESR can be inuenced by fac­tors unrelated to RA, such as age, sex, and anemia. Therefore, it is typically interpreted in conjunction with other clinical ndings and tests to provide a more accurate picture of the patient’s disease status.
v. Arthritis disease activity score measurement
The Arthritis Disease Activity Score (ADAS) is a quantitative tool used by clinicians to measure the severity of disease activity in patients with RA. It combines clinical assess­ments and, in some versions, laboratory results to produce a single numerical score that reects how active the disease is at a given time. The score helps physicians determine the level of inammation, monitors how the disease is progressing, and guides treatment decisions.
The most commonly used version is the DAS28, which stands for disease activity score using 28 joints. In this method, 28 specic joints are examined for tenderness and swelling,
usually including the shoulders, elbows, wrists, knees, and joints of the hands. The score
)
=
)
+
)
=
++
+ 0.96
)
also incorporates the patient’s general health assessment, which is typically provided by the patient on a scale of 0–100 (or sometimes 0–10), and either the ESR or CRP, which are both blood tests that measure inammation. The nal DAS28 score is calculated using a mathematical formula that combines these elements. The formula of DAS28 using ESR (Prevoo et al., 1995) is the following:
171 RNA Sequencing
AS28(ESR
where TJC28 = tender joint count (28 joints), SJC28 = swollen joint count (28 joints), ESR= erythrocyte sedimentation rate (mm/hr), GH = general health assessment (usually patient global assessment on a 0–100 mm visual analog scale).
If CRP is used instead of ESR, the formula slightly changes (Fransen & van Riel, 2005):
AS28(CRP
The resulting score usually falls between 0 and 10. Scores are interpreted in ranges: a score above 5.1 indicates high disease activity, between 3.2 and 5.1 indicates moderate activity, between 2.6 and 3.2 indicates low activity, and a score below 2.6 is considered remission. A decrease in DAS28 over time suggests improvement in disease control, while an increase suggests a worsening of RA.
Overall, DAS28 is a standardized and validated tool that allows for consistent moni­toring of rheumatoid arthritis activity across different visits, clinics, and clinical trials. It helps ensure that treatment decisions are based on objective measurements rather than subjective impressions alone.
vi. Pathotype (Lymphoid, Myeloid, Fibroid)
The term pathotype refers to distinct patterns of immune cell inltration and tissue orga­nization found in the synovial membrane, which lines the joints. These patterns reect dif­ferent underlying mechanisms of inammation and can be classied into three main types: lymphoid, myeloid, and broid. Understanding these pathotypes provides insight into the heterogeneity of RA and may eventually guide more personalized treatment strategies.
The lymphoid pathotype is characterized by the presence of organized clusters of B cells and T cells in the synovial tissue. These immune cells often form structures resem­bling GCs, which are typically found in lymphoid organs like lymph nodes. This organiza­tion suggests a highly active, antigen-driven immune response occurring directly within the joint. Patients with this pathotype often have autoantibodies such as rheumatoid factor (RF) or anti-CCP, and the lymphoid pattern is associated with more severe inammation and joint damage. These patients may respond well to therapies targeting B cells or specic cytokines like IL-6.
The myeloid pathotype is dominated by macrophages and neutrophils, which are cells of the innate immune system. This type is typically associated with strong expression of pro-inammatory cytokines such as tumor necrosis factor (TNF) and IL-1. The myeloid pattern tends to reect a different kind of immune activity that is more innate and less dependent on the adaptive immune system. Patients with a myeloid pathotype may respond particularly well to TNF inhibitors, which are a common class of biologic drugs used in RA treatment.
The broid pathotype is distinguished by a lack of signicant immune cell inltration. Instead, the synovial tissue shows broblast proliferation and tissue remodeling without the high levels of lymphocytes or macrophages seen in the other types. This suggests a lower level of active inammation, and patients with this pattern may have less joint swell­ing and fewer systemic symptoms. However, the broid pattern may still be associated with
0.56 × TJC28 + 0.28 × SJC28 + 0.70 × ln (ESR
0.56 ×TJC28 +0.28 ×SJC28 +0.36 ×ln (CRP
0.014 × GH
0.014 ×GH
1
172 Bioinformatics of Autoimmune Diseases
joint stiffness and structural changes due to chronic broblast activity and extracellular matrix deposition.
These pathotypes reect different immune processes operating in RA and highlight the complexity of the disease. Identifying which pathotype a patient has (through synovial biopsy and molecular analysis) could help tailor therapies more precisely in the future, moving beyond the one-size-ts-all approach currently used in clinical practice.
vii. Rheumatoid factor measurement
Rheumatoid factor (RF) measurement is a blood test used to detect the presence of auto­antibodies (specically, antibodies directed against the Fc portion of immunoglobulin G (IgG)) which are commonly found in people with RA. RF was one of the earliest biomark­ers associated with RA and remains a widely used tool in diagnosing and classifying the disease.
In RA, the immune system becomes dysregulated and begins to produce RF, which binds to the body’s own antibodies, forming immune complexes that contribute to inam­mation in the joints. The presence of RF in the blood suggests that the body is mounting an abnormal immune response, characteristic of autoimmune conditions. RF is detected using techniques like nephelometry, ELISA (enzyme-linked immunosorbent assay), or latex agglutination, and the results are usually reported in international units per milliliter (IU/mL). A value above the laboratory’s reference range (commonly around 20–30 IU/mL) is considered positive.
However, RF is not specic to RA. It can also be elevated in other autoimmune diseases (like Sjögren’s syndrome or lupus), chronic infections (such as hepatitis C or tuberculo­sis), and even in some healthy older adults. For this reason, RF is typically interpreted in combination with clinical ndings and other tests, especially the anti-cyclic citrullinated peptide (anti-CCP) antibody test, which is more specic for RA.
A positive RF test, especially at high levels, supports a diagnosis of rheumatoid arthritis and is often associated with more severe or erosive disease. It may also be used as part of the classication criteria for RA and to help predict disease progression. Despite its limita­tions in specicity, RF remains an important marker in the diagnostic workup and ongoing assessment of RA.
viii. Swollen Joint Count
Swollen joint count (SJC) is a clinical measure used to assess joint inammation and dis­ease activity. It refers to the number of joints that appear swollen due to synovitis, which is the inammation of the synovial membrane lining the joints. Swelling in RA is typically a sign of active disease and is associated with pain, stiffness, and long-term joint damage if not treated.
SJC is typically assessed by a trained clinician who performs a systematic physical examination of the patient’s joints. The most commonly used format is the SJC28, which involves evaluating 28 specic joints, including the shoulders, elbows, wrists, metacar­pophalangeal (MCP) joints, proximal interphalangeal (PIP) joints, and knees. Each joint is examined for the presence of visible or palpable swelling, which is indicative of uid accumulation or synovial thickening.
The total count of swollen joints is used as part of several composite disease activity scores, such as the DAS28, SDAI (Simplied Disease Activity Index), and CDAI (Clinical Disease Activity Index). A higher SJC suggests more active inammation and typically correlates with greater disease severity. Conversely, a low or zero SJC may indicate remis­sion or effective control of RA with treatment.
SJC is an essential tool in both clinical practice and research. It provides an objective measure that, when combined with other assessments like tender joint count (TJC), patient­reported symptoms, and inammatory markers like CRP or ESR, helps guide treatment decisions and track response over time.
ix. Tender Joint Count
Tender joint count (TJC) is a clinical measure used to assess how many joints are painful or sensitive to pressure during a physical examination. It reects the subjective experience of joint pain from the patient and is a key indicator of disease activity, particularly how inammation is affecting joint function and causing discomfort.
TJC is usually assessed by a healthcare provider who gently presses on specic joints to check for tenderness. The most common format is the TJC28, which evaluates 28 joints, including the shoulders, elbows, wrists, knees, and the small joints of the hands (MCP and PIP joints). If pressing on a joint causes pain or discomfort, it is recorded as tender. The total number of tender joints is then summed to produce the TJC score.
This measure is often used alongside the SJC and other indicators such as CRP (C-reactive protein) or ESR to calculate composite scores like the DAS28, SDAI, and CDAI, which provide a more comprehensive view of disease activity in RA. A high TJC indicates more joint pain, which can reect active inammation, but it may also be inu­enced by other factors such as fatigue, bromyalgia, or individual pain sensitivity.
TJC is important for monitoring disease progression and treatment response. While it is more subjective than SJC—since it depends on the patient’s pain perception—it still provides valuable information, especially when tracked over time. In clinical settings, a decrease in TJC suggests improvement and effective disease control, while an increase may signal a are or inadequate response to therapy.
5.4 DOWNLOADING THE RNA-SEQ RAW DATA
173 RNA Sequencing
Save the run identiers “runID” from Table 5.1 in a text le named “ids.txt”. Then run the Python script “download _ fastq.py”, which is designed to automate the process of downloading FASTQ les from the Sequence Read Archive (SRA) using run identiers provided in the text le. It begins by dening the name of the input le and the directory where the downloaded les will be stored, in this case, a folder called “fastqs”. The script checks whether this directory already exists. If it does, it prints a warning message to notify the user that the directory is present and then exits the program without performing any downloads. This safety check helps prevent accidental overwriting of existing data or mixing new les with previously downloaded ones.
If the fastqs directory does not exist, the script proceeds to create it, ensuring that there is a dedicated location to store the downloaded data. It then attempts to open and read the id s.t xt le, which should contain a list of SRA run accession IDs, one per line. If the le is missing or cannot be opened, the script prints an error message and exits, avoiding any further execution that would otherwise rely on non-existent input data.
Once the run IDs are successfully read, the script iterates over each ID and calls the fasterq- dump command-line tool using Python’s su b pro ce ss.r u n() function. For each run ID, the script instructs fasterq-dump to skip technical reads, split the output into separate les if paired­end sequencing is used, and utilize four threads for parallel processing to speed up the download. The output is saved in the previously created fastqs directory. The script includes basic error handling for each download, printing an error message if a specic run fails to download. After all run IDs have been processed, the script prints a nal message indicating that all downloads have completed. This script provides a practical and automated way to batch-download sequencing data from the SRA while ensuring data is organized and that users are alerted to potential conicts or errors.
5.5 RNA-SEQ DATA ANALYSIS PIPELINE
The RNA-Seq pipeline developed for RA analysis in Python is designed to be modular, robust, and adaptable to both simulated and real-world datasets. It encompasses the complete bioinformatics