Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
RNA Sequencing
5
5.1 INTRODUCTION
Gene expression is the intricate biological process through which genetic information encoded
within DNA is transcribed and translated into functional molecules such as proteins and noncoding RNAs. This process is fundamental to cellular function, development, and adaptability,
ensuring that each cell type executes its specialized role. The regulation of gene expression is highly
dynamic, involving multiple layers of control, including transcriptional, post-transcriptional, translational, and post-translational mechanisms. These regulatory networks allow cells to respond to
intrinsic and extrinsic stimuli, including environmental changes, developmental cues, and immune
challenges.
The immune system is particularly dependent on precise gene expression control, as it must
rapidly adapt to pathogens while maintaining self-tolerance. A breakdown in the regulation of gene
expression can lead to immune dysfunction, contributing to conditions such as autoimmune diseases. Autoimmune disorders arise when immune cells erroneously target the body’s own tissues,
causing chronic inammation and tissue damage. The pathogenesis of these diseases is inuenced
by a complex interplay between genetic predisposition and environmental factors, with gene expression playing a crucial role in disease initiation, progression, and severity. Dysregulated gene expression can result in the aberrant activation of immune cells, an imbalance between pro-inammatory
and anti-inammatory cytokines, and a failure to establish immune tolerance.
The advent of high-throughput sequencing technologies, particularly RNA sequencing (RNASeq), has revolutionized the study of gene expression in autoimmune diseases. These approaches
enable researchers to systematically prole transcriptomic changes in patients, uncovering gene
expression signatures associated with different disease states. For example, systemic lupus erythematosus (SLE) is characterized by an elevated expression of interferon-stimulated genes, highlighting the role of type I interferons in disease pathogenesis. Similarly, impaired expression of
genes involved in regulatory T-cell (Treg) function has been implicated in the loss of immune
tolerance, a hallmark of autoimmune diseases such as type 1 diabetes and multiple sclerosis (MS).
By analyzing these expression patterns, researchers gain insights into the molecular mechanisms driving autoimmunity and identify potential biomarkers for early diagnosis and disease
stratication.
Gene expression analysis also facilitates the development of targeted therapies, allowing for
the design of interventions that modulate specic disease-related pathways. Biologic drugs such
as monoclonal antibodies targeting tumor necrosis factor-alpha (TNF-α) or interleukin-6 (IL-6)
have been developed based on transcriptomic studies and have transformed the management of
diseases like rheumatoid arthritis and inammatory bowel disease. Moreover, the emergence of
personalized medicine underscores the importance of tailoring treatments based on individual gene
expression proles. Since patients exhibit varying transcriptomic responses to immunosuppressive
drugs, pharmacogenomic studies help optimize therapeutic strategies, reducing adverse effects and
enhancing efcacy.
Beyond genetic inuences, epigenetic modications signicantly impact gene expression in
autoimmune diseases. DNA methylation, histone modications, and non-coding RNAs serve as
regulatory elements that ne-tune gene activity in response to environmental stimuli. External
factors such as viral infections, diet, and stress can induce epigenetic changes that alter immune
function and contribute to disease susceptibility. For instance, DNA hypomethylation of immunerelated genes has been observed in SLE patients, leading to the hyperactivation of immune cells and
164 DO I: 10 .1201/9 7810 0 36 85432-5

165 RNA Sequencing
increased autoantibody production. These ndings emphasize the need to integrate epigenetic data
with transcriptomic analyses to obtain a comprehensive understanding of autoimmunity.
Experimental studies of gene expression in autoimmune diseases rely on meticulously designed
RNA-Seq workows to ensure robust and biologically meaningful results. A well-structured study
begins with the clear denition of research objectives, such as comparing gene expression proles between diseased and healthy individuals or assessing transcriptional changes in response to
treatment. Patient cohort selection is critical, requiring stringent inclusion and exclusion criteria
to minimize confounding variables. Factors such as disease severity, medication history, genetic
background, and demographic characteristics must be carefully controlled to enhance the reliability
of ndings.
Once the study design is established, biological samples are collected from relevant tissues,
including peripheral blood mononuclear cells (PBMCs), synovial tissue in rheumatoid arthritis,
or intestinal biopsies in Crohn’s disease. The choice of sample type is crucial, as autoimmune diseases exhibit cell-type-specic gene expression changes. To preserve RNA integrity, samples are
processed under standardized conditions, and RNA extraction protocols are optimized to minimize degradation. The quality of extracted RNA is assessed using bioanalyzers, ensuring high
RNA integrity numbers (RIN) before sequencing. Depending on the study’s goals, either poly(A)
enrichment or ribosomal RNA depletion is performed to focus on coding or total RNA populations,
respe ctive ly.
Library preparation involves converting RNA into cDNA, fragmenting it, and ligating sequencing adapters. To reduce amplication bias, unique molecular identiers (UMIs) may be incorporated. Sequencing is then carried out on high-throughput platforms such as Illumina NovaSeq, with
sequencing depth tailored to the experimental objectives. For differential gene expression analysis,
read depths of 30–50 million reads per sample are typically sufcient, whereas studies investigating
alternative splicing or rare transcripts may require deeper sequencing.
Post-sequencing, bioinformatics pipelines process the raw sequencing data. Reads are rst subjected to quality control using FastQC to assess sequencing accuracy, remove adapter sequences,
and lter out low-quality bases. High-quality reads are then aligned to the reference genome using
alignment tools like STAR or HISAT2, ensuring accurate mapping. Gene expression levels are
quantied using featureCounts or Salmon, and normalization techniques such as transcripts per
million (TPM) or variance-stabilizing transformation are applied to account for sequencing biases.
Differential expression analysis is conducted using statistical frameworks such as DESeq2 or
edgeR, identifying genes that exhibit signicant expression changes between disease and control
groups. Functional enrichment analyses follow, employing tools like gene set enrichment analysis (GSEA) or DAVID to determine biological pathways associated with differentially expressed
genes. By integrating transcriptomic data with proteomic and epigenomic datasets, researchers can
construct comprehensive disease models that capture the interplay between gene regulation and
immune dysfunction.
Validation of key ndings is a critical step, ensuring the reproducibility of results. Quantitative
real-time PCR (qRT-PCR) is commonly used to conrm differentially expressed genes, while functional assays, such as cytokine secretion measurements or CRISPR-based gene perturbation, elucidate the biological relevance of identied genes. These experimental approaches bridge the gap
between gene expression studies and mechanistic insights, paving the way for novel therapeutic
strategies.
RNA-Seq data visualization techniques play an essential role in interpreting complex transcriptomic datasets. Heatmaps are used to display hierarchical clustering of gene expression patterns
across samples, while volcano plots highlight signicantly dysregulated genes. Principal component
analysis (PCA) enables the identication of sample clustering and outliers, providing a global overview of gene expression variations. Such visualizations facilitate hypothesis generation and enhance
the interpretability of large-scale transcriptomic studies.

166 Bioinformatics of Autoimmune Diseases
The integration of multi-omics approaches with gene expression analysis represents the future
of autoimmune disease research. By combining RNA-Seq with single-cell sequencing, chromatin
accessibility proling (ATAC-Seq ), and proteomics, researchers can gain unprecedented insights
into cellular heterogeneity and molecular drivers of autoimmunity. Ultimately, these advances will
contribute to more precise diagnostic tools, improved therapeutic interventions, and personalized
treatment strategies, revolutionizing the management of autoimmune disorders.
RNA-Seq has revolutionized the study of autoimmune diseases by providing a powerful tool for
uncovering molecular mechanisms, identifying diagnostic biomarkers, and guiding personalized
treatments. Unlike traditional methods, RNA-Seq enables comprehensive analysis of gene expression across different immune cell populations, tissues, and disease states, leading to breakthroughs
in understanding the underlying pathology of autoimmune disorders.
One of the most signicant successes of RNA-Seq in autoimmune disease research has been the
identication of disease-specic gene expression signatures. For instance, in SLE, a chronic autoimmune disease affecting multiple organs, RNA-Seq has revealed distinct interferon-related gene
expression patterns in patients’ blood cells. These ndings have not only conrmed the central role
of type I interferons in SLE but have also led to the development of new targeted therapies, such
as anifrolumab, which specically blocks the type I interferon receptor and has shown efcacy in
clinical trials.
Similarly, in rheumatoid arthritis (RA), RNA-Seq has been instrumental in identifying inammatory pathways that drive disease progression. By analyzing synovial tissue from RA patients,
researchers have uncovered key differences in gene expression between those who respond to standard treatments and those who do not. This has paved the way for precision medicine approaches,
where RNA-Seq data can predict which patients will benet from biologic therapies such as TNF
inhibitors or JAK inhibitors. Such stratication not only improves treatment efcacy but also minimizes unnecessary exposure to ineffective drugs and their potential side effects.
Beyond diagnosis and treatment selection, RNA-Seq has contributed to the discovery of novel
therapeutic targets. In MS, a disease characterized by immune-mediated damage to the nervous
system, RNA-Seq has helped identify previously unrecognized immune cell subtypes involved in
disease progression. This has led to the exploration of new drug candidates that modulate these
cell populations, potentially offering more effective treatment options for MS patients who do not
respond to conventional therapies.
RNA-Seq has also facilitated the understanding of complex autoimmune conditions such as
inammatory bowel disease (IBD), which includes Crohn’s disease and ulcerative colitis. By proling intestinal tissue samples, researchers have identied molecular subtypes of IBD, some of which
respond differently to available treatments. This knowledge is being used to develop personalized
therapeutic strategies, ensuring that patients receive the most appropriate interventions based on
their specic gene expression proles.
Perhaps one of the most exciting applications of RNA-Seq in autoimmune disease research is
the ability to track disease progression and treatment response over time. By sequencing RNA from
patient samples at multiple time points, scientists can monitor how immune activity changes in
response to therapy, providing insights into disease remission and potential relapses. This real-time
molecular monitoring has the potential to transform disease management, allowing clinicians to
adjust treatments proactively rather than reactively.
The success of RNA-Seq in autoimmune disease research highlights its potential as a cornerstone
of future diagnostic and therapeutic approaches. As sequencing technologies continue to advance,
costs decrease, and computational methods improve, RNA-Seq is poised to become an integral
part of routine clinical practice. Its ability to provide detailed molecular insights into immune dysfunction not only enhances our understanding of autoimmune diseases but also brings us closer to
achieving truly personalized medicine for patients suffering from these complex conditions.
In the following sections, we demonstrate how to retrieve both raw data and metadata from the
NCBI SRA database. The raw data typically consists of sequencing reads in FASTQ format, while

167 RNA Sequencing
the metadata includes study design information, which is essential for meaningful RNA-Seq data
analysis. If you are working with your own data, you can substitute your own study design and make
minimal adjustments to the provided code to suit your specic objectives. To keep the focus on the
workow rather than the biological interpretation, gene names in the results, such as tables and
plots, are anonymized using labels like gene1, gene2, and so on. All the scripts used in this chapter
are available in the GitHub repository referenced in the introduction of this book.
5.2 EXTRACTING BIOPROJECT STUDY DESIGN FROM SRA DATABASE
We discussed the principles of study design in Chapter 3. A well-dened study design is essential
prior to performing RNA-Seq experiments, as it determines the structure of the experiment, the
number and nature of sample groups, and the specic biological questions to be addressed. By
the time we enter the RNA-Seq data analysis phase, critical decisions (such as sample groupings,
experimental conditions, and the hypotheses to be tested) should already be in place.
Researchers commonly deposit raw RNA-Seq data into public repositories such as the NCBI
Sequence Read Archive (SRA), along with associated metadata that describes the study, sample
characteristics, and sequencing details. This standardized archiving ensures data transparency,
reproducibility, and reusability by the broader scientic community.
In this chapter, we demonstrate how to access and retrieve RNA-Seq raw data from the NCBI
SRA database. Specically, we focus on the NCBI BioProject PRJEB52174, which contains
paired-end FASTQ les representing 184 RNA-Seq samples from human subjects diagnosed with
rheumatoid arthritis (RA). To simplify our analysis and reduce computational demands, we will
work with a representative subset of 20 samples (National Center for Biotechnology Information,
2022).
The associated study metadata (such as SRA run accession numbers, experimental conditions,
patient demographics, and disease states) can be retrieved either manually through the SRA Run
Selector interface on the NCBI website or programmatically using scripting tools such as Bash or
Python. For reproducibility and automation, we will use Python to download and parse the metadata, after which we will select the 20 samples to be included in our downstream analyses.
The Python script “sr a _ ru n s el ec to r.p y ” is designed to extract detailed study design
metadata from the SRA (Sequence Read Archive) database at NCBI, specically for a given
BioProject, and save that metadata into a CSV le. The goal is to capture not just the basic run and
sample information but also the custom annotations found in the BioSample, such as tissue type,
time point, treatment conditions, or developmental stage. These annotations are commonly used in
the SRA Run Selector interface on NCBI’s website and are critical for understanding the experimental context of the sequencing data.
To begin with, the script uses the BioPython Entrez module to search the SRA database for
all entries associated with the specied BioProject ID. The esearch function retrieves a list of
unique IDs corresponding to SRA records. These IDs are then individually processed in a loop,
where each one is fetched in XML format using efetch. However, instead of using Entrez.
re ad() to parse the XML, as was done initially, the script now uses Python’s built-in xm l.etr ee.
ElementTree module. This change is necessary because many SRA XML records do not include
a formal Document Type Denition (DTD) or XML Schema, which makes them incompatible with
BioPython’s strict parser.
For each XML record retrieved, the script navigates the tree structure to locate the relevant
sections: EXPERIMENT_PACKAGE, RUN, and SAMPLE. It extracts the run accession number,
the BioSample ID, and the scientic name of the sample. The most important section is SAMPLE_
ATTRIBUTES, where the submitter-dened metadata resides. This part of the XML contains
repeated elements with a TAG and VALUE, which represent various custom attributes describing
the sample and study design. The script iterates through these attributes and dynamically adds them
as new columns in a metadata dictionary. Each dictionary corresponds to a row in the nal dataset.

168 Bioinformatics of Autoimmune Diseases
FIGURE 5.1 Partial view of the metadata contained in the CSV le.
All rows are collected into a list and converted into a pandas DataFrame, which is then saved
as a CSV le (Figure 5.1). The output le contains standard identiers such as run and sample
accession numbers, along with any additional study-specic elds submitted by the researchers. By
using ElementTree, the script ensures compatibility with the XML format of SRA records while
capturing rich contextual metadata that can be used for downstream analysis or curation. It also
includes error handling and respects NCBI’s rate-limiting guidelines by inserting a delay between
fetch requests. This approach provides a reliable and scalable way to programmatically extract
structured metadata from the SRA for any BioProject of interest.
5.3 STUDY VARIABLES
After downloading the study metadata, we selected a set of variables, listed in Table 5.1, that will
be used in the RNA-Seq data analysis. The clinical variables and their relevance to the study are
described in detail below.
TABLE 5.1
Selected Variables of the Rheumatoid Arthritis Study
runID Age Anti-CCP ADAS cellType CRP ESR Ethnic Pathotype RF Sex SJC TJC
ERR9539313 40 Positive 75 Bpoor 39 67 African Lymphoid Positive Female 5 11
ERR9539319 56 Negative 98 Brich 11 30 African Lymphoid Positive Female 2 8
ERR9539328 46 Positive 95 Brich 29 26 African Lymphoid Negative Female 11 15
ERR9539343 45 Positive 4 Brich 14 19 African Lymphoid Positive Male 6 2
ERR9539358 24 Positive 71 Brich 36 52 African Lymphoid Positive Female 4 16
ERR9539216 67 Negative 17 Bpoor 0 5 African Fibroid Negative Female 10 12
ERR9539393 38 Positive 40 Brich 9 5 Asian Lymphoid Positive Male 7 5
ERR9539217 41 Positive 82 Bpoor 16 42 Asian Fibroid Positive Female 18 19
ERR9539241 39 Negative 21 Brich 6 52 Asian Lymphoid Positive Male 0 10
ERR9539243 43 Positive 89 Brich 89 131 Asian Lymphoid Positive Female 3 5
ERR9539246 59 Positive 85 Bpoor 0 35 Asian Fibroid Positive Female 3 11
ERR9539251 42 Negative 70 Bpoor 13 45 Asian Lymphoid Negative Female 11 28
ERR9539213 31 Positive 93 Bpoor 6 12 Caucasian Fibroid Negative Female 14 18
ERR9539320 62 Positive 55 Bpoor 0 20 Caucasian Fibroid Positive Female 5 6
ERR9539331 60 Negative 41 Brich 4 12 Caucasian Lymphoid Positive Female 3 2
ERR9539353 69 Negative 56 Brich 1 5 Caucasian Lymphoid Negative Male 1 2
ERR9539364 70 Positive 49 Bpoor 6 25 Caucasian Fibroid Positive Female 16 13
ERR9539375 59 Positive 57 Bpoor 0 8 Caucasian Myeloid Negative Female 7 13

i. Anti-cyclic citrullinated peptide antibody measurement
Anti-cyclic citrullinated peptide (anti-CCP) antibody measurement is a blood test used
primarily to help diagnose rheumatoid arthritis (RA). This test detects the presence of specic antibodies that the immune system produces against peptides or proteins that contain
the amino acid citrulline (a modied form of the amino acid arginine). In people with RA,
the immune system mistakenly targets these citrullinated proteins as if they were foreign
invaders.
The process of citrullination happens normally in the body, but in autoimmune conditions like RA, it becomes excessive or dysregulated, leading to an immune response. The
anti-CCP test looks for antibodies directed against cyclic citrullinated peptides, which
are synthetic versions of citrullinated proteins. A high level of anti-CCP antibodies in the
blood is strongly associated with RA and can even appear in the blood before symptoms
start, making the test useful for early diagnosis.
ii. B Cells
In RA, B cells play a signicant and multifaceted role in driving the autoimmune response
that leads to chronic joint inammation and tissue damage. B cells are a type of white
blood cell that are part of the adaptive immune system, normally responsible for producing
antibodies to ght infections. However, in autoimmune diseases like RA, their functions
become dysregulated, contributing to the development and persistence of the disease.
One of the hallmark contributions of B cells in RA is their production of autoantibodies, such as rheumatoid factor (RF) and anti-CCP antibodies. These autoantibodies target
the body’s own proteins, forming immune complexes that deposit in the joints and trigger
inammatory cascades. The presence of these antibodies is strongly associated with disease severity and progression, and anti-CCP antibodies in particular are highly specic to
RA. B cells also function as antigen-presenting cells, meaning they process and present
antigens to T cells, which further amplies the autoimmune response.
Beyond antibody production and antigen presentation, B cells secrete pro-inammatory
cytokines such as IL-6, which contribute to the recruitment and activation of other immune
cells in the synovial membrane. This leads to the formation of pannus tissue, an abnormal
thickening of the synovial lining that invades cartilage and bone. B cells are found in
increased numbers within the inamed synovial tissue of RA patients, often organized
into structures that resemble germinal centers (GCs), similar to those seen in lymph nodes.
The central role of B cells in RA pathogenesis has made them a therapeutic target.
Treatments such as rituximab, a monoclonal antibody that depletes B cells by targeting
CD20, have been effective in reducing disease activity in RA patients, particularly those
who are seropositive for RF or anti-CCP antibodies. This underscores the importance of
B cells not just in initiating the autoimmune process but also in sustaining the chronic
inammation characteristic of rheumatoid arthritis.
In the context of RA, the terms B-cell rich, B-cell poor, and GC refer to distinct patterns of immune cell organization and activity observed in the synovial tissue, which lines
the joints. These patterns are identied through histological and immunological analysis,
often during research studies or biopsy-based evaluations, and they help characterize the
immune microenvironment of the affected joints. Understanding these categories can provide insights into the disease mechanism and may even guide therapeutic strategies.
B-cell-rich synovitis describes synovial tissue that contains a high number of B lymphocytes. In this pattern, B cells are often found clustered together, sometimes forming
organized structures resembling GCs, which are areas where B cells proliferate, mutate
their antibody genes, and undergo selection. These structures are similar to those found
in lymphoid tissues like lymph nodes and suggest a highly active and organized immune
response within the joint. B-cell-rich synovitis is typically associated with the presence of
autoantibodies such as rheumatoid factor and anti-CCP antibodies, as well as more severe
169 RNA Sequencing

170 Bioinformatics of Autoimmune Diseases
inammation. Patients with this pattern may respond well to therapies that target B cells,
such as rituximab.
iii. C-reactive protein
C-reactive protein (CRP) is a key biomarker of inammation that is widely used to assess
disease activity and monitor response to treatment. CRP is a protein produced by the liver
in response to inammation, particularly when stimulated by cytokines like IL-6. Its levels in the blood rise rapidly during episodes of acute inammation and fall just as quickly
when the inammation subsides, making it a sensitive and dynamic indicator.
In RA, which is a chronic inammatory autoimmune disease, elevated CRP levels often
reect active joint inammation and tissue damage. Clinicians use CRP values alongside
other measures, such as joint counts, patient symptoms, and imaging, to determine how
active the disease is. A high CRP suggests ongoing inammation and possibly uncontrolled disease, while a low or normal CRP can indicate effective disease control or remission. However, CRP is a general marker of inammation and is not specic to RA; it can
be elevated in infections, injuries, or other autoimmune conditions as well.
CRP is also a component of several composite disease activity scores used in RA, such
as the Disease Activity Score 28-CRP (DAS28-CRP) and Simplied Disease Activity
Index (SDAI). In these indices, CRP contributes to the overall evaluation of disease severity. Tracking CRP over time helps rheumatologists adjust medications, evaluate ares, and
assess whether a patient is responding well to therapy. Overall, CRP is a practical, accessible, and clinically valuable tool in the diagnosis and management of rheumatoid arthritis.
iv. Erythrocyte sedimentation rate
Erythrocyte sedimentation rate (ESR) is a common blood test used to measure the level
of inammation in the body. It reects how quickly red blood cells (erythrocytes) settle to
the bottom of a test tube over a specied period, typically 1 hour. In healthy individuals,
red blood cells settle slowly, but in the presence of inammation, certain proteins such as
brinogen cause the cells to clump together and settle more rapidly. A higher ESR indicates a higher level of systemic inammation.
In RA, which is a chronic autoimmune condition characterized by persistent joint
inammation, ESR is used as a marker of disease activity. It helps clinicians assess how
active the disease is, monitor the patient’s response to therapy, and detect disease ares.
Elevated ESR is often found in patients with active RA and may correlate with symptoms
such as joint pain, swelling, stiffness, and fatigue. However, like CRP, ESR is a nonspecic
test, meaning it can also be elevated in other conditions like infections, other autoimmune
diseases, or even some cancers.
ESR is included in several composite scores used to measure disease activity in RA,
such as the Disease Activity Score 28 (DAS28-ESR). In this context, ESR is combined
with joint assessments and patient-reported symptoms to generate a numerical score that
helps guide treatment decisions. Although it is a useful tool, ESR can be inuenced by factors unrelated to RA, such as age, sex, and anemia. Therefore, it is typically interpreted in
conjunction with other clinical ndings and tests to provide a more accurate picture of the
patient’s disease status.
v. Arthritis disease activity score measurement
The Arthritis Disease Activity Score (ADAS) is a quantitative tool used by clinicians to
measure the severity of disease activity in patients with RA. It combines clinical assessments and, in some versions, laboratory results to produce a single numerical score that
reects how active the disease is at a given time. The score helps physicians determine
the level of inammation, monitors how the disease is progressing, and guides treatment
decisions.
The most commonly used version is the DAS28, which stands for disease activity score
using 28 joints. In this method, 28 specic joints are examined for tenderness and swelling,

usually including the shoulders, elbows, wrists, knees, and joints of the hands. The score
)
=
)
+
)
=
++
+ 0.96
)
also incorporates the patient’s general health assessment, which is typically provided by
the patient on a scale of 0–100 (or sometimes 0–10), and either the ESR or CRP, which are
both blood tests that measure inammation. The nal DAS28 score is calculated using a
mathematical formula that combines these elements. The formula of DAS28 using ESR
(Prevoo et al., 1995) is the following:
171 RNA Sequencing
AS28(ESR
where TJC28 = tender joint count (28 joints), SJC28 = swollen joint count (28 joints),
ESR= erythrocyte sedimentation rate (mm/hr), GH = general health assessment (usually
patient global assessment on a 0–100 mm visual analog scale).
If CRP is used instead of ESR, the formula slightly changes (Fransen & van Riel, 2005):
AS28(CRP
The resulting score usually falls between 0 and 10. Scores are interpreted in ranges:
a score above 5.1 indicates high disease activity, between 3.2 and 5.1 indicates moderate
activity, between 2.6 and 3.2 indicates low activity, and a score below 2.6 is considered
remission. A decrease in DAS28 over time suggests improvement in disease control, while
an increase suggests a worsening of RA.
Overall, DAS28 is a standardized and validated tool that allows for consistent monitoring of rheumatoid arthritis activity across different visits, clinics, and clinical trials.
It helps ensure that treatment decisions are based on objective measurements rather than
subjective impressions alone.
vi. Pathotype (Lymphoid, Myeloid, Fibroid)
The term pathotype refers to distinct patterns of immune cell inltration and tissue organization found in the synovial membrane, which lines the joints. These patterns reect different underlying mechanisms of inammation and can be classied into three main types:
lymphoid, myeloid, and broid. Understanding these pathotypes provides insight into the
heterogeneity of RA and may eventually guide more personalized treatment strategies.
The lymphoid pathotype is characterized by the presence of organized clusters of B
cells and T cells in the synovial tissue. These immune cells often form structures resembling GCs, which are typically found in lymphoid organs like lymph nodes. This organization suggests a highly active, antigen-driven immune response occurring directly within
the joint. Patients with this pathotype often have autoantibodies such as rheumatoid factor
(RF) or anti-CCP, and the lymphoid pattern is associated with more severe inammation
and joint damage. These patients may respond well to therapies targeting B cells or specic
cytokines like IL-6.
The myeloid pathotype is dominated by macrophages and neutrophils, which are cells
of the innate immune system. This type is typically associated with strong expression of
pro-inammatory cytokines such as tumor necrosis factor (TNF) and IL-1. The myeloid
pattern tends to reect a different kind of immune activity that is more innate and less
dependent on the adaptive immune system. Patients with a myeloid pathotype may respond
particularly well to TNF inhibitors, which are a common class of biologic drugs used in
RA treatment.
The broid pathotype is distinguished by a lack of signicant immune cell inltration.
Instead, the synovial tissue shows broblast proliferation and tissue remodeling without
the high levels of lymphocytes or macrophages seen in the other types. This suggests a
lower level of active inammation, and patients with this pattern may have less joint swelling and fewer systemic symptoms. However, the broid pattern may still be associated with
0.56 × TJC28 + 0.28 × SJC28 + 0.70 × ln (ESR
0.56 ×TJC28 +0.28 ×SJC28 +0.36 ×ln (CRP
0.014 × GH
0.014 ×GH
1

172 Bioinformatics of Autoimmune Diseases
joint stiffness and structural changes due to chronic broblast activity and extracellular
matrix deposition.
These pathotypes reect different immune processes operating in RA and highlight
the complexity of the disease. Identifying which pathotype a patient has (through synovial
biopsy and molecular analysis) could help tailor therapies more precisely in the future,
moving beyond the one-size-ts-all approach currently used in clinical practice.
vii. Rheumatoid factor measurement
Rheumatoid factor (RF) measurement is a blood test used to detect the presence of autoantibodies (specically, antibodies directed against the Fc portion of immunoglobulin G
(IgG)) which are commonly found in people with RA. RF was one of the earliest biomarkers associated with RA and remains a widely used tool in diagnosing and classifying the
disease.
In RA, the immune system becomes dysregulated and begins to produce RF, which
binds to the body’s own antibodies, forming immune complexes that contribute to inammation in the joints. The presence of RF in the blood suggests that the body is mounting
an abnormal immune response, characteristic of autoimmune conditions. RF is detected
using techniques like nephelometry, ELISA (enzyme-linked immunosorbent assay), or
latex agglutination, and the results are usually reported in international units per milliliter
(IU/mL). A value above the laboratory’s reference range (commonly around 20–30 IU/mL)
is considered positive.
However, RF is not specic to RA. It can also be elevated in other autoimmune diseases
(like Sjögren’s syndrome or lupus), chronic infections (such as hepatitis C or tuberculosis), and even in some healthy older adults. For this reason, RF is typically interpreted in
combination with clinical ndings and other tests, especially the anti-cyclic citrullinated
peptide (anti-CCP) antibody test, which is more specic for RA.
A positive RF test, especially at high levels, supports a diagnosis of rheumatoid arthritis
and is often associated with more severe or erosive disease. It may also be used as part of
the classication criteria for RA and to help predict disease progression. Despite its limitations in specicity, RF remains an important marker in the diagnostic workup and ongoing
assessment of RA.
viii. Swollen Joint Count
Swollen joint count (SJC) is a clinical measure used to assess joint inammation and disease activity. It refers to the number of joints that appear swollen due to synovitis, which is
the inammation of the synovial membrane lining the joints. Swelling in RA is typically
a sign of active disease and is associated with pain, stiffness, and long-term joint damage
if not treated.
SJC is typically assessed by a trained clinician who performs a systematic physical
examination of the patient’s joints. The most commonly used format is the SJC28, which
involves evaluating 28 specic joints, including the shoulders, elbows, wrists, metacarpophalangeal (MCP) joints, proximal interphalangeal (PIP) joints, and knees. Each joint
is examined for the presence of visible or palpable swelling, which is indicative of uid
accumulation or synovial thickening.
The total count of swollen joints is used as part of several composite disease activity
scores, such as the DAS28, SDAI (Simplied Disease Activity Index), and CDAI (Clinical
Disease Activity Index). A higher SJC suggests more active inammation and typically
correlates with greater disease severity. Conversely, a low or zero SJC may indicate remission or effective control of RA with treatment.
SJC is an essential tool in both clinical practice and research. It provides an objective
measure that, when combined with other assessments like tender joint count (TJC), patientreported symptoms, and inammatory markers like CRP or ESR, helps guide treatment
decisions and track response over time.

ix. Tender Joint Count
Tender joint count (TJC) is a clinical measure used to assess how many joints are painful
or sensitive to pressure during a physical examination. It reects the subjective experience
of joint pain from the patient and is a key indicator of disease activity, particularly how
inammation is affecting joint function and causing discomfort.
TJC is usually assessed by a healthcare provider who gently presses on specic joints
to check for tenderness. The most common format is the TJC28, which evaluates 28 joints,
including the shoulders, elbows, wrists, knees, and the small joints of the hands (MCP and
PIP joints). If pressing on a joint causes pain or discomfort, it is recorded as tender. The
total number of tender joints is then summed to produce the TJC score.
This measure is often used alongside the SJC and other indicators such as CRP
(C-reactive protein) or ESR to calculate composite scores like the DAS28, SDAI, and
CDAI, which provide a more comprehensive view of disease activity in RA. A high TJC
indicates more joint pain, which can reect active inammation, but it may also be inuenced by other factors such as fatigue, bromyalgia, or individual pain sensitivity.
TJC is important for monitoring disease progression and treatment response. While it
is more subjective than SJC—since it depends on the patient’s pain perception—it still
provides valuable information, especially when tracked over time. In clinical settings, a
decrease in TJC suggests improvement and effective disease control, while an increase
may signal a are or inadequate response to therapy.
5.4 DOWNLOADING THE RNA-SEQ RAW DATA
173 RNA Sequencing
Save the run identiers “runID” from Table 5.1 in a text le named “ids.txt”. Then run the Python
script “download _ fastq.py”, which is designed to automate the process of downloading
FASTQ les from the Sequence Read Archive (SRA) using run identiers provided in the text le.
It begins by dening the name of the input le and the directory where the downloaded les will
be stored, in this case, a folder called “fastqs”. The script checks whether this directory already
exists. If it does, it prints a warning message to notify the user that the directory is present and then
exits the program without performing any downloads. This safety check helps prevent accidental
overwriting of existing data or mixing new les with previously downloaded ones.
If the fastqs directory does not exist, the script proceeds to create it, ensuring that there is a
dedicated location to store the downloaded data. It then attempts to open and read the id s.t xt le,
which should contain a list of SRA run accession IDs, one per line. If the le is missing or cannot
be opened, the script prints an error message and exits, avoiding any further execution that would
otherwise rely on non-existent input data.
Once the run IDs are successfully read, the script iterates over each ID and calls the fasterq-
dump command-line tool using Python’s su b pro ce ss.r u n() function. For each run ID, the
script instructs fasterq-dump to skip technical reads, split the output into separate les if pairedend sequencing is used, and utilize four threads for parallel processing to speed up the download.
The output is saved in the previously created fastqs directory. The script includes basic error
handling for each download, printing an error message if a specic run fails to download. After all
run IDs have been processed, the script prints a nal message indicating that all downloads have
completed. This script provides a practical and automated way to batch-download sequencing data
from the SRA while ensuring data is organized and that users are alerted to potential conicts or
errors.
5.5 RNA-SEQ DATA ANALYSIS PIPELINE
The RNA-Seq pipeline developed for RA analysis in Python is designed to be modular, robust, and
adaptable to both simulated and real-world datasets. It encompasses the complete bioinformatics
Соседние файлы в папке Библиотека им академика М.И. Перельмана
