Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5529_Библиотеки_им_академика_М_И_Перельмана
.pdf
284 Bioinformatics of Autoimmune Diseases
a ligand-activated transcription factor that regulates both innate and adaptive immune responses.
Upon activation, AhR signaling can promote the generation of Tregs and suppress Th17 differentiation, similar to SCFA effects, though via a distinct molecular pathway. Dysbiosis that results in
lower levels of these tryptophan metabolites compromises AhR signaling and thereby contributes
to the expansion of inammatory T-cell subsets. Notably, studies using germ-free mice or mice
colonized with altered microbial compositions show exacerbated autoimmune symptoms, which
can be ameliorated by reintroducing specic metabolite-producing strains or supplementing the
metabolites directly.
These ndings underscore the critical role of the gut microbiota as a regulatory hub that integrates environmental, dietary, and immunological signals to orchestrate immune homeostasis.
Importantly, they reveal that immune cell differentiation is not solely a host-intrinsic process but
is dynamically inuenced by the microbial consortia residing in the gastrointestinal tract. In autoimmune diseases, the balance between tolerance and inammation is frequently disturbed, and
the loss of regulatory microbial signals tilts the immune equilibrium toward pathogenic inammation. Therapeutic strategies that aim to restore benecial microbial metabolites (either through
dietary interventions, probiotics, prebiotics, or fecal microbiota transplantation) are therefore being
explored as innovative approaches to reestablish immune tolerance.
Furthermore, recent metagenomic and metabolomic studies in patients with autoimmune diseases have consistently shown reduced levels of SCFAs and altered tryptophan metabolism. For
example, patients with ulcerative colitis and Crohn’s disease exhibit signicantly lower fecal butyrate concentrations, correlating with disease severity and increased Th17 inltration in the intestinal
mucosa. Similar patterns have been observed in MS, where the gut microbiota composition favors
species that promote Th17 over Treg differentiation. These clinical ndings are further supported
by animal models, in which administration of butyrate or tryptophan metabolites ameliorates disease symptoms and restores immune balance.
9.2.5 CHRONIC IMMUNE STIMULATION
Chronic immune stimulation by persistent bacterial antigens or endotoxins represents a powerful and
insidious driver of immune system dysregulation. The immune system is designed to recognize and
eliminate microbial invaders efciently, mounting a robust but self-limited response that resolves with
pathogen clearance. However, in certain scenarios, components of bacterial pathogens (such as LPS,
peptidoglycans, and unmethylated Cytosine–Phosphate–Guanine dinucleotide (CpG) DNA) may linger in host tissues due to persistent infection, biolm formation, or even translocation of microbial
products from mucosal barriers like the gut. These microbial molecules act as potent PAMPs, continuously engaging PRRs such as TLRs and NOD-like receptors (NLRs). This ongoing signaling
cascade sustains the activation of innate immune cells, leading to a pro-inammatory environment
marked by elevated cytokine production, tissue inltration by immune cells, and activation of APCs.
Under physiological conditions, immune tolerance mechanisms (including regulatory T cells
(Tregs), immune checkpoint pathways, and central deletion of autoreactive clones) maintain the
balance between immune defense and self-tolerance. However, when the immune system is subjected to persistent bacterial stimuli, these regulatory networks may become overwhelmed or dysfunctional. Continuous exposure to endotoxins such as LPS can desensitize TLR signaling, but
paradoxically may also lead to chronic low-grade inammation, sometimes referred to as “inammaging” in elderly individuals. Monocytes and dendritic cells exposed to persistent LPS signaling
often undergo phenotypic changes, skewing toward a semi-activated state that continuously secretes
inammatory mediators like TNF-α, IL-6, and IL-1β. This persistent activation distorts antigen
presentation and disrupts peripheral tolerance, priming the immune system to react against selfantigens in genetically or epigenetically susceptible individuals.
Over time, this shift from a regulated, homeostatic immune response to a chronically inamed
state can have profound consequences. One major pathological outcome is the breakdown of

285 Roles of Bacteria in Autoimmune Diseases
tolerance and the initiation of autoimmunity. Chronic bacterial antigen exposure can induce epitope spreading, a process by which the immune response initially directed at microbial antigens
begins to target structurally similar or co-localized self-antigens. Furthermore, repeated activation of Bcells by microbial antigens in a pro-inammatory cytokine environment promotes class
switching and somatic hypermutation, which may inadvertently result in the generation of autoreactive B cell clones. These clones can differentiate into plasma cells that secrete autoantibodies, such
as anti-nuclear antibodies (ANA), rheumatoid factor (RF), or anti-citrullinated protein antibodies
(ACPA), depending on the tissue environment and context.
The generation of autoantibodies marks a critical transition from a merely dysregulated immune
system to one actively participating in autoimmune disease pathology. These antibodies can form
immune complexes that deposit in tissues, activate complement, and recruit effector cells, leading to further inammation and tissue damage. In diseases like SLE, chronic immune stimulation
is believed to be one of the initiating events, particularly through exposure to bacterial DNA or
nucleoprotein antigens. In RA, translocation of gut bacterial products may contribute to the initial citrullination of host proteins in the lungs or synovium, subsequently targeted by the adaptive
immune system.
Recent metagenomic and transcriptomic studies have revealed that patients with autoimmune
diseases often harbor a distinct microbial signature characterized by increased abundance of pathobionts (commensal bacteria that can become harmful under certain conditions) and decreased
microbial diversity. These microbial communities may produce metabolites or antigens that perpetuate immune activation. Additionally, gut barrier dysfunction, commonly termed “leaky gut”, has
been implicated in allowing bacterial endotoxins to enter systemic circulation, acting as a chronic
stimulus for immune activation. This supports the hypothesis that the gut-immune axis is not merely
involved in nutrient absorption and digestion but also plays a central role in systemic immune regulation and the pathogenesis of autoimmunity.
At the molecular level, the persistent activation of intracellular signaling pathways such as NF-κB,
STAT3, and MAPKs in response to bacterial products promotes the transcription of genes involved
in inammation, cell survival, and antigen presentation. Chronic stimulation of these pathways
can lead to the exhaustion of regulatory cell populations and the expansion of pro-inammatory
Th subsets like Th17 cells, which are commonly found in autoimmune lesions. IL-17, a hallmark
cytokine of this lineage, promotes neutrophil recruitment, matrix degradation, and further amplies the inammatory cascade. In animal models, blocking IL-17 or its receptor has been shown to
ameliorate autoimmune pathology induced by persistent microbial stimuli.
Taken together, chronic immune stimulation by bacterial antigens or endotoxins acts as a dangerous double-edged sword. While intended as a defense mechanism, its persistent activation creates
a pathological loop where inammation breeds more inammation, self-tolerance is progressively
eroded, and the immune system begins to perceive host tissues as foreign. The chronicity of the
stimulus appears to be a key determinant; transient exposure to bacteria often leads to resolution
and memory, but persistent exposure drives dysfunction and disease. As a result, therapeutic strategies that restore mucosal barriers, modulate the microbiome, or block chronic PRR signaling are
emerging as promising avenues for preventing or managing autoimmune diseases rooted in microbial persistence.
9.2.6 EPIGENETIC MODIFICATIONS
Epigenetic modications serve as a crucial interface between environmental stimuli and gene regulation, providing a exible and reversible mechanism through which cells adapt to external cues. In
the context of autoimmunity, the role of epigenetics becomes especially signicant when considering microbial inuences. Among these, bacterial components (such as toxins, metabolites, and cell
wall elements) have emerged as powerful modulators of the host epigenome. These microbial signals do not act indiscriminately but instead interact with the host immune system in context-specic

286 Bioinformatics of Autoimmune Diseases
manners, inuencing the development, function, and fate of immune cells. Increasingly, research
indicates that bacterial factors can alter the host’s epigenetic landscape in a way that promotes
chronic inammation, immune dysregulation, and even breaks in self-tolerance, culminating in
autoimmune disease.
DNA methylation, one of the most extensively studied epigenetic modications, involves the
addition of a methyl group to cytosine residues, particularly within CpG islands in gene promoter
regions. This process typically represses gene transcription. Certain bacterial metabolites, especially SCFAs like butyrate and propionate (produced by gut microbiota during the fermentation
of dietary bers) have been shown to affect DNA methylation patterns in immune cells. Butyrate,
for example, is known to inhibit histone deacetylases (HDACs), thereby inuencing the chromatin structure and indirectly affecting DNA methylation. While SCFAs can have anti-inammatory
effects under normal conditions, imbalances in their concentration or altered microbial composition
(dysbiosis) may lead to inappropriate epigenetic changes. In dendritic cells and macrophages, for
instance, this can result in aberrant expression of co-stimulatory molecules and pro-inammatory
cytokines that drive autoimmune responses.
Bacterial toxins offer another route through which the epigenome is modulated. For example,
the cytotoxins produced by Clostridium perfringens and H. pylori can interfere with chromatin
organization and transcriptional regulation in T and B lymphocytes. Some of these toxins have
been reported to disrupt histone modication patterns (such as acetylation and methylation), leading to sustained activation of genes involved in inammation and immune activation. Specically,
alterations in histone H3 lysine 27 trimethylation (H3K27me3) or histone acetylation states may
prime immune cells for hyper-responsiveness or dysregulated memory formation. These changes
can persist even after the resolution of infection, suggesting a form of “epigenetic memory” that
predisposes the immune system to overreact to future stimuli, including self-antigens.
A striking example is the modulation of Th cell differentiation by bacterial signals through
histone modication. Studies show that exposure to bacterial LPS and peptidoglycan, both potent
immune activators, induces epigenetic remodeling in naïve CD4+ T cells. These changes skew their
differentiation toward Th17 or Th1 phenotypes, which are known to be implicated in several autoimmune diseases such as MS, RA, and inammatory bowel disease. Through the upregulation of
transcription factors like RORγt and T-bet, bacterial components can effectively reprogram immune
responses toward a pro-inammatory trajectory. Such a skewing is often supported by histone acetylation at enhancer regions of pro-inammatory cytokine genes like IL-17 and IFN-γ, reinforcing a
chronic inammatory state.
Beyond histones and DNA methylation, non-coding RNAs such as microRNAs (miRNAs) also
play an epigenetic role in shaping immune responses, and bacteria are known to inuence their
expression. Several pathogens, including Mycobacterium tuberculosis and E. coli, have been shown
to induce or repress specic host miRNAs that control genes involved in inammation and cell
death. These miRNAs may inhibit negative regulators of inammation or suppress signaling pathways that normally promote immune tolerance. For instance, bacterial modulation of miR-155 and
miR-146a has been implicated in autoimmune pathology through the enhancement of NF-κB signaling and reduction of regulatory T-cell function.
Importantly, these epigenetic modications are not isolated events but are often coordinated in
a context-dependent manner, inuenced by the tissue microenvironment, host genetic background,
and duration of bacterial exposure. In genetically susceptible individuals, the cumulative effect of
microbial-induced epigenetic changes may tip the balance toward autoimmunity. For example, individuals carrying risk alleles in genes related to epigenetic regulation, such as MECP2 or DNMT3A,
may be more vulnerable to the epigenomic insults triggered by bacterial metabolites or toxins. In
such cases, the epigenetic changes not only contribute to disease onset but may also inuence its
progression and response to therapy.
Moreover, the reversibility of epigenetic marks offers both a challenge and an opportunity.
While persistent epigenetic reprogramming can sustain autoimmune pathology long after the initial

287 Roles of Bacteria in Autoimmune Diseases
microbial insult, it also provides a potential therapeutic avenue. Interventions targeting HDACs,
DNA methyltransferases, or histone methyltransferases are being explored to reverse pathological
gene expression programs in autoimmune diseases. Understanding the specic bacterial triggers
and the epigenetic pathways they engage can lead to the development of precision microbiomebased or epigenetic therapies.
9.3 METAGENOMICS IN AUTOIMMUNE DISEASES
Metagenomics has emerged as a powerful tool for understanding the complex interplay between
microbial communities and autoimmune diseases. Two primary approaches dominate metagenomics studies: amplicon-based sequencing, often targeting the 16S ribosomal RNA (rRNA) gene for
bacteria or internal transcribed spacer (ITS) regions for fungi, and shotgun metagenomic sequenc-
ing, which provides comprehensive insights into the taxonomic and functional composition of entire
microbial communities. In autoimmune disease research, these approaches have revealed consistent
patterns of microbial dysbiosis, imbalances in the normal gut or mucosal microbiota, which may
contribute to the onset or progression of diseases such as RA, SLE, T1D, and MS.
Amplicon-based metagenomics offers a cost-effective method to prole microbial diversity, particularly useful in large cohort studies. Recent ndings using 16S rRNA gene sequencing have
identied a reduced diversity and altered abundance of specic taxa, such as P. copri, Akkermansia
muciniphila, and Bacteroides species, in individuals with autoimmune diseases. These microbial
shifts may inuence mucosal immunity, intestinal barrier function, and systemic inammation.
Studies have demonstrated that in diseases like RA, the presence of P. copri is often associated with
disease onset and elevated inammatory markers. Similarly, in T1D, early-life microbiome changes
identied through amplicon sequencing precede autoantibody production, suggesting a potential
predictive role of microbial biomarkers.
Shotgun metagenomic sequencing provides a deeper resolution, enabling researchers not only
to identify species-level changes but also to infer microbial metabolic pathways, gene content, and
potential virulence factors. In recent applications, shotgun metagenomics has been used to reveal
functional alterations in the gut microbiome of autoimmune patients, such as the depletion of shortchain fatty acid-producing genes and enrichment of genes associated with oxidative stress and
LPS biosynthesis. These functional disruptions may promote an inammatory gut environment and
aberrant immune activation. For instance, in SLE, shotgun sequencing has revealed an enrichment
of microbial genes encoding ribosomal proteins and agellar assembly, potential drivers of innate
immune stimulation via PRRs.
Both metagenomics approaches have also enabled the development of microbiome-based classiers and risk models for predicting autoimmune disease susceptibility and monitoring treatment
response. Integration with host genomic and transcriptomic data has begun to unravel host–microbe
interactions at unprecedented resolution. Furthermore, recent research is increasingly applying
machine learning models to large metagenomic datasets to identify microbial signatures predictive of disease subtypes or progression. As sequencing technologies advance and more longitudinal datasets become available, metagenomics will continue to play a critical role in uncovering
the microbial component of autoimmune pathogenesis, opening new avenues for diagnostics and
microbiota-targeted therapies.
9.3.1 AMPLICON-BASED METAGENOMICS WORKFLOW
Amplicon-based metagenomics is a targeted approach that focuses on amplifying and sequencing specic genomic regions, most commonly the 16S rRNA gene for bacteria, 18S rRNA or ITS
regions for fungi, and other taxonomically informative markers for different microbial groups. The
workow begins with the careful collection of biological samples such as stool, saliva, skin swabs,
or mucosal biopsies from both healthy individuals and those diagnosed with autoimmune diseases.

288 Bioinformatics of Autoimmune Diseases
These samples must be handled with strict protocols to prevent contamination and preserve microbial DNA integrity. DNA is then extracted using optimized lysis protocols that account for the
wide diversity in microbial cell wall structures. Once DNA is extracted, specic hypervariable
regions of the marker genes are amplied using polymerase chain reaction (PCR) with barcoded
primers, allowing multiplexing of samples. The resulting amplicons are puried and prepared for
sequencing, commonly using platforms such as Illumina MiSeq or NovaSeq, which provide highthroughput and accurate reads suitable for taxonomic classication.
The bioinformatics analysis of amplicon-based data has been greatly streamlined with the use
of QIIME 2 (Bolyen et al., 2019), which is a widely used, open-source bioinformatics platform
designed for analyzing and interpreting amplicon-based metagenomic data, particularly from
marker genes such as the 16S rRNA gene in bacteria and archaea, the ITS region in fungi, or 18S
rRNA in eukaryotes. The pipeline begins with importing raw sequence data, typically demultiplexed FASTQ (FAST Quality) les, which are read into the QIIME 2 environment using specic formats that recognize metadata structures and sequencing formats. Once imported, the reads
undergo a quality control step where sequences are ltered, trimmed, and denoised using tools like
DADA2 or Deblur, which remove noise and chimeras while producing high-resolution amplicon
sequence variants (ASVs). This denoising process is critical because it resolves sequences at the
single-nucleotide level, increasing taxonomic resolution beyond traditional Operational Taxonomic
Unit (OTU) clustering approaches (Callahan et al., 2016).
Following denoising, a feature table is generated that records the frequency of each ASV in
every sample. This table serves as the foundation for downstream analyses. Each ASV is then classied taxonomically by aligning its representative sequence to a reference database such as SILVA
ribosomal RNA gene database (SILVA) (Quast et al., 2013) or Greengenes (DeSantis et al., 2006),
depending on the targeted marker gene. This classication allows researchers to map each ASV
to a known microbial lineage, enabling taxonomic composition analysis across different samples.
QIIME 2 provides robust visualization tools that facilitate the inspection of quality plots, taxonomic
bar charts, and diversity metrics.
Diversity analysis in QIIME 2 includes both alpha diversity, which assesses the richness
and evenness of microbial communities within a sample, and beta diversity, which compares
microbial communities between samples. These analyses can be visualized through principal
coordinates analysis (PCoA) plots and statistically compared using tests such as PERMANOVA.
QIIME 2 also supports phylogenetic analysis by constructing a rooted tree from ASV sequences,
which is often required for certain diversity metrics like Faith’s Phylogenetic Diversity or UniFrac
distances.
Finally, QIIME 2‘s plugin-based architecture allows for the integration of external tools and the
export of results for custom statistical analysis. It supports the use of metadata for grouping samples
by experimental conditions, which is essential for identifying microbial shifts associated with diseases, treatments, or environmental factors. All analyses performed within QIIME 2 are tracked
in a provenance system that ensures reproducibility, allowing researchers to trace each output back
to its inputs and parameters. This transparency, combined with its broad community support and
extensibility, makes QIIME 2 an essential tool for amplicon-based metagenomic analysis in microbiome research, including studies on autoimmune diseases.
A powerful and extensible microbiome bioinformatics platform, QIIME 2 supports end-to-end
processing, beginning with importing raw sequencing data in FASTQ format. The rst crucial
step involves demultiplexing and quality control, where reads are ltered, trimmed, and denoised.
The DADA2 plugin in QIIME 2 models sequencing error proles to recover accurate microbial
sequences, referred to as ASVs, improving taxonomic resolution. These ASVs are then assigned
taxonomy using trained classiers and reference databases such as SILVA or Greengenes. Once taxonomic proles are established, QIIME 2 offers a wide range of analytical and visualization tools
to evaluate alpha diversity (within-sample diversity), beta diversity (between-sample dissimilarity),
and differential abundance of microbial taxa across groups (Figure 9.1).

FIGURE 9.1 Amplicon-based metagenomics data analysis workow.
289 Roles of Bacteria in Autoimmune Diseases
In the context of autoimmune disease research, QIIME 2 has been instrumental in revealing
patterns of microbial dysbiosis and potential microbial triggers. The reproducibility and modularity of QIIME 2 make it particularly suited for multi-site studies and longitudinal monitoring of
microbiome changes in response to treatment, such as dietary interventions or immunosuppressive
therapies. Additionally, its integration with plugins for phylogenetic analysis, machine learning, and
statistical modeling allows researchers to link microbial signatures with clinical phenotypes and
immune proles. Through its transparency, community support, and robust documentation, QIIME
2 continues to be a central tool in advancing our understanding of how microbial communities
shape the immunological landscape in autoimmune disorders.
9.3.1.1 Software Installation for Amplicon Metagenomics
To run this pipeline successfully, several software packages and dependencies must be installed
and properly congured. The primary requirement is QIIME 2, a powerful platform for microbiome analysis that provides the core functionality used throughout the pipeline. QIIME 2 should
be installed within a compatible Conda environment to ensure all dependencies are satised and
version conicts are avoided. It is essential to use a QIIME 2 release that matches the version of
the pre-trained taxonomy classier (such as SILVA), as mismatched versions can result in compatibility errors during taxonomy assignment. In addition to QIIME 2, the environment must include
Python 3 (typically version 3.8 or 3.10, depending on the QIIME 2 release), and Conda itself
should be installed as either Miniconda or Anaconda. The pipeline relies on QIIME 2 plugins
like dada2, demux, feature-table, phylogeny, diversity, composition, taxa,
emperor, and metadata, which are bundled with standard QIIME 2 distributions but must
be properly activated via the environment. The system should also have the Bash shell available
for command execution, as subprocess calls depend on a UNIX-like shell environment. Optional
tools like wget or curl may be useful for downloading external les, including classiers and
example datasets. Furthermore, a reasonably powerful machine is required for practical execution, especially for large datasets; multi-core CPUs, at least 16 GB of RAM, and several tens of
gigabytes of free disk space are recommended to ensure smooth performance during denoising,

290 Bioinformatics of Autoimmune Diseases
phylogenetic tree construction, and diversity analyses. The script assumes a typical directory
structure and requires standard read/write le system access, so permissions must be congured
accordingly.
9.3.1.2 Acquiring Raw Data
For demonstration purposes, we will use NCBI BioProject PRJNA321051, which encompasses a
comprehensive study investigating the alterations in the human gut microbiome associated with
MS. Conducted by researchers at Baylor College of Medicine, the study involved sequencing the
16S rRNA gene from stool samples of MS patients and healthy controls using both Roche 454
GS FLX+ and Illumina MiSeq platforms. The sequencing data are publicly available in the NCBI
Sequence Read Archive (SRA) under BioProject PRJNA321051. For this analysis, we selected ten
healthy samples and ten MS samples, and we will apply an amplicon-based metagenomics pipeline
to explore the microbial proles in these groups (Table 9.1).
To download FASTQ les from the NCBI SRA database, save the following Bash script to a le
named d ata/r aw/dow nlo a d _ fastq.sh:
#!/bin/bash
# Check if input file is provided
if [ "$#" -ne 1 ]; then
echo "Usage: $0 sra_ids.txt"
exit 1
fi
input_file="$1"
# Check if fastq-dump is available
if ! command -v fastq-dump &> /dev/null; then
echo "Error: fastq-dump is not installed or not in PATH."
exit 1
fi
# Download paired-end FASTQ files
while IFS= read -r sra_id; do
if [ -n "$sra_id" ]; then
echo "Downloading $sra_id..."
fastq-dump --split-files "$sra_id"
fi
done < "$input_file"
echo "Download completed."
TABLE 9.1
The Selected Samples from NCBI BioProject PRJNA321051
# Sample ID Condition # Sample ID Condition
1 SRR3501908 Healthy 11 SRR3501917 Sclerosis
2 SRR3501909 Healthy 12 SRR3501918 Sclerosis
3 SRR3501910 Healthy 13 SRR3501919 Sclerosis
4 SRR3501911 Healthy 14 SRR3501920 Sclerosis
5 SRR3501912 Healthy 15 SRR3501921 Sclerosis
6 SRR3501913 Healthy 16 SRR3501922 Sclerosis
7 SRR3501914 Healthy 17 SRR3501923 Sclerosis
8 SRR3501915 Healthy 18 SRR3501924 Sclerosis
9 SRR3501916 Healthy 19 SRR3501926 Sclerosis
10 SRR3501925 Healthy 20 SRR3501927 Sclerosis

291 Roles of Bacteria in Autoimmune Diseases
This Bash script is designed to automate the download of paired-end FASTQ les from the NCBI
SRA using a list of SRA accession identiers (IDs) provided in a text le. When executed, the script
expects exactly one argument: the name of a text le containing SRA IDs, one per line. It begins
by checking whether this input argument has been provided. If not, it outputs a usage message to
inform the user of the correct command format and then exits with an error code, halting further
execution.
Once the input le is conrmed, the script veries that the fastq-dump utility (part of the SRA
Toolkit) is installed and accessible through the system’s PATH. If fastq-dump is not available, the
script noties the user and exits to prevent runtime errors from occurring due to the missing tool.
The core logic of the script is a loop that reads each line from the provided input le. Each line
is assumed to contain a single SRA accession ID. For every non-empty line, the script prints a message indicating which SRA ID is currently being processed, then uses the fastq-dump command
with the --split-files option to download the corresponding paired-end FASTQ les. This
option ensures that the forward and reverse reads are saved separately for downstream bioinformatics analysis. After all SRA IDs have been processed, the script prints a nal message indicating that
the download operation is complete.
Save the SRA run IDs in a text le located at d ata/r aw/sra _ id s.txt, making sure that
each ID is on a separate line. Once the le is ready, navigate to the data/r aw directory and run the
Bash script using the following command in your terminal:
bash download_fastq.sh sra_ids.txt
Be aware that FASTQ les are typically large, and downloading them may take a considerable
amount of time depending on your internet speed and available system memory. Also, make sure
you have sufcient storage space before beginning the download. Once all les have been downloaded, you can compress the FASTQ les using the following command:
gzip *.fastq
9.3.1.3 The Pipeline Inputs
The pipeline program requires specic inputs that must be provided in the correct format; otherwise, it will not complete successfully. These inputs include the following:
9.3.1.3.1 Manifest File
The program is designed to process amplicon-based metagenomic sequencing data using QIIME
2, and it relies on a few essential input les and directories. The most important input is the mani-
fest le located at meta/manifest.csv. This TSV le denes the locations of the paired-end
FASTQ les for each sample, mapping sample IDs to the corresponding le paths. It follows the
PairedEndFastqManifestPhred33V2 format, which is required by QIIME 2 to correctly recognize
and import paired-end sequencing data. The quality encoding used in the FASTQ les must also
conform to Phred33, which is standard for most modern Illumina datasets. You can use the Python
script “create _ ma nifest.py ” to create the manifest le:
import os
import glob
raw_dir = os.path.abspath("data/raw")
manifest_file = "meta/manifest.tsv"
samples = {}
for f in sorted(glob.glob(os.path.join(raw_dir, "*.fastq.gz"))):
basename = os.path.basename(f)
if "_1.fastq.gz" in basename:
sample_id = basename.replace("_1.fastq.gz", "")

292 Bioinformatics of Autoimmune Diseases
samples.setdefault(sample_id, {})["forward"] = f
elif "_2.fastq.gz" in basename:
sample_id = basename.replace("_2.fastq.gz", "")
samples.setdefault(sample_id, {})["reverse"] = f
os.makedirs(os.path.dirname(manifest_file), exist_ok=True)
with open(manifest_file, "w") as out:
out.write(
"sample-id\tforward-absolute-filepath\treverse-absolute-filepath\n"
)
for sample, paths in samples.items():
if "forward" in paths and "reverse" in paths:
out.write(f"{sample}\t{paths['forward']}\t{paths['reverse']}\n")
else:
print(f"Skipping {sample}: missing forward or reverse read.")
print(f"Manifest saved to: {manifest_file}")
The manifest le is structured as follows:
sample-id forward-absolute-filepath reverse-absolute-filepath
SRR12714143 abs_path/SRR12714143_1.fastq.gz abs_path/SRR12714143_2.fastq.gz
SRR12714147 abs_path/SRR12714147_1.fastq.gz abs_path/SRR12714147_2.fastq.gz
SRR12714128 abs_path/SRR12714128_1.fastq.gz abs_path/SRR12714128_2.fastq.gz
SRR12714131 abs_path/SRR12714131_1.fastq.gz abs_path/SRR12714131_2.fastq.gz
SRR12714132 abs_path/SRR12714132_1.fastq.gz abs_path/SRR12714132_2.fastq.gz
SRR12714136 abs_path/SRR12714136_1.fastq.gz abs_path/SRR12714136_2.fastq.gz
SRR12714139 abs_path/SRR12714139_1.fastq.gz abs_path/SRR12714139_2.fastq.gz
SRR12714140 abs_path/SRR12714140_1.fastq.gz abs_path/SRR12714140_2.fastq.gz
SRR12714141 abs_path/SRR12714141_1.fastq.gz abs_path/SRR12714141_2.fastq.gz
SRR12714142 abs_path/SRR12714142_1.fastq.gz abs_path/SRR12714142_2.fastq.gz
9.3.1.3.2 Metadata File
Another key input is the metadata le found at meta/metadata.tsv. This tab-separated le
contains sample-specic metadata, including experimental groupings such as treatment or control conditions. The metadata is critical for many downstream analyses, including diversity group
signicance testing and differential abundance analysis. The metadata le must follow QIIME 2‘s
formatting conventions, with a header row and sample IDs matching those dened in the manifest
le. The metadata le is structured as follows:
#SampleID condition
SRR3501908 healthy
SRR3501909 healthy
SRR3501910 healthy
SRR3501911 healthy
SRR3501912 healthy
SRR3501913 healthy
SRR3501914 healthy
SRR3501915 healthy
SRR3501916 healthy
SRR3501925 healthy
SRR3501917 sclerosis
SRR3501918 sclerosis
SRR3501919 sclerosis
SRR3501920 sclerosis

293 Roles of Bacteria in Autoimmune Diseases
SRR3501921 sclerosis
SRR3501922 sclerosis
SRR3501923 sclerosis
SRR3501924 sclerosis
SRR3501926 sclerosis
SRR3501927 sclerosis
9.3.1.3.3 Pre-Trained Classier
Additionally, the program expects a pre-trained Naive Bayes classier in the form of a QIIME 2
artifact. You can download a pre-trained classier from https://resources.qiime2.org/ or train your
own classier by following the QIIME 2 documentation. This classier is used to assign taxonomy
to the representative sequences derived from the denoising process. It must be trained on a compatible set of sequences and primers that match the target region of the input amplicon data. You can
download the SILVA 138 99% pre-trained Naive Bayes classier for QIIME 2 directly from the
ofcial QIIME 2 data repository. However, the classier must be trained using the same version of
scikit-learn that is installed on your machine.
9.3.1.3.4 FASTQ Files
The FASTQ, downloaded in a previous step, referenced in the manifest, the sample metadata, and
the classier artifact collectively serve as the core inputs required to run the pipeline. These les
need to be present and correctly formatted for the analysis to proceed without errors. All output
generated by the program is stored in the results directory, which is automatically created if it
does not already exist.
The directory tree structure, including input les and scripts, is shown in Figure 9.2.
9.3.1.4 QIIME 2 Artifacts and Visualizations
In QIIME 2, artifacts and visualizations are central components of the data analysis framework,
each serving a distinct but complementary role. Artifacts, represented by les with the .q z a exten-
sion, are data containers that store raw or processed data along with essential metadata and provenance information. These artifacts are not simple les but rather structured archives that include
the data, a record of how it was generated, and the software environment used, ensuring full traceability and reproducibility. For example, an imported sequence dataset, a denoised feature table, or a
phylogenetic tree are all stored as artifacts, encapsulating the actual data output of various pipeline
stages.
In contrast, visualizations are represented by les with the .qz v extension and are designed
specically to present results to the user in a human-readable, interactive form. These visualizations
include tables, summaries, plots, and statistics generated from artifacts. Unlike artifacts, which are
primarily intended for computational use in downstream steps, visualizations are meant for interpretation, inspection, and decision-making. For instance, after generating a feature table artifact,
one might create a corresponding visualization to explore sequencing depth per sample, which
assists in determining whether to lter out low-abundance samples.
The key distinction between artifacts and visualizations lies in their purpose and usage. Artifacts
are data carriers that can be passed into other QIIME 2 commands, acting as inputs for further
processing. Visualizations, however, are terminal products of analysis steps and cannot be reused
as inputs for other commands. They are instead opened using the QIIME 2 viewer or within a
QIIME2 environment to facilitate data exploration and result validation. This separation enforces
clarity in workow structure: artifacts build the computational history, while visualizations support
the interpretative process.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
