Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5423_Библиотеки_им_академика_М_И_Перельмана
.pdf
13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 325
https://t.me/medicina_free
that AAV integration is a low risk [49–53]. The summary basis of approval for
Luxturna indicates that the manufacturer of this GTx justified not conducting
genotoxicity studies on the basis that the current scientific literature reported a
low integration frequency of AAV vectors and delivery to a post mitotic cell population[54]. For Glybera, after extensive evaluation of integration in animals, the
EMA concluded: “Data do not substantiate a concern for tumorigenicity.
Theoretically, the product could integrate and cause a tumor, however no further
animal testing or experiments can usefully address these concerns”[49]. Given
the remaining uncertainty on the relevance of AAV integration for human risk,
the assessment of integration events in nonclinical and clinical studies will be a
discussion point with regulatory authorities during the development and post‐
approval monitoring of AAV GTx.
13.2.1 Factors to Consider in the Design of Nonclinical Studies
Evaluating AAV Integration
Drug product that is evaluated in pivotal nonclinical studies to support administration to humans should be comparable to that administered to subjects in clinical trials. One of the parameters assessed for comparability that is relevant to
characterizing AAV integration is the DNA content and form of the administered
vector. Where studied, integrated rAAV DNA is often found to be rearranged, and
this is also true of rAAV DNA in vector particles prior to administration[55, 56].
While sequencing technologies for characterizing different forms of DNA that are
packaged in AAV vector particles are in their infancy, the ability to characterize
these forms may become an important criterion in assessing comparability of
AAV preparations[55–57]. Rearranged rAAV genomes and production plasmid
sequences have been found to be integrated in the genome of human hepatocytes
transduced exvivo or invivo[19]. Consequently, the molecular characterization
of AAV vector preparations may be important in understanding the effects on
genome integration and functional implications of an AAV’s integration profile.
The need to assess DNA integration may vary between regulatory agencies in
different geographies and may be requested by regulatory authorities at the latter
stages of development (author’s experience). Consequently, this topic should be a
discussion point with regulatory agencies. Considerations should be given to
collecting appropriate samples from toxicology studies and/or retaining DNA
extracted from tissues for evaluating vector biodistribution, so that DNA integration can be assessed, at the request of a regulatory authority, as a program develops over time or if an observation of concern arises in a preclinical study[45, 58].
The proactive collection of tissues will avoid the need to repeat studies to address
this issue, thus minimizing animal use.

13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
326
13.2.2 Methods for rAAV Integration Analysis
Integration site analyses are based on the identification of host cell genomic
sequences flanking the integrated vector genome in order to determine the position of the insertion event. The approaches for integration site screening traditionally consisted of the enrichment for vector‐genome junctions present within a
sample to enable the subsequent identification process. In this context, rAAV vectors with a predominance of vector genomes persisting episomally, that will be
co‐purified during the junction retrieval process, together with rare integration
frequencies due to the lack of integrase activity, require highly sensitive approaches
to detect infrequent integration events.
rAAV integration has been predominantly evaluated by ligation‐mediated (LM‐)
PCR [59] or linear amplification‐mediated (LAM‐)PCR [60] approaches using
restriction endonucleases. In essence, LAM‐PCR consists of an initial linear
amplification performed with a biotinylated primer located at the vector end, followed by purification utilizing streptavidin magnetic beads. Subsequent steps are
performed on the DNA immobilized on the bead surface in order to cut the flanking genomic sequence with a restriction endonuclease enabling the ligation of a
linker cassette of known sequence. This provides the binding site for a reverse
primer that together with a forward oligonucleotide binding to the vector is used
in the exponential amplification steps. LAM‐PCR represented the state‐of‐the‐art
for vector integration site retrieval for almost two decades due to its superior sensitivity. Nonetheless, LM‐PCR approaches have been widely used and their general principle was the digestion of the genomic DNA with a restriction
endonuclease followed by a linker cassette ligation and exponential amplification.
Initially, both approaches were followed by Sanger sequencing of the amplicons,
sometimes including an interim cloning step[15]. However, all methods are now
combined with next‐generation sequencing, allowing for a time‐ and cost‐
effective high‐throughput screening of vector insertion sites[13].
The limitations associated with the usage of restriction enzymes were recognized early, since integration site retrieval was limited to those events having a
suitable restriction site nearby, and initially approached by the usage of multiple
restriction endonucleases [61, 62]. This, in addition to biases introduced by
deep sequencing of amplicons of divergent sizes, has encouraged the usage of
approaches utilizing random DNA shearing[63, 64]. Accordingly, a new generation of LM‐PCR‐based methods utilizing DNA sonication has emerged and been
successfully used to retrieve AAV vector insertion sites[27].
In addition to the determination of the integration site position, the quantification of the number of integration events constitutes a key part of integration site
analyses. Different strategies are being utilized in order to correct insertion
site frequencies for PCR‐associated as well as other technical biases [65].

13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 327
https://t.me/medicina_free
Although these mainly occur at the data processing stage as it will be detailed
later, laboratory protocols contemplating the addition of molecular barcodes
allow for broader dynamic ranges by overcoming eventual saturation phenomena
due to the limited availability of diverse shearing sites in scenarios of clonal dominance. These molecular barcodes consist of unique molecular identifiers introduced within the linker cassettes ligated to the DNA fragments prior to the
exponential amplification steps[66, 67].
PCR‐based approaches have allowed the identification of AAV vector IS and
clonal expansions. However, their dependence on the presence of the selected
primer binding site limits the ability to detect all components of the vector that
may integrate into host cell DNA. To address integrated fragments undergoing
internal breakage and rearrangements, multiplexed LAM‐PCR approaches were
developed also allowing for the retrieval of insertion events arising from internal
vector regions[3]. Improved protocols for multiplex PCR‐based techniques are
still needed to overcome the LAM‐PCR limitations detailed above. Recently, target enrichment (TE) approaches offer a promising alternative [68]. Sequence
adaptors are ligated to fragmented genomic DNA, and the enrichment is performed by hybridization with a set of RNA or DNA baits complementary to the
entire vector sequence. This approach pulls down any DNA molecule‐bearing
vector fragments, including both episomal and integrated genomes. Considering
the lower number of amplification cycles and the relatively small size of the targeted region (the AAV genome), TE approaches are considered to present a lower
sensitivity when compared to PCR‐based methods, although systematic comparisons are still needed. Nonetheless, these approaches allow to capture and sequence
internal vector fragments, thus providing complete information on persisting vector genomes. Indeed, a study coupling a capture approach with long‐read sequencing has showed the presence of integrated concatemeric sequences and a high
degree of rearrangements within the vector genomes revealing the presence of
vector structures that may have been missed by PCR‐based technologies[19].
Besides the vector‐targeted approaches, whole genome sequencing (WGS) has
also been utilized to retrieve vector IS[69]. The procedure is analog to the one
used in classical WGS studies and the difference relies on the sequencing data
analyses. Despite its unbiased nature, this approach presents a low sensitivity,
given by the absence of enrichment steps, thus limiting its usage in samples with
rare and low‐frequency integration events. Despite the advances in sequencing
technologies, performing integration site analyses through WGS still remains
time and cost‐intensive. Nevertheless, it still constitutes a valid orthogonal method
when eventual clonal expansions are observed.
A consensus on the precise range of AAV integration frequency detected by different approaches is still missing. The diversity of approaches utilized and the

13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
328
absence of reference materials, impose limitations to systematic comparisons. For
integrating vectors (e.g. lentivirus vectors), reference materials have been generated by cloning cells with a limited and known number of insertion events.
However, such process would be challenging for AAV vectors considering their
passive integration nature and the high diversity of vector designs, with only the
ITRs being almost universal throughout different vector designs. Therefore, the
definition of method performance parameters and novel strategies enabling inter‐
study comparison are still required.
13.2.3 AAV Data Analysis Methods
In the past decade, all the approaches evaluating rAAV genome interaction with
the host DNA employed next‐generation sequencing technologies. Nowadays,
standardized procedures automatically and cost‐effectively reveal thousands of
IS in a few days. Figure13.1 presents a simplified flowchart that generalizes the
stages required to generate the common safety artifacts concerning the scientific
and regulatory GTx community.
For rAAV, safety studies incorporating an assessment of DNA integration,
sequencing technologies, and bioinformatics have become an important consideration. Low‐throughput technologies are frequently used in studies with low
numbers of samples to permit manual analysis[15, 26, 70]. The analysis does not
necessitate specialized bioinformatics and can be performed using online alignment tools [71, 72] and visually inspecting the results. This approach was frequently used in older studies with low numbers of samples[15, 26, 70].
Advancements in the wet lab, sequencing, and data management technologies
require developing progressively enhanced data analysis procedures. For example, the different chemistries employed for library preparation require distinct
data processing strategies. This section combines the library preparation methods
into three main categories: PCR, TE or viral capture, and WGS (Figure13.2). All
three, are described in Section 13.2.2, yield datasets requiring specialized data
analysis workflows to enable integration site retrieval and quantification.
WGS and whole transcriptome sequencing (WTS) libraries are utilized for analysis like germinal/somatic variants detection and copy number variation, where
the IS analysis is typically an accessory safety assessment primarily used in large
diagnostic projects [69]. This data type does not require specific preprocessing
steps, such as trimming‐specific sequences or combining reads, and the alignment stage directly processes the reads [73, 74]. Some tools, such as
VirusFinder[73], Virus‐clip [74], and VirusBreakend [69], are explicitly developed for WGS data. They are not AAV‐specialized tools and do not make any
assumptions regarding the reads structure or the integration process. To achieve
the high sensitivity required for unspecialized methods, the tools perform a

Start
https://t.me/medicina_free
Sequencing
Generic
preprocessing
Manual analysis?
Figure13.1 Common analysis steps in AAV GTx safety analysis. In light blue are the steps that are common to all
the analyses. The three main categories are: impurities detection (red), AAV genome rearrangement (green), and
integration site analysis (blue). White boxes refer to the analyses that detect integrations inspecting the reads
manually. “Other libraries” represents emerging protocols and technologies as long-range reads. “Other Analyses”
represent analyses that are not discussed in this chapter as the nuclease on/off-target activity detection.
No
Target
enrichment
library?
No
WGS library?
No
PCR library?
No
Yes
Yes
Yes
Yes
Other libraries
Demultiplexing?
No
Sample sorting
and
trimming
Visual inspection
Alignment
Yes
Demultiplexing
Alignment
Evaluate DNA content
of AAV particles
No
Detects AAV
rearrangements?
No
Detect integration
Sites?
No
Other analyses
Yes
Yes
Yes
Extract IS clonality
Detect impurities
Detect rearranged
AAV genomes
Integration sites table
Stop
Integration
site list
Impurity table
Impurity coverage
Episomal AAV
Rearranged AAV
Detect integration
hotspots
Detect genotoxic
integrations

Methods
https://t.me/medicina_free
Whole Genome Sequencing
– Used in large diagnostic projects
– Integration Site Analysis
as an accessory analysis
– Extermely low sensitivity
Figure13.2 Library preparation methods for vector integration site analysis yielding sequencing datasets
requiring specific data analysis workflows.
Target Enrichment (TE, TES)
– Detection of integration sites involving
internal regions of the AAV genome
– Lower sensitivity than PCR methods
PCR
– High sensitivity
– Blind to integrations involving
internal part of the AAV genome
Ligation Mediated (LM-)PCR
– Usage of Restriction Enzymes (abandoned)
– Sonication
Linear Amplification Mediated (LAM-)PCR
– Restriction Enzymes (abandoned)

13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 331
https://t.me/medicina_free
preselection step where the reads are aligned on the vector genome. In this way,
any vector signal coming from the data can be further elaborated for extracting IS.
TE is a method developed to enrich the concentration of fragments that contain
a preselected portion of the vector. Like WGS libraries, the read structure is simple, and the strategies used for WGS analysis are still usable. Given that the sample is enriched in the AAV genome fragments, some pipelines[19, 20, 39, 75]
analyze TE data by implementing a strategy that uses this information aligning
the reads directly with a hybrid genome that combines the organism assembly and
the vector sequence as an additional chromosome. TE was used to study AAV
integration[20] and to characterize HCC development after AAV infection[39].
rAAV integrations and the genomic rearrangements in transduced human hepatocytes, expanded in a mouse model, were recently discovered by employing TE in
combination with PacBio long‐read sequencing[19].
The PCR methods are probably the most widely used to assess rAAV integration
as they are considered more sensitive than WGS and TE [3, 13, 18, 27, 76–78].
Medium‐throughput sequencers, such as MiSeq, are sufficient for analyzing several samples in one run. Mixing the samples in a single library requires advanced
barcoding strategies for the sample fragments. As an example, a unique combination of two barcodes (short sequences) are ligated at the 5′ and 3′of the sample
fragment. This combination is recognized on the forward and reverse read by the
sorting tool and is used to group the reads in the corresponding sample. An essential preprocessing stage requires assigning (sorting) the reads to the original sample. The AAV genome sequence used as a template for the PCR primer is employed
as an anchor point for preselecting the vector positive reads. Typically, this
sequence is removed in the trimming phase.
13.2.3.1 AAV Primary Analysis
An alignment tool, independent from the algorithm implemented for finding hits,
returns the coordinates of the alignment along with the information helpful in
evaluating the match. The most straightforward score is the identity and the percentage of identical bases in the alignment. It is customary to consider valid a
match when the identity is more significant than a predetermined threshold (usually 95%)[3, 13, 19, 27, 76, 77]. Moreover, under the simple hypothesis that all the
bases are equiprobable, we expect that any random fragment of 16 bases should
appear about once into a random string of the dimension of the human genome.
For this reason in many approaches, the aligned region must be longer than a
minimum[75] to reduce the risk of false‐positive hits. Some alignment tools present a supplementary score that combines different quality metrics, simplifying
the match quality estimation. For example, BLAST returns a score that estimates
the number of times an alignment with the same characteristics is expected by
chance (E‐value). The information about the alignment is stored in large data files

13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
332
and is the primary source for answering one or more questions addressed by the
study. A de facto standard is the Sequence Alignment Map format (SAM) and most
of the modern alignment tools return the alignment in this or in the corresponding binary format (BAM).
13.2.3.2 Impurity Analysis
One area of emerging interest is the evaluation of rearranged vector DNA and
DNA‐process‐related contaminates that are packaged in the AAV viral particle.
The bioinformatics methods are in active development. However, the fundamental analysis can be reduced to detect potential contaminants or misarranged AAV
encapsidated during the virus production. From the alignment file, where all the
potential DNA contaminants, derived from the production process, are used as
the target genome, the number of reads and the position of the contaminants are
returned and presented as count table and coverage figures[57, 79, 80].
13.2.3.3 AAV Genome Rearrangements
In postmitotic cells, rAAV mainly persists as an extrachromosomal element able
to stably express the transgene for a long period of time[6, 81–83]. The inter and
intra‐molecular recombination of the ITRs produces three primary types of circles: 5′ITR‐3′ITR (head to tail), 5′ITR‐5′ITR (head to head), and 3′ITR‐3′ITR (tail
to tail). While in dividing cells, the mechanism that ensures the persistence of the
rAAV DNA seems to be the integration into the host DNA. Assessing the amount
of episomal circular/concatemeric forms of the vector and distinguishing the elements from the integrated vector genome is a challenge for the analysis tools.
Searching for reads that map discontinuously on the AAV genome is the primary
strategy for determining the frequency of the episomal forms in the samples
[22, 77] assuming that rearrangement among vector is an indication of an extrachromosomal AVV genome. It has been shown that rearranged vectors can integrate into the host genome[3, 19, 84], and that numerous rearrangement events
can lead to complicated structures of concatenated AAV fragments. For this reason, it is not possible to univocally distinguish between integrated/nonintegrated
AAV DNA without more specialized methods that make use of long‐read sequencing[19, 85] and new bioinformatics tools[3, 19].
13.2.3.4 Integration Site Analysis
The most common question answered by advanced bioinformatics methods is the
location of IS. Whilst only a few reads are inspected visually in the manual analysis, in
all the modern analysis tools, a multitude of criteria are introduced to select, evaluate,
and report integration events. Here, we will discuss the most relevant standard analysis.
It is not uncommon that a single DNA fragment aligns with multiple genomic
regions, and these analyses must be handled with caution to minimize signal loss.

13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 333
https://t.me/medicina_free
For example, integrations into a repetitive area of the host genome cannot be
identified uniquely and represents potential unresolved IS. Different strategies
can be adopted to reduce the impact of unresolved IS. A straightforward but not
always implementable experimental solution is to increase the reads’ length to
decrease the number of unresolved IS. Alternatively, a computational strategy
that increases the reads’ dimension in silico can be implemented as established in
gamma‐TRIS aligner [86], a tool for detecting retroviral integrations. More
straightforward but less efficient approaches discard the reads that show alignments with the same score on multiple locations[13] or when the score between
the best match and the second best is more significant than a threshold [75].
Additionally, the mapping of IS into low complexity and repetitive regions
depends on the alignment tool sensitivity. For example, by default, BLAST masks
low complexity regions from the search[87], making it practically blind in a not
negligible portion of the human genome, which results in a loss of sensitivity
[3, 13, 15, 26, 77].
Besides host genome mapping, identification of the junction, i.e. the region
where the AAV genome and the reference are fused can be challenging. Ideally,
when the AAV sequence ends, the genomic begins so that the junction region is
clearly defined. While preparing the sequencing libraries, different operation on
the samples introduces noise in the form of (1) minor variations, (2) biased amplification rate, and (3) invitro recombination of the DNA fragments. The last one,
producing biological artifacts[88], affects the integration site detection. For this
reason, the analysis tool should control the formation of the chimeric reads by
requiring that the vector‐genome junction sequences are well‐formed.
In WGS and TE methods, the reads aligned to the vector are parsed to validate
the alignment structure. The Concise Idiosyncratic Gapped Alignment Report
(CIGAR) string in the Sequence Alignment/Map (SAM) format is used to authenticate the alignment structure[19, 75]. A less stringent requirement sometimes
adopted is the presence in the read pairs of at least one read aligned to the AAV
genome[20, 39]. Analysis tools for PCR methods use filtering strategies that may
involve a) a limited number of unaligned bases between the vector and the
genomic regions[27, 75, 77] and the absence of more than one vector‐genome
junction per fragment[3, 22]. This last strategy decreases the noise due to artifacts, such as chimeras generated during the library preparation and sequencing
but makes the analysis unable to detect rearranged AAV genome integration (see
Section13.2.3.3).
13.2.3.5 Clonality Analysis
Due to the stochastic nature of the integration process, a single integration site
marks univocally a single clone. The number of IS present at a particular time point
in a given sample is thus a representation of the clonal configuration at that moment.

13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
334
If some IS are found more frequently than others, then it can be inferred that
those IS may be a sign of clonal expansion of the cell harboring that integration.
The assessment that addresses the reconstruction and comparison of the clonality
of a set of samples is composed of two procedures: (1) the rebuilding of the clonality of any single IS detected within a sample (IS clonality) and (2) the calculation
of the clonality of the sample (clonal diversity).
The first approach associates with each IS a number representing the total number of times each clone is present in the sample. The straightforward derivation of
IS clonality is by assessing the number of times a particular IS is detected in the
given sample[3, 18, 89]. This approach, however, may be unable to distinguish
between reads from different clones or PCR amplification. Other counting strategies can reduce this issue by exploiting experimental evidence and adding a few
assumptions[19, 20, 27, 39]. For example, libraries that originate from random
DNA shearing (e.g. by sonication) are less prone to contain identical fragments
from the same clones, thus reducing the count bias introduced by the PCR[63].
On the other hand, this approach requires adjustment when the sequencing depth
increases to a level where the sonicated fragments are saturated[63].
The clonal diversity is calculated from the IS clonality employing methods used
in ecology for measuring biodiversity. Its main aim is to compare the clonal repertoire in the same individual and, as an example, identify the surge of a clone,
intercepting potential adverse events caused by insertional mutagenesis as soon
as possible. The simplest and frequently used measure of diversity is clonal abundance[27, 40, 78], which determines the number of times a clone is present in the
sample compared to the total number of clones. When the clonal abundance of a
specific IS increases over a defined threshold, a red flag is raised, triggering more
detailed investigations. With this regard, several diversity indexes are employed,
such as the Shannon index[90], the polyclonal monoclonal diversity (PMD)[91],
and the shape‐constrained splines (SCS) method [92] in GTx for tracking and
comparing clonal diversity across samples and studies.
13.2.3.6 Genotoxic Integrations
The emergence of a clone that expands during longitudinal sampling triggers a
major concern for an insertional mutagenesis event[93]. It is a common practice
to assign each IS to the closest gene within a specific range[1, 13, 15, 18, 20, 27].
In this way, the analysis can focus only on the integrations close to a subset of
cancer‐associated genes. This may seem a straightforward process, however,
selecting a comprehensive and meaningful set of detrimental genes is challenging. Aside from gamma‐retroviral studies that established a specific (limited)
group of genes as a source of potential insertional mutagenesis[94, 95], for more
comprehensive analysis different, not overlapping collections of cancer genes, are
curated by multiple databases[96–98]. An additional layer of complexity in the
Соседние файлы в папке Библиотека им академика М.И. Перельмана
