Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5943_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 325
https://t.me/medicina_free
that AAV integration is a low risk [49–53]. The summary basis of approval for Luxturna indicates that the manufacturer of this GTx justified not conducting genotoxicity studies on the basis that the current scientific literature reported a low integration frequency of AAV vectors and delivery to a post mitotic cell popu­lation[54]. For Glybera, after extensive evaluation of integration in animals, the EMA concluded: “Data do not substantiate a concern for tumorigenicity. Theoretically, the product could integrate and cause a tumor, however no further animal testing or experiments can usefully address these concerns”[49]. Given the remaining uncertainty on the relevance of AAV integration for human risk, the assessment of integration events in nonclinical and clinical studies will be a discussion point with regulatory authorities during the development and post‐ approval monitoring of AAV GTx.
13.2.1 Factors to Consider in the Design of Nonclinical Studies Evaluating AAV Integration
Drug product that is evaluated in pivotal nonclinical studies to support adminis­tration to humans should be comparable to that administered to subjects in clini­cal trials. One of the parameters assessed for comparability that is relevant to characterizing AAV integration is the DNA content and form of the administered vector. Where studied, integrated rAAV DNA is often found to be rearranged, and this is also true of rAAV DNA in vector particles prior to administration[55, 56]. While sequencing technologies for characterizing different forms of DNA that are packaged in AAV vector particles are in their infancy, the ability to characterize these forms may become an important criterion in assessing comparability of AAV preparations[55–57]. Rearranged rAAV genomes and production plasmid sequences have been found to be integrated in the genome of human hepatocytes transduced exvivo or invivo[19]. Consequently, the molecular characterization of AAV vector preparations may be important in understanding the effects on genome integration and functional implications of an AAV’s integration profile.
The need to assess DNA integration may vary between regulatory agencies in different geographies and may be requested by regulatory authorities at the latter stages of development (author’s experience). Consequently, this topic should be a discussion point with regulatory agencies. Considerations should be given to collecting appropriate samples from toxicology studies and/or retaining DNA extracted from tissues for evaluating vector biodistribution, so that DNA integra­tion can be assessed, at the request of a regulatory authority, as a program devel­ops over time or if an observation of concern arises in a preclinical study[45, 58]. The proactive collection of tissues will avoid the need to repeat studies to address this issue, thus minimizing animal use.
13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
326
13.2.2 Methods for rAAV Integration Analysis
Integration site analyses are based on the identification of host cell genomic sequences flanking the integrated vector genome in order to determine the posi­tion of the insertion event. The approaches for integration site screening tradition­ally consisted of the enrichment for vector‐genome junctions present within a sample to enable the subsequent identification process. In this context, rAAV vec­tors with a predominance of vector genomes persisting episomally, that will be co‐purified during the junction retrieval process, together with rare integration frequencies due to the lack of integrase activity, require highly sensitive approaches to detect infrequent integration events.
rAAV integration has been predominantly evaluated by ligation‐mediated (LM‐) PCR [59] or linear amplification‐mediated (LAM‐)PCR [60] approaches using restriction endonucleases. In essence, LAM‐PCR consists of an initial linear amplification performed with a biotinylated primer located at the vector end, fol­lowed by purification utilizing streptavidin magnetic beads. Subsequent steps are performed on the DNA immobilized on the bead surface in order to cut the flank­ing genomic sequence with a restriction endonuclease enabling the ligation of a linker cassette of known sequence. This provides the binding site for a reverse primer that together with a forward oligonucleotide binding to the vector is used in the exponential amplification steps. LAM‐PCR represented the state‐of‐the‐art for vector integration site retrieval for almost two decades due to its superior sen­sitivity. Nonetheless, LM‐PCR approaches have been widely used and their gen­eral principle was the digestion of the genomic DNA with a restriction endonuclease followed by a linker cassette ligation and exponential amplification. Initially, both approaches were followed by Sanger sequencing of the amplicons, sometimes including an interim cloning step[15]. However, all methods are now combined with next‐generation sequencing, allowing for a time‐ and cost‐ effective high‐throughput screening of vector insertion sites[13].
The limitations associated with the usage of restriction enzymes were recog­nized early, since integration site retrieval was limited to those events having a suitable restriction site nearby, and initially approached by the usage of multiple restriction endonucleases [61, 62]. This, in addition to biases introduced by deep sequencing of amplicons of divergent sizes, has encouraged the usage of approaches utilizing random DNA shearing[63, 64]. Accordingly, a new genera­tion of LM‐PCR‐based methods utilizing DNA sonication has emerged and been successfully used to retrieve AAV vector insertion sites[27].
In addition to the determination of the integration site position, the quantifica­tion of the number of integration events constitutes a key part of integration site analyses. Different strategies are being utilized in order to correct insertion site frequencies for PCR‐associated as well as other technical biases [65].
13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 327
https://t.me/medicina_free
Although these mainly occur at the data processing stage as it will be detailed later, laboratory protocols contemplating the addition of molecular barcodes allow for broader dynamic ranges by overcoming eventual saturation phenomena due to the limited availability of diverse shearing sites in scenarios of clonal domi­nance. These molecular barcodes consist of unique molecular identifiers intro­duced within the linker cassettes ligated to the DNA fragments prior to the exponential amplification steps[66, 67].
PCR‐based approaches have allowed the identification of AAV vector IS and clonal expansions. However, their dependence on the presence of the selected primer binding site limits the ability to detect all components of the vector that may integrate into host cell DNA. To address integrated fragments undergoing internal breakage and rearrangements, multiplexed LAM‐PCR approaches were developed also allowing for the retrieval of insertion events arising from internal vector regions[3]. Improved protocols for multiplex PCR‐based techniques are still needed to overcome the LAM‐PCR limitations detailed above. Recently, tar­get enrichment (TE) approaches offer a promising alternative [68]. Sequence adaptors are ligated to fragmented genomic DNA, and the enrichment is per­formed by hybridization with a set of RNA or DNA baits complementary to the entire vector sequence. This approach pulls down any DNA molecule‐bearing vector fragments, including both episomal and integrated genomes. Considering the lower number of amplification cycles and the relatively small size of the tar­geted region (the AAV genome), TE approaches are considered to present a lower sensitivity when compared to PCR‐based methods, although systematic compari­sons are still needed. Nonetheless, these approaches allow to capture and sequence internal vector fragments, thus providing complete information on persisting vec­tor genomes. Indeed, a study coupling a capture approach with long‐read sequenc­ing has showed the presence of integrated concatemeric sequences and a high degree of rearrangements within the vector genomes revealing the presence of vector structures that may have been missed by PCR‐based technologies[19].
Besides the vector‐targeted approaches, whole genome sequencing (WGS) has also been utilized to retrieve vector IS[69]. The procedure is analog to the one used in classical WGS studies and the difference relies on the sequencing data analyses. Despite its unbiased nature, this approach presents a low sensitivity, given by the absence of enrichment steps, thus limiting its usage in samples with rare and low‐frequency integration events. Despite the advances in sequencing technologies, performing integration site analyses through WGS still remains time and cost‐intensive. Nevertheless, it still constitutes a valid orthogonal method when eventual clonal expansions are observed.
A consensus on the precise range of AAV integration frequency detected by dif­ferent approaches is still missing. The diversity of approaches utilized and the
13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
328
absence of reference materials, impose limitations to systematic comparisons. For integrating vectors (e.g. lentivirus vectors), reference materials have been gener­ated by cloning cells with a limited and known number of insertion events. However, such process would be challenging for AAV vectors considering their passive integration nature and the high diversity of vector designs, with only the ITRs being almost universal throughout different vector designs. Therefore, the definition of method performance parameters and novel strategies enabling inter‐ study comparison are still required.
13.2.3 AAV Data Analysis Methods
In the past decade, all the approaches evaluating rAAV genome interaction with the host DNA employed next‐generation sequencing technologies. Nowadays, standardized procedures automatically and cost‐effectively reveal thousands of IS in a few days. Figure13.1 presents a simplified flowchart that generalizes the stages required to generate the common safety artifacts concerning the scientific and regulatory GTx community.
For rAAV, safety studies incorporating an assessment of DNA integration, sequencing technologies, and bioinformatics have become an important consid­eration. Low‐throughput technologies are frequently used in studies with low numbers of samples to permit manual analysis[15, 26, 70]. The analysis does not necessitate specialized bioinformatics and can be performed using online align­ment tools [71, 72] and visually inspecting the results. This approach was fre­quently used in older studies with low numbers of samples[15, 26, 70].
Advancements in the wet lab, sequencing, and data management technologies require developing progressively enhanced data analysis procedures. For exam­ple, the different chemistries employed for library preparation require distinct data processing strategies. This section combines the library preparation methods into three main categories: PCR, TE or viral capture, and WGS (Figure13.2). All three, are described in Section 13.2.2, yield datasets requiring specialized data analysis workflows to enable integration site retrieval and quantification.
WGS and whole transcriptome sequencing (WTS) libraries are utilized for anal­ysis like germinal/somatic variants detection and copy number variation, where the IS analysis is typically an accessory safety assessment primarily used in large diagnostic projects [69]. This data type does not require specific preprocessing steps, such as trimming‐specific sequences or combining reads, and the align­ment stage directly processes the reads [73, 74]. Some tools, such as VirusFinder[73], Virus‐clip [74], and VirusBreakend [69], are explicitly devel­oped for WGS data. They are not AAV‐specialized tools and do not make any assumptions regarding the reads structure or the integration process. To achieve the high sensitivity required for unspecialized methods, the tools perform a
Start
https://t.me/medicina_free
Sequencing
Generic
preprocessing
Manual analysis?
Figure13.1  Common analysis steps in AAV GTx safety analysis. In light blue are the steps that are common to all
the analyses. The three main categories are: impurities detection (red), AAV genome rearrangement (green), and integration site analysis (blue). White boxes refer to the analyses that detect integrations inspecting the reads manually. “Other libraries” represents emerging protocols and technologies as long-range reads. “Other Analyses” represent analyses that are not discussed in this chapter as the nuclease on/off-target activity detection.
No
Target
enrichment
library?
No
WGS library?
No
PCR library?
No
Yes
Yes
Yes
Yes
Other libraries
Demultiplexing?
No
Sample sorting
and
trimming
Visual inspection
Alignment
Yes
Demultiplexing
Alignment
Evaluate DNA content
of AAV particles
No
Detects AAV
rearrangements?
No
Detect integration
Sites?
No
Other analyses
Yes
Yes
Yes
Extract IS clonality
Detect impurities
Detect rearranged
AAV genomes
Integration sites table
Stop
Integration
site list
Impurity table
Impurity coverage
Episomal AAV
Rearranged AAV
Detect integration
hotspots
Detect genotoxic
integrations
Methods
https://t.me/medicina_free
Whole Genome Sequencing – Used in large diagnostic projects – Integration Site Analysis as an accessory analysis – Extermely low sensitivity
Figure13.2  Library preparation methods for vector integration site analysis yielding sequencing datasets
requiring specific data analysis workflows.
Target Enrichment (TE, TES) – Detection of integration sites involving internal regions of the AAV genome – Lower sensitivity than PCR methods
PCR – High sensitivity – Blind to integrations involving internal part of the AAV genome
Ligation Mediated (LM-)PCR – Usage of Restriction Enzymes (abandoned) – Sonication
Linear Amplification Mediated (LAM-)PCR – Restriction Enzymes (abandoned)
13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 331
https://t.me/medicina_free
preselection step where the reads are aligned on the vector genome. In this way, any vector signal coming from the data can be further elaborated for extracting IS.
TE is a method developed to enrich the concentration of fragments that contain a preselected portion of the vector. Like WGS libraries, the read structure is sim­ple, and the strategies used for WGS analysis are still usable. Given that the sam­ple is enriched in the AAV genome fragments, some pipelines[19, 20, 39, 75] analyze TE data by implementing a strategy that uses this information aligning the reads directly with a hybrid genome that combines the organism assembly and the vector sequence as an additional chromosome. TE was used to study AAV integration[20] and to characterize HCC development after AAV infection[39]. rAAV integrations and the genomic rearrangements in transduced human hepat­ocytes, expanded in a mouse model, were recently discovered by employing TE in combination with PacBio long‐read sequencing[19].
The PCR methods are probably the most widely used to assess rAAV integration as they are considered more sensitive than WGS and TE [3, 13, 18, 27, 76–78]. Medium‐throughput sequencers, such as MiSeq, are sufficient for analyzing sev­eral samples in one run. Mixing the samples in a single library requires advanced barcoding strategies for the sample fragments. As an example, a unique combina­tion of two barcodes (short sequences) are ligated at the 5′ and 3′of the sample fragment. This combination is recognized on the forward and reverse read by the sorting tool and is used to group the reads in the corresponding sample. An essen­tial preprocessing stage requires assigning (sorting) the reads to the original sam­ple. The AAV genome sequence used as a template for the PCR primer is employed as an anchor point for preselecting the vector positive reads. Typically, this sequence is removed in the trimming phase.
13.2.3.1 AAV Primary Analysis
An alignment tool, independent from the algorithm implemented for finding hits, returns the coordinates of the alignment along with the information helpful in evaluating the match. The most straightforward score is the identity and the per­centage of identical bases in the alignment. It is customary to consider valid a match when the identity is more significant than a predetermined threshold (usu­ally 95%)[3, 13, 19, 27, 76, 77]. Moreover, under the simple hypothesis that all the bases are equiprobable, we expect that any random fragment of 16 bases should appear about once into a random string of the dimension of the human genome. For this reason in many approaches, the aligned region must be longer than a minimum[75] to reduce the risk of false‐positive hits. Some alignment tools pre­sent a supplementary score that combines different quality metrics, simplifying the match quality estimation. For example, BLAST returns a score that estimates the number of times an alignment with the same characteristics is expected by chance (E‐value). The information about the alignment is stored in large data files
13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
332
and is the primary source for answering one or more questions addressed by the study. A de facto standard is the Sequence Alignment Map format (SAM) and most of the modern alignment tools return the alignment in this or in the correspond­ing binary format (BAM).
13.2.3.2 Impurity Analysis
One area of emerging interest is the evaluation of rearranged vector DNA and DNA‐process‐related contaminates that are packaged in the AAV viral particle. The bioinformatics methods are in active development. However, the fundamen­tal analysis can be reduced to detect potential contaminants or misarranged AAV encapsidated during the virus production. From the alignment file, where all the potential DNA contaminants, derived from the production process, are used as the target genome, the number of reads and the position of the contaminants are returned and presented as count table and coverage figures[57, 79, 80].
13.2.3.3 AAV Genome Rearrangements
In postmitotic cells, rAAV mainly persists as an extrachromosomal element able to stably express the transgene for a long period of time[6, 81–83]. The inter and intra‐molecular recombination of the ITRs produces three primary types of cir­cles: 5′ITR‐3′ITR (head to tail), 5′ITR‐5′ITR (head to head), and 3′ITR‐3′ITR (tail to tail). While in dividing cells, the mechanism that ensures the persistence of the rAAV DNA seems to be the integration into the host DNA. Assessing the amount of episomal circular/concatemeric forms of the vector and distinguishing the ele­ments from the integrated vector genome is a challenge for the analysis tools. Searching for reads that map discontinuously on the AAV genome is the primary strategy for determining the frequency of the episomal forms in the samples [22, 77] assuming that rearrangement among vector is an indication of an extra­chromosomal AVV genome. It has been shown that rearranged vectors can inte­grate into the host genome[3, 19, 84], and that numerous rearrangement events can lead to complicated structures of concatenated AAV fragments. For this rea­son, it is not possible to univocally distinguish between integrated/nonintegrated AAV DNA without more specialized methods that make use of long‐read sequenc­ing[19, 85] and new bioinformatics tools[3, 19].
13.2.3.4 Integration Site Analysis
The most common question answered by advanced bioinformatics methods is the location of IS. Whilst only a few reads are inspected visually in the manual analysis, in all the modern analysis tools, a multitude of criteria are introduced to select, evaluate, and report integration events. Here, we will discuss the most relevant standard analysis.
It is not uncommon that a single DNA fragment aligns with multiple genomic
regions, and these analyses must be handled with caution to minimize signal loss.
13.2 Review of Regulatory Guidance and Discussion Points that Are Raised on AAV Carcinogenesis 333
https://t.me/medicina_free
For example, integrations into a repetitive area of the host genome cannot be identified uniquely and represents potential unresolved IS. Different strategies can be adopted to reduce the impact of unresolved IS. A straightforward but not always implementable experimental solution is to increase the reads’ length to decrease the number of unresolved IS. Alternatively, a computational strategy that increases the reads’ dimension in silico can be implemented as established in gamma‐TRIS aligner [86], a tool for detecting retroviral integrations. More straightforward but less efficient approaches discard the reads that show align­ments with the same score on multiple locations[13] or when the score between the best match and the second best is more significant than a threshold [75]. Additionally, the mapping of IS into low complexity and repetitive regions depends on the alignment tool sensitivity. For example, by default, BLAST masks low complexity regions from the search[87], making it practically blind in a not negligible portion of the human genome, which results in a loss of sensitivity [3, 13, 15, 26, 77].
Besides host genome mapping, identification of the junction, i.e. the region where the AAV genome and the reference are fused can be challenging. Ideally, when the AAV sequence ends, the genomic begins so that the junction region is clearly defined. While preparing the sequencing libraries, different operation on the samples introduces noise in the form of (1) minor variations, (2) biased ampli­fication rate, and (3) invitro recombination of the DNA fragments. The last one, producing biological artifacts[88], affects the integration site detection. For this reason, the analysis tool should control the formation of the chimeric reads by requiring that the vector‐genome junction sequences are well‐formed.
In WGS and TE methods, the reads aligned to the vector are parsed to validate the alignment structure. The Concise Idiosyncratic Gapped Alignment Report (CIGAR) string in the Sequence Alignment/Map (SAM) format is used to authen­ticate the alignment structure[19, 75]. A less stringent requirement sometimes adopted is the presence in the read pairs of at least one read aligned to the AAV genome[20, 39]. Analysis tools for PCR methods use filtering strategies that may involve a) a limited number of unaligned bases between the vector and the genomic regions[27, 75, 77] and the absence of more than one vector‐genome junction per fragment[3, 22]. This last strategy decreases the noise due to arti­facts, such as chimeras generated during the library preparation and sequencing but makes the analysis unable to detect rearranged AAV genome integration (see Section13.2.3.3).
13.2.3.5 Clonality Analysis
Due to the stochastic nature of the integration process, a single integration site marks univocally a single clone. The number of IS present at a particular time point in a given sample is thus a representation of the clonal configuration at that moment.
13 rAAV Integration: Detection and Risk Assessment
https://t.me/medicina_free
334
If some IS are found more frequently than others, then it can be inferred that those IS may be a sign of clonal expansion of the cell harboring that integration. The assessment that addresses the reconstruction and comparison of the clonality of a set of samples is composed of two procedures: (1) the rebuilding of the clonal­ity of any single IS detected within a sample (IS clonality) and (2) the calculation of the clonality of the sample (clonal diversity).
The first approach associates with each IS a number representing the total num­ber of times each clone is present in the sample. The straightforward derivation of IS clonality is by assessing the number of times a particular IS is detected in the given sample[3, 18, 89]. This approach, however, may be unable to distinguish between reads from different clones or PCR amplification. Other counting strate­gies can reduce this issue by exploiting experimental evidence and adding a few assumptions[19, 20, 27, 39]. For example, libraries that originate from random DNA shearing (e.g. by sonication) are less prone to contain identical fragments from the same clones, thus reducing the count bias introduced by the PCR[63]. On the other hand, this approach requires adjustment when the sequencing depth increases to a level where the sonicated fragments are saturated[63].
The clonal diversity is calculated from the IS clonality employing methods used in ecology for measuring biodiversity. Its main aim is to compare the clonal reper­toire in the same individual and, as an example, identify the surge of a clone, intercepting potential adverse events caused by insertional mutagenesis as soon as possible. The simplest and frequently used measure of diversity is clonal abun­dance[27, 40, 78], which determines the number of times a clone is present in the sample compared to the total number of clones. When the clonal abundance of a specific IS increases over a defined threshold, a red flag is raised, triggering more detailed investigations. With this regard, several diversity indexes are employed, such as the Shannon index[90], the polyclonal monoclonal diversity (PMD)[91], and the shape‐constrained splines (SCS) method [92] in GTx for tracking and comparing clonal diversity across samples and studies.
13.2.3.6 Genotoxic Integrations
The emergence of a clone that expands during longitudinal sampling triggers a major concern for an insertional mutagenesis event[93]. It is a common practice to assign each IS to the closest gene within a specific range[1, 13, 15, 18, 20, 27]. In this way, the analysis can focus only on the integrations close to a subset of cancer‐associated genes. This may seem a straightforward process, however, selecting a comprehensive and meaningful set of detrimental genes is challeng­ing. Aside from gamma‐retroviral studies that established a specific (limited) group of genes as a source of potential insertional mutagenesis[94, 95], for more comprehensive analysis different, not overlapping collections of cancer genes, are curated by multiple databases[96–98]. An additional layer of complexity in the