Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5665_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
 
https://t.me/medicina_free
356
been linearized will have ends that are available for adapter ligation and sequencing. CHANGE‐seq is a streamlined high‐throughput improvement on the original CIRCLE‐seq method that utilizes Tn5 tagmentation to eliminate a large number of molecular biology steps. In the SITE‐seq protocol, sequencing adapters are sequentially added to high‐molecular‐weight gDNA that has been treated with Cas9 digestion and fragmentation, respectively. Biotinylated adapter containing DNA is pulled down via streptavidin affinity purification (Figure 14.4C).
In direct comparisons of cellular GUIDE‐seq and biochemical CHANGE‐seq methods, we found that sites detected exclusively by CHANGE‐seq could be con­firmed in cells and concluded that CHANGE‐seq is more sensitive, a likely gen­eral property of biochemical methods where superphysiologic ratios of enzyme to gDNA substrates can be used to maximize detection sensitivity.
Currently, using a combination of cellular and biochemical methods can pro­vide the best mix of direct detection and sensitivity. In the future, optimizing methods like GUIDE‐seq for more cell types will be important to support the increasing range of cells that are being edited for clinical applications.
14.3.3 Targeted Approaches toMeasure Short Insertions and Deletions
Insertions and deletions are the footprints of nuclease activity–editing efficiency and specificity can be assessed by analyzing indels at the on‐target and off‐target genomic locations. Editing efficiency is usually expressed as the percentage of DNA alleles in the study sample that possess indels, i.e. are edited. Since NHEJ mechanism of repair is often imprecise, it results in a polyclonal cell mixture with a variety of different indels. Detection of indel sequences in conjunction with quantification of each clone’s frequency is a safety evaluation of clonal amplification.
There are several molecular biology techniques that were either repurposed or specifically developed to detect the short indels introduced by engineered nucle­ases. All these methods rely on enrichment of the genomic locus containing the edited site. A majority of indels introduced by engineered nucleases are under
bp[48, 49]. Most commonly the enrichment is performed by PCR using prim-
50 ers flanking the edited site. Once the edited site is enriched it can be analyzed using four main groups of approaches – Next‐Generation Sequencing (NGS), Endonuclease Mismatch Cleavage Assays (EMC), digital droplet PCR (ddPCR), Sanger sequencing combined with sequence trace decomposition and Indel Detection by Amplicon Analysis (IDAA) (Table14.1). This section will provide an overview of each technique, the context of use, and technique’s advantages and disadvantages.
GUIDE-seq Digenome-seq CHANGE-seq
q
(a) (b) (c)
https://t.me/medicina_free
dsODN
DSB
dsODN integration in live cells
Shearing, adapter ligation and
tag specic amplication
High-throughput sequencing
Figure14.4 Techniques to identify genome-wide off-target sites. (a) GUIDE-seq is based on the principle of efficient integration of
short, end-protected oligodeoxynucleotide (dsODN) tags into the sites of nuclease-induced DSBs. Genomic DNA shearing, adapter ligation, and tag-specific amplification are performed to map cellular DSBs. GUIDE-seq reads are proportional to indel mutation frequencies in cells. (b) Genomic DNA is treated with Cas9 ribonucleoprotein complexes, sheared, ligated to adapters for whole­genome sequencing. After sequencing, sites with uniform ends that are likely created by Cas9invitro cleavage are identified with a bioinformatic algorithm. (c) CHANGE-seq is based on the principle of selective sequencing of nuclease-modified genomic DNA. Genomic DNA is tagmented with Tn5 to add adapters for circularization. Genomic DNA is circularized by intramolecular ligation and excess linear DNA removed by exonuclease treatment. Genomic DNA circles are treated with Cas9 and only circles that have been cut will have ends available for adapter ligation, PCR, and high-throughput sequencing. CHANGE-seq reads map outward from
the breakpoint location.
In vitro DSB
Shearing and adapter ligation
Whole-genome sequencing
gDNA
Tagmented
gDNA
Uncut
Tn5
Tagmentation
Circularization and residual linear DNA degradation
Cut
Cas9 cutting
Paired-end
high-throughput sequencing
Adapter ligation and PCR
DSB
CHANGE-se reads
https://t.me/medicina_free
         359
https://t.me/medicina_free
Targeted Next-generation Sequencing forIndel Detection
NGS is currently accepted in the field as the gold standard for high‐throughput detection of genome editing events and is likely the most accurate analysis platform. It is the only method that allows for simultaneous quantification of both on‐ and off‐target edits in multiple samples. Since its first use in detecting CCR5 modifica­tion introduced by ZFNs using Roche 454[50] and Illumina sequencing[51], it is now widely adopted for all other genome editing platforms.
Whole genome sequencing is rarely used to assess genome editing due to lim­ited sensitivity and high cost. More commonly a targeted sequencing approach is employed where the edited sites are first enriched and then sequenced, followed by a bioinformatics analysis that determines the exact sequences and frequencies of the indels. Targeting the specific editing sites allows a higher depth of sequenc­ing, resulting in a higher sensitivity and confidence of adjudicating indels. In addition, the number of individual samples to be pooled into one sequencing run can be increased, resulting in lower sequencing costs per sample and shorter turn­around time[52, 53].
Sequencing is commonly performed using sequencing‐by‐synthesis on MiSeq or NextSeq instruments from Illumina though the use of other technologies such as Oxford Nanopore and PacBio sequencing is increasing [54, 55]. The use of PacBio sequencing is limited due to specialized sample preparation and high arti­factual rate of long‐range PCR (PCR chimeras and amplicon length‐dependent biases).
This section will focus on targeted NGS performed using sequencing‐by‐ synthesis from Illumina for products enriched by two commonly used techniques– amplicon sequencing and hybrid capture.
Amplicon Sequencing
Amplicon sequencing relies on two sequential PCR reactions (Figure14.5A). The first PCR reaction is a locus‐specific PCR where the edited site and flanking sequences (approximately 75 bp on each side) are amplified with primers containing overhang sequences. An aliquot of the first PCR reaction is then taken into a second PCR reaction where overhang sequences are recognized by the second primer pair, which adds sample specific barcodes and the P5 and P7 universal Illumina sequencing adaptors. The resulting PCR products are then analyzed for purity using gel electrophoresis, pooled before purification and normalization and sequenced. The sample specific barcodes are 6–8nucleotide sequences that are unique for each PCR reaction. Their presence allows for the pooling of multiple samples into the same sequencing run. Post sequencing, the barcodes allow the assignment of sequencing reads to individual samples bioinformatically. This enables amplicon sequencing to analyze multiple samples in a single run.
Amplicon sequencing
Amplicons analysis by capillary
Hybrid capture-based
Sanger sequencing with
Indel Detection by
(a) (b) (c) (d) (e) (f)
https://t.me/medicina_free
Locus specific PCR
WT
indel
Adding sample specific barcodes
and sequencing adapters
Pooling, purification
Sequencing and
bioinformatics
Figure14.5 Targeted approaches to measure short insertions and deletions. (a) Amplicon sequencing; (b) Hybrid capture-based
sequencing; (c) Droplet Digital PCR (one of various approaches for indel detection is illustrated); (d) Endonuclease Mismatch Cleavage Assays; (e) Sanger sequencing combined with sequence trace decomposition (TIDE workflow is illustrated); (f) Indel Detection by Amplicon Analysis (IDAA). Note, all methods require extracted gDNA which is not illustrated here. Source: Created with
BioRender.com.
sequencing
gDNA sonication
End repair, sequencing adapters ligation, PCR
Probe hybridization,
washing, hybrids capture
Captured products elution,
PCR and purification
PCR in nanoliter-sized droplets
WT
indel
Droplet Digital
with TaqMan probes
Ref
Ref
Droplet counting and
signal detection
Ref signal
NHEJ signal
empty (Ref-/NHEJ-)
NHEJ
TM
NHEJ
WT (Ref+/
NHEJ+)
indel (Ref+/ NHEJ-)
PCR
Endonuclease Mismatch
Cleavage Assay
PCR,denaturing and annealing
WT
indel
Homoduplex
Heteroduplex
Endonuclease treatment
and analysis
Heteroduplex
Homoduplex
sequence trace decomposition
PCR and Sanger sequencing
WT
mixed
pool
Trace decomposition
80
60
40
deletion
insertion
20
% of sequence
0
–10 –5 0+5 +10
indel position
Amplicon Analysis
Tri-primer PCR
FAM
WT
indel
electrophoresis
amplicon length
WT
deletions
amplicon amount
insertions
         361
https://t.me/medicina_free
When performing multiplexed amplicon sequencing, the design of the primers requires special attention. Primers should be highly specific to the queried sites without cross‐reacting with each other to minimize primer‐dimer formation and amplification at sites other than target site. In addition, to ensure equal sequenc­ing coverage across all amplicons, the ratio between different primer pairs needs to be optimized to prevent preferential amplification of some amplicons. There are several technologies that are addressing these concerns by using highly devel­oped primer design algorithms. rhAmpSeq combines primer design algorithm with increased PCR specificity by using RNase H‐dependent PCR[56, 57]. Though this technology is marketed for CRISPR‐based editing, rhAmpSeq can be readily applied to detect indels introduced by other nucleases if alternative sequencing analysis software is used. This technology enables the examination of hundreds of edited sites in the same sample. In rhAmpSeq each primer contains an RNA base at or near the 3′‐end of the primer and lacks the free 3′‐OH group, thereby needing “activation” to be extended by DNA polymerase. Once perfect complementarity between primer and its target DNA is achieved, the primer:target heteroduplex is cleaved at the 5′‐side of the RNA base by thermally stable RNase H2 supplied as a part of PCR master mix. The resulting 3′‐OH group then allows DNA polymerase to carry out the primer extension. Due to the requirement for primers to be “acti­vated” by first binding to the target sequence, primer‐dimer formation and false amplification of similar sequences are greatly reduced.
The advantages of amplicon sequencing techniques are relative ease of design and implementation. In addition, the library preparation protocol is simple, quick, and not labor‐intensive. Perhaps the main disadvantage of amplicon sequencing is the difficulty of integrating unique molecular identifies (UMIs) into the work­flow. UMIs are molecular barcodes that comprise of short sequences used to uniquely tag each molecule in a sample prior to PCR amplification. UMIs can be used to remove sequencing reads that arise from PCR duplicates as PCR dupli­cates will have the same UMI. This allows for accurate allele frequency detection as PCR duplicates can falsely overrepresent different allele frequencies if more PCR copies were generated for one variant compared to other variants present in a sample.
A recent publication showed that UMI tagging can be performed using the primer extension reaction with one UMI primer prior to PCR. The extension product is then purified using AMPure XP beads and PCR‐amplified using the universal primer and the gene‐specific reverse primer[54]. This application has yet to be broadly used.
Hybrid Capture-based Sequencing
A second enrichment technique is referred to as hybrid capture[52] (Figure14.5B). Briefly, gDNA is sheared into 120–200 bp DNA fragments using sonication. DNA
 
https://t.me/medicina_free
362
fragments are then end‐repaired and ligated with sequencing adaptors that include sample specific barcodes and P5 and P7 universal Illumina sequencing adaptors. The ligation products are then subjected to PCR to generate sufficient material for the capture reaction. The library of amplified genomic fragments is hybridized with biotinylated RNA or DNA probes tiling across the candidate edited site and flanking sequences. The hybridization products are then purified using streptavidin beads, amplified, purified, and sequenced. This hybrid capture approach is compatible with UMI tagging, which can address artifacts arising from PCR and sequencing errors. The disadvantages of hybrid capture approaches are requirements for substantial assay optimization, including gDNA sheering conditions and PCR conditions. In addition, DNA capture is less efficient for GC‐ and AT‐rich sequences. Since off‐target sites are often located in intronic and intergenic regions the efficient probe development to accommodate all edited sites can be challenging.
NGS Sensitivity for Indel Detection and Quantification
In general, the sensitivity of indel detection using NGS depends on sequencing depth (how many sequencing reads correspond to each analyzed sample) and background signal (modifications measured in non‐edited sample).
There is an inverse correlation between level of editing and sequencing depth requirements–samples with a low editing frequency require a higher number of sequencing reads for accurate detection. So far, no systematic study has been pub­lished to investigate the exact relationship between modification frequency and number of sequencings reads that are needed for accurate quantification. A cover­age of >1000 paired reads per edited site was shown to allow for detection of indels down to 0.5% with a 40% coefficient of variation (CV)[57]. In another pub­lication, the coverage of >5000 paired reads per edited site was suggested to detect indels down to 0.2%, but the precision of detection was not specified[58].
NGS background is defined as detected modifications in unedited samples. Such background can arise due to errors introduced during PCR amplification and sequencing. The background noise reported for Illumina sequencing is
0.1–0.3%[59, 60]. Background noise was also shown to increase when poor qual­ity gDNA was used[61]. Investigation of background modification levels across 273 CRISPR edited sites showed that the background varies from 0% to 1.0%, with 98% of sites having background indels ranging from 0% to 0.4%[57]. One of the ways to decrease the background bioinformatically is to shorten the sequence “window” where indels are quantified. Several published software packages stress the importance of quantifying indels in a specific sequence “window” where modification is expected to occur when high sensitivity of detection is desired. Such a “window” is specific to each type of engineered nuclease and needs to be experimentally established. When the optimal “window” was applied, a 60%
         363
https://t.me/medicina_free
decrease in background (false‐positive indel signal) was seen while retaining >98% indels in treated samples [57]. The optimal sequence “window” can be established by careful analysis of oligo incorporation using unbiased genome‐ wide sequencing methods e.g. GUIDE‐seq. The exact genomic position of incor­porated oligo can be used to establish the correct sequence “window” to analyze and quantify indels.
The sensitivity of indel quantification can be thought to depend on the particu­lar NGS approach (e.g. amplicon sequencing vs. hybrid capture). Head‐to‐head comparison of hybrid capture and multiplexed amplicon sequencing (rhAmpSeq) for edited sites introduced by CRISPR‐Cas9 showed high data similarity; the rhAmpSeq analysis was more sensitive likely due to the higher sequencing depth used[58]. Recently, Kurgan etal. investigated the sensitivity of the rhAmpSeq assay[57]. They prepared a mixture of synthetic DNA containing insertions, dele­tions, and SNPs to simulate the editing of the HPRT1 genomic locus by CRISPR‐ Cas9nuclease. This mixture was then serially diluted with wild‐type synthetic DNA to create samples with 0–100% “modifications,” then analyzed using the rhAmpSeq protocol with a sequencing depth of >40,000 reads per sample. The authors found that at 1% modification and lower, there was a ≥20% deviation in detected modification levels compared to that expected. The accuracy of detection was improved by applying a background correction. This allowed the detected modifications down to 0.1%, within 10% of expected values. Whether such high sensitivity can be achieved in real biological samples for all investigated sites requires additional investigation. Indeed, the authors emphasized the need for improving library construction chemistry such as incorporating UMIs and coming up with sophisticated background correction techniques to improve the accuracy and precision of indel calling.
The assay LOD is often defined as the indel frequency that can be detected as a statistically significant difference between edited and unedited samples. In two publications, the edited sites were defined as confirmed when the difference in indel frequencies between treated and untreated samples was higher than
0.16%[62] and 0.20%[58]. In the scenario where untreated samples show no back­ground indels, the highest possible theoretical sensitivity would be 0.16% and
0.20% indels, respectively.
Miller et al. demonstrated high indel sensitivity (≈0.001–0.025% indels with 50% CV) at 100 AAVS1 off‐target sites assessed as independent PCR reactions and sequenced using optimized conditions on NextSeq Illumina[63]. High sensitivity was achieved through increasing sequencing depth and decreasing background signal to 0.01% indels by implementing multiple improvements. Sequencing depth was increased so that each input DNA allele was sequenced at least ten times (≥200,000 reads per replicate per target site), and multiple technical replicates of edited samples and up to 24 replicates of unedited background samples were
 
https://t.me/medicina_free
364
run. The occurrence of sequencing artifacts was decreased by reducing the cluster density on the flow cell during the sequencing reaction and precisely defining the sequence “window” where editing is expected to occur. This method is amenable for invitro studies where multiple biological replicates of treated and untreated samples can be run.
Bioanalytical Characterization of NGS for Indel Detection and Quantification
The FDA guidance for bioanalytical method validation[64] underlines the impor­tance of analytical method characterization supporting regulatory submissions. This includes bioanalytical parameters, such as sensitivity, specificity, accuracy, precision, and description of the reference materials. There is currently no spe­cific FDA guidance for analytical method characterization for off‐target detection. This section will focus on our case study of indel quantification in ZFN‐edited hematopoietic stem and progenitor cells using NGS, specifically, how to generate material for making quality controls (QCs) . Reference standards and QCs are important reagents for quantification of most detection methodologies. QCs are commonly used to develop and characterize the assay and to monitor run‐to‐run variability when drug product or post‐infusion patient samples are analyzed. There are several approaches for QC creation that have been explored in the field. One approach is to use engineered cell lines by introducing edits at the same genomic locations as expected in study samples (cell therapy drug product as well as animal and human samples post‐infusion). Such cell lines can be generated by transfecting immortalized cells with studied engineered nucleases to cause a high level of indels, followed by the propagation and banking of these cell lines. The level of indels can be determined by NGS and, ideally, by an orthogonal methodology e.g. ddPCR. QC samples containing different levels of indels then can be gener­ated by mixing a cell line with a known frequency of editing together with WT cells at different ratios[25] or isolated gDNA.
An alternative approach relies on preparing a mixture of synthetic DNA con­taining the anticipated edits with wild‐type synthetic DNA to create samples with various modification levels[57]. The synthetic DNA mixtures can then be spiked into gDNA isolated from WT animal or human tissues to include background matrix of the study samples.
The synthetic DNA approach and edited cells approach both have their respec­tive advantages and limitations. Engineered cell lines can be cumbersome to gen­erate and maintain, which substantially extend assay development timelines, especially if multiple editing sites are explored (on‐ and off‐target editing in the same assay). Multiple cell lines with individual edits can be pooled or the same cell line can be sequentially edited. However, the clear advantage of engineered control cell lines is that they better represent the complexity of study samples and editing outcomes. They can be easily propagated, and more QC materials can be
         365
https://t.me/medicina_free
made as needed using the same stock material. Though synthetic DNA can be implemented faster than engineered cells, this approach suffers from its artificial nature. In cell or tissue study samples, indel sequences at a known editing site typically are heterogenous in nucleotide sequence and length. Usually, only a few indel sequences are mimicked by the synthetic DNA, while others are not repre­sented. In addition, synthetic DNA is usually short and therefore an easier target for PCR amplification when compared to the more complex gDNA matrix present in study samples. Some DNA sequences might be difficult to synthesize due to their sequence composition (homopolymers, stem‐loop structures, etc.) which may result in a low yield and high level of sequence errors in synthetic DNA. This can be a common issue for off‐target sites that are often found in non‐coding regions with sequence compositions that may be difficult to synthesize. When synthetic DNA is ordered, it should be carefully evaluated how much material is required to support testing for the entire clinical trial to avoid batch effects.
Therefore, the engineered control cell line approach may be better suited when establishing method feasibility and accuracy if timeline permits. A synthetic DNA approach can be implemented at the sample testing stage with the purpose of assessing run‐to‐run variability.
In summary, targeted NGS sequencing provides the frequency and exact sequence of indels. It can be used for high‐throughput quantification of both on‐ and off‐target edits with high sensitivity in one biological sample. There are sev­eral targeted NGS technologies available: a particular approach should be chosen based on required sensitivity, number of analyzed sites, and the amount of sample available for analysis. To support clinical sample monitoring, NGS assays can be analytically characterized to establish assay sensitivity, variability, and acceptance criteria.
14.3.3.1 Droplet Digital™ PCR
Droplet Digital™ PCR (ddPCR™) is a technique that allows for absolute quantita­tion of target DNA in complex biological matrices. The PCR reaction is fraction­ated into 20,000nanoliter‐size droplets. As a result of partitioning, target DNA molecules get distributed across the droplets so that each droplet contains zero, one, or a few copies of DNA templates. PCR amplification is then carried out to the plateau phase and the droplets are counted as positive and negative reactions by a droplet reader. The quantity of target DNA can be calculated using the frac­tion of positive droplets and Poisson statistics[65, 66]. This methodology was suc­cessfully adopted to assess NHEJ‐derived indels and implemented in various research studies (GEF‐dPCR – gene‐editing frequencies digital PCR or DSB‐ ddPCR–DSB ddPCR)[67–69] (Figure14.5B). These techniques utilize two differ­ent TaqMan probes carrying two different fluorophores (FAM and HEX) targeting the same PCR product and simultaneously detecting wild‐type and edited alleles.