Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5943_Библиотеки_им_академика_М_И_Перельмана
.pdf
https://t.me/medicina_free
356
been linearized will have ends that are available for adapter ligation and sequencing.
CHANGE‐seq is a streamlined high‐throughput improvement on the original
CIRCLE‐seq method that utilizes Tn5 tagmentation to eliminate a large number
of molecular biology steps. In the SITE‐seq protocol, sequencing adapters are
sequentially added to high‐molecular‐weight gDNA that has been treated with
Cas9 digestion and fragmentation, respectively. Biotinylated adapter containing
DNA is pulled down via streptavidin affinity purification (Figure 14.4C).
In direct comparisons of cellular GUIDE‐seq and biochemical CHANGE‐seq
methods, we found that sites detected exclusively by CHANGE‐seq could be confirmed in cells and concluded that CHANGE‐seq is more sensitive, a likely general property of biochemical methods where superphysiologic ratios of enzyme to
gDNA substrates can be used to maximize detection sensitivity.
Currently, using a combination of cellular and biochemical methods can provide the best mix of direct detection and sensitivity. In the future, optimizing
methods like GUIDE‐seq for more cell types will be important to support the
increasing range of cells that are being edited for clinical applications.
14.3.3 Targeted Approaches toMeasure Short Insertions
and Deletions
Insertions and deletions are the footprints of nuclease activity–editing efficiency
and specificity can be assessed by analyzing indels at the on‐target and off‐target
genomic locations. Editing efficiency is usually expressed as the percentage of
DNA alleles in the study sample that possess indels, i.e. are edited. Since NHEJ
mechanism of repair is often imprecise, it results in a polyclonal cell mixture with
a variety of different indels. Detection of indel sequences in conjunction with
quantification of each clone’s frequency is a safety evaluation of clonal
amplification.
There are several molecular biology techniques that were either repurposed or
specifically developed to detect the short indels introduced by engineered nucleases. All these methods rely on enrichment of the genomic locus containing the
edited site. A majority of indels introduced by engineered nucleases are under
bp[48, 49]. Most commonly the enrichment is performed by PCR using prim-
50
ers flanking the edited site. Once the edited site is enriched it can be analyzed
using four main groups of approaches – Next‐Generation Sequencing (NGS),
Endonuclease Mismatch Cleavage Assays (EMC), digital droplet PCR (ddPCR),
Sanger sequencing combined with sequence trace decomposition and Indel
Detection by Amplicon Analysis (IDAA) (Table14.1). This section will provide an
overview of each technique, the context of use, and technique’s advantages and
disadvantages.

GUIDE-seq Digenome-seq CHANGE-seq
q
(a) (b) (c)
https://t.me/medicina_free
dsODN
DSB
dsODN integration in live cells
Shearing, adapter ligation and
tag specic amplication
High-throughput sequencing
Figure14.4 Techniques to identify genome-wide off-target sites. (a) GUIDE-seq is based on the principle of efficient integration of
short, end-protected oligodeoxynucleotide (dsODN) tags into the sites of nuclease-induced DSBs. Genomic DNA shearing, adapter
ligation, and tag-specific amplification are performed to map cellular DSBs. GUIDE-seq reads are proportional to indel mutation
frequencies in cells. (b) Genomic DNA is treated with Cas9 ribonucleoprotein complexes, sheared, ligated to adapters for wholegenome sequencing. After sequencing, sites with uniform ends that are likely created by Cas9invitro cleavage are identified with a
bioinformatic algorithm. (c) CHANGE-seq is based on the principle of selective sequencing of nuclease-modified genomic
DNA. Genomic DNA is tagmented with Tn5 to add adapters for circularization. Genomic DNA is circularized by intramolecular ligation
and excess linear DNA removed by exonuclease treatment. Genomic DNA circles are treated with Cas9 and only circles that have
been cut will have ends available for adapter ligation, PCR, and high-throughput sequencing. CHANGE-seq reads map outward from
the breakpoint location.
In vitro DSB
Shearing and adapter ligation
Whole-genome sequencing
gDNA
Tagmented
gDNA
Uncut
Tn5
Tagmentation
Circularization and
residual linear DNA degradation
Cut
Cas9 cutting
Paired-end
high-throughput sequencing
Adapter ligation and PCR
DSB
CHANGE-se
reads

https://t.me/medicina_free

359
https://t.me/medicina_free
Targeted Next-generation Sequencing forIndel Detection
NGS is currently accepted in the field as the gold standard for high‐throughput
detection of genome editing events and is likely the most accurate analysis platform.
It is the only method that allows for simultaneous quantification of both on‐ and
off‐target edits in multiple samples. Since its first use in detecting CCR5 modification introduced by ZFNs using Roche 454[50] and Illumina sequencing[51], it is
now widely adopted for all other genome editing platforms.
Whole genome sequencing is rarely used to assess genome editing due to limited sensitivity and high cost. More commonly a targeted sequencing approach is
employed where the edited sites are first enriched and then sequenced, followed
by a bioinformatics analysis that determines the exact sequences and frequencies
of the indels. Targeting the specific editing sites allows a higher depth of sequencing, resulting in a higher sensitivity and confidence of adjudicating indels. In
addition, the number of individual samples to be pooled into one sequencing run
can be increased, resulting in lower sequencing costs per sample and shorter turnaround time[52, 53].
Sequencing is commonly performed using sequencing‐by‐synthesis on MiSeq
or NextSeq instruments from Illumina though the use of other technologies such
as Oxford Nanopore and PacBio sequencing is increasing [54, 55]. The use of
PacBio sequencing is limited due to specialized sample preparation and high artifactual rate of long‐range PCR (PCR chimeras and amplicon length‐dependent
biases).
This section will focus on targeted NGS performed using sequencing‐by‐
synthesis from Illumina for products enriched by two commonly used techniques–
amplicon sequencing and hybrid capture.
Amplicon Sequencing
Amplicon sequencing relies on two sequential PCR reactions (Figure14.5A). The
first PCR reaction is a locus‐specific PCR where the edited site and flanking
sequences (approximately 75 bp on each side) are amplified with primers containing
overhang sequences. An aliquot of the first PCR reaction is then taken into a second
PCR reaction where overhang sequences are recognized by the second primer
pair, which adds sample specific barcodes and the P5 and P7 universal Illumina
sequencing adaptors. The resulting PCR products are then analyzed for purity
using gel electrophoresis, pooled before purification and normalization and
sequenced. The sample specific barcodes are 6–8nucleotide sequences that are
unique for each PCR reaction. Their presence allows for the pooling of multiple
samples into the same sequencing run. Post sequencing, the barcodes allow
the assignment of sequencing reads to individual samples bioinformatically. This
enables amplicon sequencing to analyze multiple samples in a single run.

Amplicon sequencing
Amplicons analysis by capillary
Hybrid capture-based
Sanger sequencing with
Indel Detection by
(a) (b) (c) (d) (e) (f)
https://t.me/medicina_free
Locus specific PCR
WT
indel
Adding sample specific barcodes
and sequencing adapters
Pooling, purification
Sequencing and
bioinformatics
Figure14.5 Targeted approaches to measure short insertions and deletions. (a) Amplicon sequencing; (b) Hybrid capture-based
sequencing; (c) Droplet Digital PCR (one of various approaches for indel detection is illustrated); (d) Endonuclease Mismatch
Cleavage Assays; (e) Sanger sequencing combined with sequence trace decomposition (TIDE workflow is illustrated); (f) Indel
Detection by Amplicon Analysis (IDAA). Note, all methods require extracted gDNA which is not illustrated here. Source: Created with
BioRender.com.
sequencing
gDNA sonication
End repair, sequencing
adapters ligation, PCR
Probe hybridization,
washing, hybrids capture
Captured products elution,
PCR and purification
PCR in nanoliter-sized droplets
WT
indel
Droplet Digital
with TaqMan probes
Ref
Ref
Droplet counting and
signal detection
Ref signal
NHEJ signal
empty
(Ref-/NHEJ-)
NHEJ
TM
NHEJ
WT
(Ref+/
NHEJ+)
indel
(Ref+/
NHEJ-)
PCR
Endonuclease Mismatch
Cleavage Assay
PCR,denaturing and annealing
WT
indel
Homoduplex
Heteroduplex
Endonuclease treatment
and analysis
Heteroduplex
Homoduplex
sequence trace decomposition
PCR and Sanger sequencing
WT
mixed
pool
Trace decomposition
80
60
40
deletion
insertion
20
% of sequence
0
–10 –5 0+5 +10
indel position
Amplicon Analysis
Tri-primer PCR
FAM
WT
indel
electrophoresis
amplicon length
WT
deletions
amplicon amount
insertions

361
https://t.me/medicina_free
When performing multiplexed amplicon sequencing, the design of the primers
requires special attention. Primers should be highly specific to the queried sites
without cross‐reacting with each other to minimize primer‐dimer formation and
amplification at sites other than target site. In addition, to ensure equal sequencing coverage across all amplicons, the ratio between different primer pairs needs
to be optimized to prevent preferential amplification of some amplicons. There
are several technologies that are addressing these concerns by using highly developed primer design algorithms. rhAmpSeq combines primer design algorithm
with increased PCR specificity by using RNase H‐dependent PCR[56, 57]. Though
this technology is marketed for CRISPR‐based editing, rhAmpSeq can be readily
applied to detect indels introduced by other nucleases if alternative sequencing
analysis software is used. This technology enables the examination of hundreds of
edited sites in the same sample. In rhAmpSeq each primer contains an RNA base
at or near the 3′‐end of the primer and lacks the free 3′‐OH group, thereby needing
“activation” to be extended by DNA polymerase. Once perfect complementarity
between primer and its target DNA is achieved, the primer:target heteroduplex is
cleaved at the 5′‐side of the RNA base by thermally stable RNase H2 supplied as a
part of PCR master mix. The resulting 3′‐OH group then allows DNA polymerase
to carry out the primer extension. Due to the requirement for primers to be “activated” by first binding to the target sequence, primer‐dimer formation and false
amplification of similar sequences are greatly reduced.
The advantages of amplicon sequencing techniques are relative ease of design
and implementation. In addition, the library preparation protocol is simple, quick,
and not labor‐intensive. Perhaps the main disadvantage of amplicon sequencing
is the difficulty of integrating unique molecular identifies (UMIs) into the workflow. UMIs are molecular barcodes that comprise of short sequences used to
uniquely tag each molecule in a sample prior to PCR amplification. UMIs can be
used to remove sequencing reads that arise from PCR duplicates as PCR duplicates will have the same UMI. This allows for accurate allele frequency detection
as PCR duplicates can falsely overrepresent different allele frequencies if more
PCR copies were generated for one variant compared to other variants present in
a sample.
A recent publication showed that UMI tagging can be performed using the
primer extension reaction with one UMI primer prior to PCR. The extension
product is then purified using AMPure XP beads and PCR‐amplified using the
universal primer and the gene‐specific reverse primer[54]. This application has
yet to be broadly used.
Hybrid Capture-based Sequencing
A second enrichment technique is referred to as hybrid capture[52] (Figure14.5B).
Briefly, gDNA is sheared into 120–200 bp DNA fragments using sonication. DNA

https://t.me/medicina_free
362
fragments are then end‐repaired and ligated with sequencing adaptors that
include sample specific barcodes and P5 and P7 universal Illumina sequencing
adaptors. The ligation products are then subjected to PCR to generate sufficient
material for the capture reaction. The library of amplified genomic fragments is
hybridized with biotinylated RNA or DNA probes tiling across the candidate
edited site and flanking sequences. The hybridization products are then purified
using streptavidin beads, amplified, purified, and sequenced. This hybrid capture
approach is compatible with UMI tagging, which can address artifacts arising
from PCR and sequencing errors. The disadvantages of hybrid capture approaches
are requirements for substantial assay optimization, including gDNA sheering
conditions and PCR conditions. In addition, DNA capture is less efficient for
GC‐ and AT‐rich sequences. Since off‐target sites are often located in intronic and
intergenic regions the efficient probe development to accommodate all edited
sites can be challenging.
NGS Sensitivity for Indel Detection and Quantification
In general, the sensitivity of indel detection using NGS depends on sequencing
depth (how many sequencing reads correspond to each analyzed sample) and
background signal (modifications measured in non‐edited sample).
There is an inverse correlation between level of editing and sequencing depth
requirements–samples with a low editing frequency require a higher number of
sequencing reads for accurate detection. So far, no systematic study has been published to investigate the exact relationship between modification frequency and
number of sequencings reads that are needed for accurate quantification. A coverage of >1000 paired reads per edited site was shown to allow for detection of
indels down to 0.5% with a 40% coefficient of variation (CV)[57]. In another publication, the coverage of >5000 paired reads per edited site was suggested to detect
indels down to 0.2%, but the precision of detection was not specified[58].
NGS background is defined as detected modifications in unedited samples.
Such background can arise due to errors introduced during PCR amplification
and sequencing. The background noise reported for Illumina sequencing is
0.1–0.3%[59, 60]. Background noise was also shown to increase when poor quality gDNA was used[61]. Investigation of background modification levels across
273 CRISPR edited sites showed that the background varies from 0% to 1.0%, with
98% of sites having background indels ranging from 0% to 0.4%[57]. One of the
ways to decrease the background bioinformatically is to shorten the sequence
“window” where indels are quantified. Several published software packages stress
the importance of quantifying indels in a specific sequence “window” where
modification is expected to occur when high sensitivity of detection is desired.
Such a “window” is specific to each type of engineered nuclease and needs to
be experimentally established. When the optimal “window” was applied, a 60%

363
https://t.me/medicina_free
decrease in background (false‐positive indel signal) was seen while retaining
>98% indels in treated samples [57]. The optimal sequence “window” can be
established by careful analysis of oligo incorporation using unbiased genome‐
wide sequencing methods e.g. GUIDE‐seq. The exact genomic position of incorporated oligo can be used to establish the correct sequence “window” to analyze
and quantify indels.
The sensitivity of indel quantification can be thought to depend on the particular NGS approach (e.g. amplicon sequencing vs. hybrid capture). Head‐to‐head
comparison of hybrid capture and multiplexed amplicon sequencing (rhAmpSeq)
for edited sites introduced by CRISPR‐Cas9 showed high data similarity; the
rhAmpSeq analysis was more sensitive likely due to the higher sequencing depth
used[58]. Recently, Kurgan etal. investigated the sensitivity of the rhAmpSeq
assay[57]. They prepared a mixture of synthetic DNA containing insertions, deletions, and SNPs to simulate the editing of the HPRT1 genomic locus by CRISPR‐
Cas9nuclease. This mixture was then serially diluted with wild‐type synthetic
DNA to create samples with 0–100% “modifications,” then analyzed using the
rhAmpSeq protocol with a sequencing depth of >40,000 reads per sample. The
authors found that at 1% modification and lower, there was a ≥20% deviation in
detected modification levels compared to that expected. The accuracy of detection
was improved by applying a background correction. This allowed the detected
modifications down to 0.1%, within 10% of expected values. Whether such high
sensitivity can be achieved in real biological samples for all investigated sites
requires additional investigation. Indeed, the authors emphasized the need for
improving library construction chemistry such as incorporating UMIs and coming
up with sophisticated background correction techniques to improve the accuracy
and precision of indel calling.
The assay LOD is often defined as the indel frequency that can be detected as a
statistically significant difference between edited and unedited samples. In two
publications, the edited sites were defined as confirmed when the difference in
indel frequencies between treated and untreated samples was higher than
0.16%[62] and 0.20%[58]. In the scenario where untreated samples show no background indels, the highest possible theoretical sensitivity would be 0.16% and
0.20% indels, respectively.
Miller et al. demonstrated high indel sensitivity (≈0.001–0.025% indels with
50% CV) at 100 AAVS1 off‐target sites assessed as independent PCR reactions and
sequenced using optimized conditions on NextSeq Illumina[63]. High sensitivity
was achieved through increasing sequencing depth and decreasing background
signal to 0.01% indels by implementing multiple improvements. Sequencing
depth was increased so that each input DNA allele was sequenced at least ten
times (≥200,000 reads per replicate per target site), and multiple technical replicates
of edited samples and up to 24 replicates of unedited background samples were

https://t.me/medicina_free
364
run. The occurrence of sequencing artifacts was decreased by reducing the cluster
density on the flow cell during the sequencing reaction and precisely defining the
sequence “window” where editing is expected to occur. This method is amenable
for invitro studies where multiple biological replicates of treated and untreated
samples can be run.
Bioanalytical Characterization of NGS for Indel Detection and Quantification
The FDA guidance for bioanalytical method validation[64] underlines the importance of analytical method characterization supporting regulatory submissions.
This includes bioanalytical parameters, such as sensitivity, specificity, accuracy,
precision, and description of the reference materials. There is currently no specific FDA guidance for analytical method characterization for off‐target detection.
This section will focus on our case study of indel quantification in ZFN‐edited
hematopoietic stem and progenitor cells using NGS, specifically, how to generate
material for making quality controls (QCs) . Reference standards and QCs are
important reagents for quantification of most detection methodologies. QCs are
commonly used to develop and characterize the assay and to monitor run‐to‐run
variability when drug product or post‐infusion patient samples are analyzed.
There are several approaches for QC creation that have been explored in the field.
One approach is to use engineered cell lines by introducing edits at the same
genomic locations as expected in study samples (cell therapy drug product as well
as animal and human samples post‐infusion). Such cell lines can be generated by
transfecting immortalized cells with studied engineered nucleases to cause a high
level of indels, followed by the propagation and banking of these cell lines. The
level of indels can be determined by NGS and, ideally, by an orthogonal methodology
e.g. ddPCR. QC samples containing different levels of indels then can be generated by mixing a cell line with a known frequency of editing together with WT
cells at different ratios[25] or isolated gDNA.
An alternative approach relies on preparing a mixture of synthetic DNA containing the anticipated edits with wild‐type synthetic DNA to create samples with
various modification levels[57]. The synthetic DNA mixtures can then be spiked
into gDNA isolated from WT animal or human tissues to include background
matrix of the study samples.
The synthetic DNA approach and edited cells approach both have their respective advantages and limitations. Engineered cell lines can be cumbersome to generate and maintain, which substantially extend assay development timelines,
especially if multiple editing sites are explored (on‐ and off‐target editing in the
same assay). Multiple cell lines with individual edits can be pooled or the same
cell line can be sequentially edited. However, the clear advantage of engineered
control cell lines is that they better represent the complexity of study samples and
editing outcomes. They can be easily propagated, and more QC materials can be

365
https://t.me/medicina_free
made as needed using the same stock material. Though synthetic DNA can be
implemented faster than engineered cells, this approach suffers from its artificial
nature. In cell or tissue study samples, indel sequences at a known editing site
typically are heterogenous in nucleotide sequence and length. Usually, only a few
indel sequences are mimicked by the synthetic DNA, while others are not represented. In addition, synthetic DNA is usually short and therefore an easier target
for PCR amplification when compared to the more complex gDNA matrix present
in study samples. Some DNA sequences might be difficult to synthesize due to
their sequence composition (homopolymers, stem‐loop structures, etc.) which
may result in a low yield and high level of sequence errors in synthetic DNA. This
can be a common issue for off‐target sites that are often found in non‐coding
regions with sequence compositions that may be difficult to synthesize. When
synthetic DNA is ordered, it should be carefully evaluated how much material is
required to support testing for the entire clinical trial to avoid batch effects.
Therefore, the engineered control cell line approach may be better suited when
establishing method feasibility and accuracy if timeline permits. A synthetic DNA
approach can be implemented at the sample testing stage with the purpose of
assessing run‐to‐run variability.
In summary, targeted NGS sequencing provides the frequency and exact
sequence of indels. It can be used for high‐throughput quantification of both on‐
and off‐target edits with high sensitivity in one biological sample. There are several targeted NGS technologies available: a particular approach should be chosen
based on required sensitivity, number of analyzed sites, and the amount of sample
available for analysis. To support clinical sample monitoring, NGS assays can be
analytically characterized to establish assay sensitivity, variability, and acceptance
criteria.
14.3.3.1 Droplet Digital™ PCR
Droplet Digital™ PCR (ddPCR™) is a technique that allows for absolute quantitation of target DNA in complex biological matrices. The PCR reaction is fractionated into 20,000nanoliter‐size droplets. As a result of partitioning, target DNA
molecules get distributed across the droplets so that each droplet contains zero,
one, or a few copies of DNA templates. PCR amplification is then carried out to
the plateau phase and the droplets are counted as positive and negative reactions
by a droplet reader. The quantity of target DNA can be calculated using the fraction of positive droplets and Poisson statistics[65, 66]. This methodology was successfully adopted to assess NHEJ‐derived indels and implemented in various
research studies (GEF‐dPCR – gene‐editing frequencies digital PCR or DSB‐
ddPCR–DSB ddPCR)[67–69] (Figure14.5B). These techniques utilize two different TaqMan probes carrying two different fluorophores (FAM and HEX) targeting
the same PCR product and simultaneously detecting wild‐type and edited alleles.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
