Seminario 11 Técnicas de Biología Molecular aplicadas a Entidades Monogénicas - Sebastián Giusti — Transcript
Full transcript
- 0:00Hello, how are you? In today's class,
- 0:04we are going to analyze various
- 0:06molecular biology techniques used in
- 0:09the context of medical genetics, which
- 0:12aim to detect different allelic
- 0:14variants present in our genome. In the
- 0:19first part of the class, we will focus
- 0:22on variants of the PCR technique to
- 0:25understand the foundation of this
- 0:27method and its variations in different
- 0:30scenarios that may arise in the context
- 0:33of medical genetics. And in the second
- 0:36part, we will analyze techniques based
- 0:39on DNA sequencing. In both cases, the
- 0:43goal will be to understand, on one hand
- 0:47, the molecular foundations of each of
- 0:50these techniques, and also the medical
- 0:54and clinical context where these
- 0:56techniques become necessary,
- 0:59understanding in each case what
- 1:02information they provide in each of
- 1:05these contexts. We will begin by
- 1:09reviewing what you have surely seen in
- 1:13previous classes, which are the
- 1:15fundamentals of the PCR technique,
- 1:18which is the acronym in English for
- 1:21polymerase chain reaction. This
- 1:25technique, which was developed in the
- 1:2880s by the biochemist Kary Mullis, who
- 1:31received the Nobel Prize in Chemistry
- 1:35in 1993 for it. 3 is a technique that
- 1:40allows for the amplification of a
- 1:43specific DNA segment starting from a
- 1:46complex mixture. Hm. It is, therefore,
- 1:50the reproduction of a chemical reaction
- 1:54where what occurs is the polymerization
- 1:58of nucleic acids. One aspect, an
- 2:04advantage of this technique, is that
- 2:07one can design the PCR technique to
- 2:09amplify a specific segment of the
- 2:12genome, even if the starting template
- 2:15used is total genomic DNA. Hm. And this
- 2:21technique can then, despite starting
- 2:24from an extensive template, total
- 2:27genomic DNA, for example, extracted
- 2:30from a peripheral blood sample or other
- 2:33biological material, specifically
- 2:36amplify a single DNA segment from that
- 2:39sample. In this polymerase chain
- 2:44reaction, the elements introduced to
- 2:47produce it are, as we already saw, the
- 2:50genomic DNA extracted from a patient
- 2:53that constitutes the template for the
- 2:56polymerization reaction. The DNA
- 3:01polymerase enzyme used is a DNA
- 3:03polymerase of bacterial origin, Taq
- 3:06polymerase, because it is
- 3:08heat-resistant. And for this enzyme,
- 3:13which is the enzyme that will
- 3:16polymerize the substrates in the
- 3:19amplification reactions, it needs dNTPs
- 3:22as substrates, that is,
- 3:24deoxynucleotides of adenine, thymine,
- 3:28cytosine, and guanine. And it will also
- 3:32need primers, which we will briefly
- 3:34look at and review to understand the
- 3:36role they play in this reaction. For
- 3:41DNA polymerase to function, this enzyme
- 3:44requires a buffer, a stabilizing
- 3:46solution that keeps the pH of the
- 3:49reaction stable, and divalent magnesium
- 3:52cations, which act as a cofactor
- 3:57preserving the active three-dimensional
- 4:00structure of the DNA polymerase. Let's
- 4:05focus once again on dNTPs and primers
- 4:09as molecular requirements for DNA
- 4:12polymerase activity. We just mentioned
- 4:18that these deoxyribonucleoside
- 4:20triphosphates of adenine, thymine,
- 4:23guanine, and cytosine are the
- 4:25substrates that DNA polymerase will
- 4:28polymerize, depending on the template
- 4:30used during replication. And these
- 4:35molecules are found in large excess in
- 4:38the reaction so that it proceeds
- 4:41quickly and effectively. Let's remember
- 4:46that for all DNA polymerases, including
- 4:49Taq polymerase, which is the polymerase
- 4:52of bacterial origin used here, they
- 4:55cannot synthesize de novo. Hm. On one
- 5:01hand, they need a single-stranded DNA
- 5:04template for polymerization, and on the
- 5:07other hand, since they do not
- 5:10synthesize de novo, this means they
- 5:13cannot start a new strand from scratch
- 5:16even with a template; rather, the
- 5:19enzymatic activity is restricted to
- 5:22extending a pre-existing strand
- 5:24previously paired to the template. In
- 5:29particular, the aspect—the chemical
- 5:32group necessary for this polymerase
- 5:35polymerization reaction—is that the
- 5:38last nucleotide incorporated into this
- 5:41pre-paired strand must have a free 3 '
- 5:44hydroxyl end, since this hydroxyl group
- 5:48at the 3' position is the necessary
- 5:51requirement for the next nucleotide to
- 5:55be incorporated in the polymerization
- 5:58reaction, and the new nucleotide will
- 6:01again have this free 3 'hydroxyl end
- 6:04that will allow the incorporation of
- 6:07the subsequent nucleotide. Precisely,
- 6:13primers fulfill this function: being a
- 6:17segment of DNA. Since the primers used
- 6:22during the polymerization that occurs
- 6:25in PCR are made of DNA to be more
- 6:29stable, they provide this previously
- 6:32paired strand that will then be the
- 6:36site for the polymerization reaction.
- 6:41But that is not the only function of
- 6:44primers; in addition to meeting that
- 6:46requirement that allows the
- 6:48functionality of the DNA polymerases in
- 6:51the reaction, primers fulfill another
- 6:54function: determining which region is
- 6:57going to be amplified. That is to say,
- 7:01they also have the function of
- 7:03conferring specificity as to which DNA
- 7:08segment will be amplified in that
- 7:10reaction. This is because in each PCR
- 7:15reaction, at least a pair of primers is
- 7:19used. Hm. One of those primers binds to
- 7:23one of the DNA strands, and the other
- 7:27primer binds to the complementary
- 7:30strand. Hm. When these primers are
- 7:35annealed during the polymerization
- 7:37reaction, they set the limit for the
- 7:40new strands, and therefore, when they
- 7:43are used as templates themselves in the
- 7:46next amplification cycle, they will
- 7:49provide a fixed endpoint for that
- 7:51amplification reaction. The PCR
- 7:58reaction consists of a series of
- 8:00amplification cycles where each of
- 8:03these cycles has three steps. Hm. The
- 8:06first step involves DNA denaturation,
- 8:09bringing the reaction to 95 ° for a
- 8:12short time. This allows genomic DNA,
- 8:17for example, and later all
- 8:21amplification products as the reactions
- 8:25progress, to denature. That is, the
- 8:28double-stranded DNA separates into
- 8:31single strands, since the template must
- 8:33be single-stranded to be a substrate
- 8:35for amplification by DNA polymerase. In
- 8:41the second step, the temperature drops
- 8:44to between 50 and 70 °, depending on
- 8:47the composition of the primers, because
- 8:50this lower temperature allows the
- 8:53primers to hybridize, to bind to the
- 8:56DNA strands. And then in the third step
- 9:01, which typically occurs at 72 degrees,
- 9:04which is the optimal polymerization
- 9:07temperature for Taq polymerase,
- 9:09elongation occurs, the effective
- 9:11polymerization of those fragments. Hm.
- 9:16When these cycles are repeated many
- 9:19times, what ends up being generated is
- 9:22an exponential amplification. And the
- 9:26exponentiality is due to the fact that
- 9:29in each of these cycles, the amount of
- 9:32template of that region to be
- 9:34specifically amplified doubles. So, in
- 9:37each cycle, you have double the amount
- 9:40of templates you had in the previous
- 9:42one. Hm. Since dNTPs, Taq polymerase
- 9:47molecules, and primers are placed in
- 9:50excess inside that tube, the reaction
- 9:54occurs, amplifying exponentially, and
- 9:57after 30 cycles, for example, there
- 10:01will be 1 billion specifically
- 10:03amplified DNA molecules, making the
- 10:07genomic DNA we initially added a
- 10:10minuscule contaminant. compared to the
- 10:14enormous amplification product
- 10:16throughout these cycles. Therefore, it
- 10:19is considered that by the end of the
- 10:23PCR reaction, we have a segment of
- 10:27amplified DNA with billions of copies
- 10:31that is virtually pure. How are these
- 10:36amplification products visualized?
- 10:41Usually, a technique called nucleic
- 10:43acid gel electrophoresis is used.
- 10:46Electrophoresis involves subjecting
- 10:48charged molecules to an electric field
- 10:51so that they migrate based on their
- 10:53charges. As you know, nucleic acid
- 10:56molecules are negatively charged
- 10:59because the phosphate groups that
- 11:02contain their ribonucleotides have a
- 11:06net negative electrical charge and will
- 11:09migrate toward the positive electrode
- 11:13in this reaction of molecule separation
- 11:18. This electrophoresis does not occur
- 11:21in a liquid medium, but in a gel. And
- 11:24this gel, if we were to see it
- 11:27microscopically, is built by a matrix
- 11:30of—it's a polymer that forms a matrix
- 11:33of pores, a porous matrix. And so, when
- 11:39one subjects nucleic acid molecules to
- 11:42migrate in a solution through this
- 11:46porous matrix, what happens is that the
- 11:49smaller fragments can pass through this
- 11:54network of pores more easily than the
- 11:57larger fragments, which will collide a
- 11:59greater number of times against the
- 12:01walls of this gel and will be slowed
- 12:03down in their speed. The final effect,
- 12:07then, is that electrophoresis allows
- 12:10for the separation of DNA fragments
- 12:12according to their size. To visualize
- 12:17the DNA fragments within the gel, a
- 12:21molecule like ethidium bromide is used,
- 12:25which is a base intercalator. It is a
- 12:29molecule of such a size that it can fit
- 12:32in between two successive DNA base
- 12:34pairs. Hm. And one of the properties of
- 12:38this ethidium bromide molecule is that
- 12:40when it is illuminated by ultraviolet
- 12:42light, it emits fluorescence. So, if we
- 12:46introduced molecules of a dye like
- 12:48ethidium bromide into the gel matrix
- 12:51and then developed the gel with
- 12:53ultraviolet light, we will see that
- 12:56they have separated; that is, we will
- 12:58see fluorescent marks where there is
- 13:01DNA. In this photo of this gel, we
- 13:05observe that different samples have
- 13:07been loaded into each of the lanes of
- 13:10this gel. For example, in the first
- 13:13lane and the last lane, molecular
- 13:18weight markers were loaded, that is, a
- 13:21pre-made mixture of different sizes of
- 13:23DNA molecules that then serve to
- 13:26estimate the size of the molecule that
- 13:28was the product of the PCR reaction.
- 13:31Let's imagine that we ran four
- 13:35reactions here: 1, 2, 3, and 4. We can
- 13:40see that in the first three there was
- 13:42an amplification product and we could
- 13:44estimate the size of that amplification
- 13:47product by comparing it with the
- 13:49molecular weight markers, while in the
- 13:52last reaction there was no
- 13:53amplification product at all. Sometimes
- 13:58, these images obtained from developing
- 14:01the gels are often digitally inverted,
- 14:04showing us dark bands on a light
- 14:07background, as this is usually more
- 14:10comfortable for our eyes for
- 14:12visualization; but both types of
- 14:15visualizations are common in research
- 14:18papers and textbooks. I would like to
- 14:24tell you that there are alternatives in
- 14:28electrophoresis to visualize
- 14:30amplification products, and
- 14:32electrophoresis can sometimes be
- 14:35performed in a capillary. Hm. That is,
- 14:39the system is equivalent; there is an
- 14:42electrophoresis in the sense that an
- 14:45electric field is applied and there is
- 14:48a negative pole and a positive pole,
- 14:51but now the gel is it is contained
- 14:56within a small, thin capillary tube. Hm
- 14:59. But it is also a porous matrix. And
- 15:04when DNA molecules are separated by
- 15:07size, using the same principle we just
- 15:10saw, visualization occurs because at a
- 15:13specific point in this capillary there
- 15:16is a detector that will provide a
- 15:19series of fluorescence peaks as the
- 15:22fluorescently labeled DNA molecules
- 15:24pass through, either by a base
- 15:27intercalator or another detection
- 15:30method. And the result of this
- 15:34capillary electrophoresis is usually
- 15:37observed digitally in the form of an
- 15:39electropherogram. Hm. But, analogously,
- 15:44these peaks represent what would be two
- 15:47bands in a gel. Hm. Where the height of
- 15:52the peaks implies the fluorescence
- 15:55intensity of the bands, and what is
- 15:58represented on the X-axis is the
- 16:00separation according to size. In other
- 16:04words, both types of information can be
- 16:07obtained by gel electrophoresis or by
- 16:09capillary electrophoresis. In the first
- 16:13case, it is the visualization of the
- 16:15gel and the interpretation of the
- 16:17banding pattern that is informative,
- 16:19and in the second case, it is the
- 16:20interpretation of the corresponding
- 16:22electropherogram. Here we see another
- 16:29example where we see at the top an
- 16:31electropherogram corresponding to a PCR
- 16:35reaction where there was a single
- 16:38amplification product, a single
- 16:40amplicon. That is the technical name we
- 16:43give it, and that allows us to explain
- 16:45this single peak. On the other hand, in
- 16:48a second example, let's imagine that in
- 16:50another PCR reaction two different
- 16:52amplification products were generated,
- 16:54and this is observed as a series of two
- 16:56peaks. Hm. Next, having seen the
- 17:04foundations of the PCR technique, the
- 17:07reagents a basic reaction of this type
- 17:10entails, and the modes of visualizing
- 17:13PCR products, we will try to understand
- 17:17the uses of this technique in the
- 17:20context of medical genetics. That is to
- 17:25say, thinking that we are going to use
- 17:27this technique to determine the
- 17:29presence or absence of specific allelic
- 17:32variants. Something important to
- 17:36consider is that PCR is a
- 17:38hypothesis-based technique, meaning
- 17:41that if we are using it in a medical or
- 17:45clinical context, it is because we are
- 17:48looking for a particular allelic
- 17:50variant. That is, the clinical context
- 17:54should suggest that we are thinking of
- 17:57looking for allelic variants of a
- 17:59specific gene, and even within the same
- 18:02gene, that we are looking for a certain
- 18:05mutation in particular. Hm. Since the
- 18:10design of the primers in this technique
- 18:13requires having hypothesized about the
- 18:17nature of these causes of our clinical
- 18:20picture. Let's see an example in the
- 18:25context of a monogenic entity such as
- 18:28cystic fibrosis. We won't get into the
- 18:32details of this pathology, uh, but let
- 18:35me tell you in advance that we already
- 18:38know there is a symptomatology in
- 18:41certain boys and girls that allows us
- 18:43to suspect this condition, which is
- 18:46characterized by mucus plugging of
- 18:48multiple ducts. And we also know that
- 18:53it originates from pathogenic variants
- 18:56in a chloride channel gene called CFTR,
- 19:02and that the most frequent mutation in
- 19:05the population, uh, that causes this
- 19:09pathology, is a mutation called
- 19:12phenylalanine 508, where a deletion of
- 19:16three nucleotides in the gene, in an
- 19:19exon, in a strictly coding region,
- 19:22correlates to the deletion of a codon
- 19:26that codes for the amino acid.
- 19:30phenylalanine, and that this absence of
- 19:32this amino acid has negative
- 19:34consequences for protein function. So,
- 19:39if there is a patient who has clinical
- 19:45signs compatible with cystic fibrosis,
- 19:48it could be useful, since it is known
- 19:51that in many patients, uh, this
- 19:54particular allelic variant that
- 19:56involves the deletion of the
- 19:59phenylalanine is the one present, for
- 20:02one to go and look if this
- 20:05phenylalanine deletion is present or
- 20:08absent at this specific position of the
- 20:11CFTR gene. So, in this context, I would
- 20:17like to show you how conventional PCR
- 20:20can be very useful for detecting
- 20:23insertions or deletions as in this case
- 20:26and guided by a clinical hypothesis.
- 20:32For that, since we know which gene we
- 20:37want to explore to see if it contains
- 20:40this insertion or deletion, we can
- 20:43design a pair of primers, because we
- 20:46know the human genome sequence and can
- 20:49then access it via databases and use it
- 20:53to, uh, design our primers; and we will
- 20:56design them in a region flanking the
- 20:59position where the deletion may or may
- 21:02not be. Here in this example, we are,
- 21:06uh, sketching at the top an allelic
- 21:10variant that does not contain the
- 21:13deletion in question and another that
- 21:16does. Note that since the primers are
- 21:20in a flanking region, in both cases, in
- 21:22both allelic variants, the primers can
- 21:25bind and will be able to generate
- 21:27amplicons, that is, amplification
- 21:29reactions in both cases. Hm. But the
- 21:34amplification products will be slightly
- 21:36different, since if we compare them
- 21:38with each other, the amplification
- 21:40product of the variant that does not
- 21:42contain the deletion will have three
- 21:45nucleotides more than the variant that
- 21:47does possess the deletion. Although
- 21:51this difference of three nucleotides is
- 21:54small, polyacrylamide gels, which are
- 21:59different from agarose gels, can be
- 22:01used to discriminate these differences.
- 22:04Hm. That is to say, in a polyacrylamide
- 22:08gel matrix, electrophoresis, which will
- 22:10separate the molecules according to
- 22:12their size, can occur in the same way.
- 22:16And let's analyze then Uh, the result
- 22:21of this PCR, imagining that we obtained
- 22:24biological samples, perhaps through
- 22:26blood, peripheral blood samples, from
- 22:29members of a family where some of them
- 22:32showed clinical signs of cystic
- 22:35fibrosis and we wanted to explore the
- 22:38presence or absence of this particular
- 22:41deletion in the CFTR gene. Let's focus,
- 22:46for example, on individual number four.
- 22:52Hm. Individual number four has a single
- 22:57amplification product. Hm. And if we
- 23:01compare it with individual number two,
- 23:04which also has a single, uh, amplicon
- 23:06fragment, an amplification product, if
- 23:08we compare them with each other, we can
- 23:11say that the amplification product of
- 23:13individual four is larger in base pairs
- 23:16than the amplification product of two;
- 23:18because, as we had said, the smaller
- 23:20fragments migrate further in the gel,
- 23:23they can pass through the porous matrix
- 23:26with less difficulty than larger
- 23:28fragments. So, considering, we would
- 23:33now have to make considerations
- 23:35regarding the structure of the human
- 23:37genome to know what has happened there.
- 23:42We have to consider, for example, the
- 23:44diploid nature of our genome, that we
- 23:46have two allelic variants, one of
- 23:48maternal origin and another of paternal
- 23:50origin for each of the genes. If we
- 23:53have, in both individual two and
- 23:55individual four, a single amplification
- 23:57product, we will say that the maternal
- 24:01and paternal variants are behaving in
- 24:03the same way, they are generating the
- 24:05same type of amplification product. Hm.
- 24:07Now then, how to understand a larger
- 24:11amplification product versus a smaller
- 24:15one? In the context of this PCR, where
- 24:18we use primers to amplify a certain
- 24:20segment of the CFTR gene, one could say
- 24:23that individual 4, since it has the
- 24:28larger sizes, its two allelic variants,
- 24:30the maternal and the paternal, were
- 24:32allelic variants without the
- 24:34phenylalanine deletion. 508, while
- 24:37individual 2 has both of its allelic
- 24:41variants with the deletion. That is why
- 24:44we see a single amplification product
- 24:46resulting from those two allelic
- 24:50variants that were used as a template
- 24:52during the PCR, but which generated the
- 24:54same amplification product. In contrast
- 24:58to these cases, individuals 1 and 3
- 25:01have two amplification products, one
- 25:04larger and one smaller. Hm. And so we
- 25:08can think that in those individuals,
- 25:10one of the allelic variants has the
- 25:13deletion and the other does not. Hm. In
- 25:15other words, the deletion would be in
- 25:18heterozygosis in these two individuals.
- 25:25With what we have seen so far, we could
- 25:28think that since insertions and
- 25:30deletions can generate amplicons, that
- 25:33is, amplification products of different
- 25:36sizes that will later be separated by
- 25:39gel electrophoresis, we can understand
- 25:42why PCR can be used to detect
- 25:45insertions and deletions. Now, what
- 25:48would happen if we want to detect other
- 25:51types of allelic variants? For example,
- 25:56those caused by a substitution mutation
- 25:59where there will be no size differences
- 26:03between the amplification products. To
- 26:09detect allelic variants caused by
- 26:12substitution mutations, you must use
- 26:16variations of the PCR technique
- 26:19specifically designed to detect these
- 26:22substitutions. And one of these
- 26:27variations is allele-specific PCR,
- 26:30which is specifically designed for this
- 26:34purpose. What is the design feature
- 26:39that allows this? Imagine we have two
- 26:42allelic variants that differ by a point
- 26:46mutation, meaning that in the same
- 26:49region where one allelic variant has an
- 26:52AT pair, another allelic variant at
- 26:56that same position of the same gene has
- 26:59a CG base pair. The key here is to
- 27:05design a pair of primers where one of
- 27:09them binds specifically at its 3'
- 27:11hydroxyl end, which must necessarily be
- 27:15base-paired to the template for
- 27:17amplification to occur. We have to
- 27:22design that primer to be
- 27:24allele-specific as well. In other words
- 27:28, if we design it so that it can only
- 27:30bind to one of the allelic variants,
- 27:32there will be a mismatch at this point
- 27:34in the other allelic variant. And this
- 27:38implies that even if the other primer
- 27:41is bound, one of the allelic variants
- 27:43will not be able to amplify during the
- 27:46PCR reaction, while the other will. Hmm
- 27:49. So, depending on whether there is an
- 27:53amplification product or not, one could
- 27:56infer which of the two allelic variants
- 27:58was present. For example, if we take
- 28:05samples from five patients to determine
- 28:08which of the two allelic variants they
- 28:11possessed and we have these results,
- 28:13here we see a lane as a result of the
- 28:16gel electrophoresis revealing a
- 28:18molecular weight marker; but the PCR of
- 28:21patient 1, and of patient 2 and 5, give
- 28:24us an amplification product. We can say
- 28:29that they had at least one CG-type
- 28:31allelic variant in their genome. Now,
- 28:36what we can also say is that
- 28:38individuals 3 and 4 apparently did not
- 28:40have any CG variant in their genome.
- 28:44Usually, these allele-specific PCRs
- 28:47need to perform another PCR to detect
- 28:50the other variant to have complete
- 28:53information. Hmm. For example, if we
- 28:58use primers with the same patients '
- 29:00genomic DNA to detect the other variant
- 29:03, that is, primers that can bind when
- 29:06the AT dinucleotide is present, but not
- 29:09the CG dinucleotide, we complete the
- 29:12information here. For instance, we see
- 29:16that individuals 3 and 4, which
- 29:17previously showed no amplification
- 29:19signal, now show one. Then we can say
- 29:23that the two allelic variants, both
- 29:25maternal and paternal in origin, of
- 29:28individuals 3 and 4 had an AT allelic
- 29:30variant and did not have any CG. But
- 29:35what can we say, for example, about
- 29:38individual 1, who has an amplification
- 29:41product in both one PCR and the other?
- 29:45In this case, this individual is
- 29:47probably heterozygous, meaning they
- 29:49have one allelic variant of parental
- 29:52origin of the CG type and another AT
- 29:54variant. Therefore, there was an
- 29:56amplification product in both cases. Hm
- 29:59. I leave it to you to think about what
- 30:03the genetic configuration will be in
- 30:04terms of these variants for individual
- 30:075, who presents amplification here and
- 30:11not here. Let me tell you that there is
- 30:18also another variant of the PCR
- 30:21technique called PCR-RFLP which is used
- 30:28to detect allelic variants with point
- 30:32substitution mutations, that is, that
- 30:35differ from one another by a
- 30:38substitution mutation. This acronym
- 30:44RFLP stands for restriction fragment
- 30:50length polymorphisms. And this refers
- 30:55to the use in the research context of
- 30:58endonucleases, called restriction
- 31:01enzymes, which are bacterial
- 31:03endonucleases that can cleave DNA at
- 31:06specific sequences. In this case, if
- 31:11one wants to differentiate two allelic
- 31:13variants, for example, let's see that
- 31:15this variant differs at this point. Hm.
- 31:18Here there is AT and here there is TA.
- 31:22That is to say, there is a substitution
- 31:24at a particular point. Whoever designs
- 31:29this technique has to infer if in the
- 31:32context where the variation one wants
- 31:34to detect is located, it contains any
- 31:37recognition sequence for a restriction
- 31:40enzyme. If one goes to the databases
- 31:44and compares the recognition sequences
- 31:47of restriction enzymes and finds that
- 31:50this variation affects one of these
- 31:52sites, one could say the following. For
- 31:56example, in this case there is a
- 31:57restriction enzyme which is D1 that can
- 32:02cut, that is, it has a recognition site
- 32:06in the allelic variant that had AT at
- 32:09this position, but when that nucleotide
- 32:12is substituted by its inverse TA, the
- 32:15restriction enzyme can no longer cut.
- 32:21Realizing this aspect allows us to
- 32:24design a PCR in combination with a
- 32:26cleavage reaction using this
- 32:28restriction enzyme, which will
- 32:30eventually allow us—and this is our
- 32:32goal—to distinguish between these
- 32:34allelic variants. In particular, to
- 32:38determine which allelic variants an
- 32:40individual has, um, that we are
- 32:44analyzing. For that, we simply design
- 32:49primers. that flank the restriction
- 32:52site where the variant we are looking
- 32:54for is located. And since these primers
- 32:57do not affect or bind to the
- 32:59potentially polymorphic region, they
- 33:02bind to both one allelic variant and
- 33:05the other. When amplifying by PCR, in
- 33:09both cases, we will obtain
- 33:11amplification products of the same size
- 33:15. And as a second characteristic step
- 33:18of this technique, we treat the PCR
- 33:20amplification products with the
- 33:22restriction enzyme. So, let's imagine
- 33:27what happens in the first step. The two
- 33:32allelic variants are amplified by PCR.
- 33:35We have millions and millions of copies
- 33:37of these amplification products in each
- 33:40PCR tube, which are equal in size. Now,
- 33:43when we treat them with the restriction
- 33:47enzyme, only one of the fragments can
- 33:50be cut and the other will remain uncut
- 33:53because the allelic variant precisely
- 33:56affected this enzyme's restriction site
- 33:59. And so now we have a polymorphism in
- 34:02the length of the fragments. Hm. Hence
- 34:06the name of the technique, and we know
- 34:09that DNA fragments of different sizes
- 34:11can be separated using the
- 34:13electrophoresis technique. Let's
- 34:18interpret once again five patients
- 34:20whose genetic makeup we have explored
- 34:23thanks to this technique. Let us recall
- 34:28, then, that the allelic variant we
- 34:30call AT here was the one that retained
- 34:33the restriction site and could be
- 34:35cleaved, generating these two fragments
- 34:38, while the TA variant was not going to
- 34:40be a substrate for this restriction
- 34:43reaction and would generate a longer
- 34:45fragment. For example, if we think of
- 34:50patient number four, who has a single
- 34:52large-sized amplification product, Hm.
- 34:56because it is the one that migrates the
- 34:57least. We can say that individual four
- 35:00had a maternal and paternal allelic
- 35:02variant with the TA-type variant, since
- 35:05the largest fragments are the ones that
- 35:08are not substrates for the restriction
- 35:10enzyme. Hm. On the other hand, if we
- 35:15think of individual one, where we have
- 35:18two restriction fragments and both are
- 35:21smaller in size than the one we had
- 35:23found in individual four, we can think
- 35:26that this individual only had AT-type
- 35:29allelic variants that generated these
- 35:32two fragments that are different from
- 35:35each other, but in both cases smaller
- 35:38than the other allelic variant. Hm. And
- 35:43the most complex case to interpret is
- 35:45found in individuals 3 and 5, where we
- 35:48see three amplification products on the
- 35:51gel. And in this case, we interpret
- 35:56that these individuals had a single TA
- 35:58allelic variant, which was not cut by
- 36:01the restriction enzyme and generated
- 36:03the larger fragment. And the other
- 36:07allelic variant was of the AT type,
- 36:09which was fragmented and generated the
- 36:12smaller fragments. I would now like to
- 36:20tell you about another application of
- 36:22PCR, which is multiplex PCR, in a very
- 36:27particular context: the context of
- 36:29forensic genetics, where, in this case,
- 36:32it is not a pathological context, but
- 36:35rather an attempt to determine genetic
- 36:37variants that allow for the
- 36:39establishment of biological
- 36:41relationships between individuals, for
- 36:44example, filiation relationships. Hm.
- 36:48And in this context, usually what is
- 36:52explored are highly polymorphic regions
- 36:56of our genome, which are short tandem
- 36:59repeats, or STRs, usually represented
- 37:03by elements called microsatellites that
- 37:07we analyzed in seminar number nine.
- 37:13These, uh, these microsatellites, these
- 37:17short tandem repeats, behave like
- 37:19allelic variants, that is to say, all
- 37:22human beings at the same position of
- 37:25certain chromosomes possess a variant
- 37:29of maternal origin and a variant of
- 37:31paternal origin of these repetitive
- 37:34regions. And since they are highly
- 37:38polymorphic, they differ from each
- 37:40other very frequently in the exact
- 37:42number of repeats that each one
- 37:44possesses. Hm. And since each one of us
- 37:50received these variants, one of
- 37:52maternal origin and one of paternal
- 37:54origin, if one explores, for example,
- 37:57for one, uh, for one of these
- 37:59microsatellites in particular, we
- 38:01should probably find one of those
- 38:04allelic variants in our biological
- 38:06mother and the other allelic variant we
- 38:09should find in our biological father.
- 38:15To increase the fidelity of these
- 38:18analyses, for example, of filiation, a
- 38:21single microsatellite is not analyzed,
- 38:25but rather, uh, in the year 1997, the
- 38:30FBI standardized a procedure that
- 38:35analyzes 13 sequences, uh, of mini
- 38:40microsatellites scattered throughout
- 38:43different human chromosomes. Therefore,
- 38:47the inferences will not only be based
- 38:50on a single microsatellite, but on
- 38:52multiple, uh, loci, multiple
- 38:54chromosomal positions that are on
- 38:56different chromosomes. These names that
- 39:01we see here, TPOX, D8S 1179, VWA, etc.,
- 39:05are the specific names of tandem repeat
- 39:12regions, which are highly polymorphic
- 39:15in the human population, but which all
- 39:18humans have at those positions on those
- 39:21chromosomes. So, I wanted to tell you
- 39:28that this multiplex PCR technique is
- 39:31applied to analyze the length of these
- 39:35regions, to determine the number of
- 39:39repetitions that each of these
- 39:42microsatellites has in a single
- 39:45reaction at the same time. What is the
- 39:51key? What specific aspect of multiplex
- 39:56PCR, which, unlike classical PCRs where
- 39:59we used a pair of primers to amplify
- 40:02those regions of the genome that we
- 40:05wanted to amplify. In this case, we are
- 40:10going to use many different pairs of
- 40:13primers, each one of them designed to
- 40:15amplify one of these microsatellites in
- 40:18particular, that is, the flanking
- 40:21region that allows amplifying that
- 40:23microsatellite which is on each of the
- 40:26chromosomes we analyze. Precisely
- 40:31because this reaction is complex,
- 40:34because amplicons from different genes,
- 40:37from different pairs of primers, are
- 40:40going to be mixed in the same tube,
- 40:43fluorophores are usually used to
- 40:45visualize the amplification products,
- 40:49molecules linked to the primers, that
- 40:52will emit fluorescence in different
- 40:55colors, thus allowing the different PCR
- 40:58products to be differentiated by color.
- 41:03Since this reaction is very complex,
- 41:05because we have at least 13 different
- 41:08amplification products, what is done is
- 41:11to use the same fluorophores for some
- 41:14genes that have amplicons of different
- 41:17sizes. So that by combining these size
- 41:21differences with differences in
- 41:23fluorophores, we can obtain the total
- 41:25information. Usually, this multiplex
- 41:30PCR is combined with capillary
- 41:32electrophoresis to then detect the
- 41:35different products as we saw. And in
- 41:40the detector, the fluorescences coming
- 41:42from the different colors that we used
- 41:45in the fluorophores of the different
- 41:47pairs of primers can be detected. I
- 41:51show you an example here in this
- 41:54electropherogram. As a result of having
- 41:59used multiple pairs of primers with
- 42:01different fluorophores, we obtain a
- 42:04series of peaks with different colors.
- 42:09Let's remember that on the vertical
- 42:11axis, the height of the peaks implies
- 42:13the fluorescence intensity of each of
- 42:15those amplification products, and on
- 42:17the horizontal axis, these peaks are
- 42:20separated based on the size of the
- 42:22amplicon. Hm. Digitally, one can
- 42:26separate these signals of different
- 42:27colors to perform a more detailed
- 42:29analysis. For example, here we see how
- 42:33we separate the channel of those
- 42:35amplification products that had a blue
- 42:38fluorophore in the upper part and in
- 42:40the lower part the channel of those
- 42:43green fluorophores. Note that because
- 42:49of how the primers were designed, one
- 42:50can know in which region we will obtain
- 42:55the amplification products of a
- 42:56particular microsatellite. For example,
- 42:59for this particular microsatellite we
- 43:01have two amplification products. We see
- 43:03two peaks. In this case, for this
- 43:06particular microsatellite, this
- 43:08individual we analyzed had a
- 43:10microsatellite with 12 repetitions and
- 43:12another with 13 of this one in
- 43:15particular. Conversely, for this other
- 43:18microsatellite from the same individual
- 43:21, we have a single peak. What does this
- 43:23mean? That the maternal and paternal
- 43:26allelic variants had the same number of
- 43:28repeats. Hm. So, by combining the
- 43:31information from these multiple primer
- 43:34pairs associated with different
- 43:36fluorophores, one can, uh, know the
- 43:41exact number of repeats an individual
- 43:44has for each of these tandem repeat
- 43:47regions and compare it, then, with
- 43:51potential, uh, likely subjects, uh,
- 43:54biological fathers or mothers, and
- 43:58determine the probability of whether or
- 44:01not a biological affiliation exists
- 44:04between them. Finally, I would like to
- 44:12tell you about this last variant of the
- 44:15PCR technique we are going to analyze
- 44:19today, which is TP-PCR, or Triplet
- 44:22Primed PCR, used in the genetic
- 44:24diagnosis of triplet repeat expansion
- 44:28diseases. This is a set of
- 44:31non-classical monogenic diseases that
- 44:34will be analyzed in depth. in an
- 44:37upcoming seminar, but they are
- 44:40characterized by the expansion of a
- 44:42three-nucleotide repeat within the
- 44:45affected gene. So, in the population,
- 44:49there are allelic variants that have a
- 44:52normal range of repeats, meaning they
- 44:55are non-pathogenic. And these, for
- 44:59example, in this particular case,
- 45:01unaffected individuals have between 5
- 45:03and 50 repeats of this trinucleotide in
- 45:05their maternal and/or paternal allelic
- 45:08variant. A characteristic of this
- 45:13disease is that some individuals have
- 45:15an allelic variant called a premutation
- 45:18with a higher number of repeats. And
- 45:21the premutation is characterized by the
- 45:24fact that those who carry this
- 45:27premutation do not show signs and
- 45:29symptoms of the pathological entity,
- 45:32but there is a high probability that in
- 45:35their gametes, an expansion of the
- 45:37mutation occurs, h, of a higher range
- 45:40that will then be associated, in the
- 45:43next generation, with the appearance of
- 45:46signs and symptoms. As you will see,
- 45:50this technique, what it tries to
- 45:52determine in the context where there is
- 45:55a hypothesis that there is a triplet
- 45:57expansion disease behind the signs and
- 46:00symptoms of some family members, is to
- 46:02determine if an individual is a carrier
- 46:05of allelic variants with expansions
- 46:07within the normal range, within the
- 46:09premutation range, or the pathogenic
- 46:12mutation range. And you might ask
- 46:16yourselves, why couldn't conventional
- 46:18PCR discriminate this? Since we have
- 46:21seen that insertions, for example, that
- 46:23variations in length could be detected
- 46:25by the PCR technique. But although from
- 46:31a theoretical point of view that is
- 46:33true, there are specific problems that
- 46:36arise for conventional PCR when it
- 46:39comes to detecting, uh, triplet
- 46:41expansions. In this case, we see the
- 46:46FMR1 gene as an example, which is one
- 46:50that can have this huge triplet
- 46:53expansion in its 5-prime UTR region. We
- 47:01mentioned that the limitation of
- 47:04conventional PCR when detecting
- 47:07amplification products of these
- 47:09expanded alleles is due to the fact
- 47:12that these regions containing multiple
- 47:15repeats are often a difficult context
- 47:18for the DNA polymerase used in these
- 47:21reactions to effectively amplify the
- 47:24entire fragment. And this causes a
- 47:29series of errors, a series of
- 47:31incomplete fragments of the potentially
- 47:33expanded fragment to be amplified,
- 47:36which results in the template not being
- 47:39able to be effectively duplicated from
- 47:41cycle to cycle, and towards the end of
- 47:44the PCR, we observe no amplification
- 47:47product if there was an expanded allele
- 47:50in the template. Simply put, the PCR
- 47:54fails and fails to amplify. When we
- 47:59only use a pair of primers flanking the
- 48:02potentially expanded region, it fails
- 48:05to amplify that fragment. Hm. So, if a
- 48:10diploid individual has an allelic
- 48:12variant from a parental origin without
- 48:15amplification, this will be amplified.
- 48:18But if it has an expanded allele as the
- 48:21other allelic variant, there will be no
- 48:24amplification product, and since only
- 48:27the amplification product of the
- 48:29non-expanded variant will be seen on
- 48:32the gel, there is a masked result.
- 48:36These problems were originally solved
- 48:39with a different technique, which was
- 48:41the Southern blot technique, which we
- 48:43will not cover here because it is a
- 48:46technique that is falling into disuse,
- 48:48but variations of the PCR technique
- 48:50allowed the problem to be solved. To
- 48:54avoid the problems of lack of
- 48:57amplification that occur with triplet
- 49:00expansion when we try to determine it
- 49:02with conventional PCR, this variant of
- 49:05PCR uses three primers. One of them,
- 49:10acting as a forward primer, that is, on
- 49:13one of the strands, is placed in the
- 49:15flanking region of the amplification.
- 49:19Then a second primer is used that will
- 49:25be complementary to the repeat region
- 49:27and a third primer that will function
- 49:30as a reverse primer, that is, in the
- 49:32opposite direction, which will bind to
- 49:35the complementary strand. Let us
- 49:39consider that the use of these three
- 49:41primers, in particular, the primer that
- 49:44binds to the repeat regions, has the
- 49:47following peculiarity. This primer can
- 49:51bind anywhere in the repeat region,
- 49:54generating diverse amplification
- 49:57products when it creates amplicons in
- 50:00combination with the primer facing in
- 50:03the opposite direction. So, let's think
- 50:08that in the PCR reaction we have
- 50:12millions of molecules acting as
- 50:14templates and many reactions occurring
- 50:17simultaneously. Consequently, one would
- 50:21expect from this configuration that,
- 50:24for example, amplification products are
- 50:27generated, uh, of different lengths, as
- 50:33we were just saying, when the repeat
- 50:35primer acts, and some of full length
- 50:37when both flanking primers act.
- 50:41Consequently, uh, this technique has
- 50:44the advantage that it will allow for
- 50:47estimating the presence or absence of
- 50:50these expanded allelic variants, uh, by
- 50:54two methodologies. Hm. On one hand, uh,
- 50:58we will try to see if there is
- 51:00amplification of the full expansion
- 51:03product, but we will have, uh, as an
- 51:06aid in this technique, the presence of
- 51:09a series of fragments of different
- 51:12sizes that span the different regions
- 51:15of the repeat. TP-PCR results are
- 51:20usually analyzed by capillary
- 51:23electrophoresis. Let us remember that,
- 51:28uh, these capillary electrophoresis
- 51:30results are usually visualized in these
- 51:33graphs called electropherograms, where
- 51:36on the vertical axis we observe that
- 51:39the height of the peaks, uh, reflects
- 51:41the level of signal fluorescence
- 51:44intensity, uh, the presence of
- 51:46amplification products and their
- 51:48intensity. And on the horizontal axis,
- 51:52what we see is the variation in the
- 51:55sizes of those amplification products.
- 51:58In this case, uh, we are observing
- 52:01different length ranges of those
- 52:03products that allow us to classify
- 52:06these allelic variants into the normal
- 52:10zone, premutation, or fully expanded
- 52:13mutation. Let us analyze how the
- 52:17results look in the case of an allelic
- 52:19variant that possesses trinucleotide
- 52:22repeats in the normal range. Here we
- 52:25see, on one hand, an intense
- 52:27amplification product of the allelic
- 52:30variant in the normal range, which is
- 52:33the amplification product of the two
- 52:35pairs of primers that are in the
- 52:38flanking regions. Hm. And we also see a
- 52:42series of peaks in the form of a ladder
- 52:44, which are those produced by the
- 52:46amplification products where the primer
- 52:49acts that can bind to the repeat region
- 52:51and, therefore, can yield different
- 52:53amplification products. But this entire
- 52:58ladder of peaks is also restricted to
- 53:00the normal range because the range of
- 53:03repeats that these alleles possess is
- 53:05relatively limited. How is it observed?
- 53:10the the result of the electropherogram
- 53:14when there is an individual who
- 53:16possesses a premutated variant, that is
- 53:18, with an intermediate range of triplet
- 53:21expansion. In the first place, we see
- 53:24that the series of peaks, uh, that is
- 53:27produced as a consequence of the
- 53:29different fragments amplified from the
- 53:32primer that binds to the repeat region
- 53:34is wider. Yes, we have a larger range
- 53:38of amplification of those products and,
- 53:41in addition, we see a stronger signal,
- 53:44uh, of the larger amplification product
- 53:46, the one produced by those primers
- 53:49that flank the repeat region, which is
- 53:52centered at a higher molecular weight.
- 53:56Hm. Here we don't see a single peak
- 53:59because usually, the longer the
- 54:01fragment to be proliferated, the
- 54:03greater the number of errors that occur
- 54:06in the PCR products. Hm. We also see
- 54:11that the peak height has decreased
- 54:13compared to the previous time,
- 54:14precisely because that amplification
- 54:16product is less efficient and the total
- 54:18amount of product we will obtain will
- 54:20be smaller, and its fluorescence
- 54:21intensity, the peak height, will also
- 54:23be lower. Finally, let's compare this
- 54:28with a carrier of an allelic variant
- 54:31with full expansion, where we see that
- 54:34the larger the expanded region, the
- 54:37region of peaks—hm—produced as a
- 54:40product of the amplification of the
- 54:42primer that bound to the repeat region
- 54:45covers a much wider range of molecular
- 54:48weights, and we see that the peak
- 54:52corresponding to the amplification of
- 54:55the full region has a lower height.
- 54:58This means that the efficiency of its
- 55:00production is lower. It was produced in
- 55:03a smaller quantity, but the size of
- 55:06that total product is larger than
- 55:09before and is in the full mutation
- 55:12region. In summary, this TP-PCR
- 55:15technique is a specific adaptation of
- 55:19the PCR technique to detect allelic
- 55:22variants due to triplet expansion. In
- 55:28this second part of the lecture, we are
- 55:30going to study the techniques based on
- 55:33DNA sequencing that are used in the
- 55:35context of medical genetics. This set
- 55:39of techniques focuses on determining
- 55:42the order of base pairs that are
- 55:45present in a given DNA fragment. And
- 55:48within this set of methodologies, we
- 55:51will analyze, first, the sequencing
- 55:54method developed by Frederik Sanger,
- 55:57who won the Nobel Prize in Chemistry in
- 56:001980 thanks to this development, and
- 56:03which, as we will see, was the
- 56:05methodology used to produce the first
- 56:08sequencing of the human genome in the
- 56:11context of the Human Genome Project.
- 56:16Then we will address subsequent
- 56:18developments from the last 15 to 20
- 56:20years that, based on Sanger sequencing
- 56:22methods, allowed it to be expanded and
- 56:25are collectively called high-throughput
- 56:27sequencing methodologies, also known as
- 56:30next-generation sequencing or
- 56:32high-throughput sequencing. To
- 56:36understand the Sanger sequencing
- 56:39methodology, we have to re-analyze some
- 56:44molecular mechanisms that are involved
- 56:46in the polymerization of nucleic acids,
- 56:49since it is the particular biochemistry
- 56:52of polymerization that Sanger used as
- 56:55the foundation for his technique. Let's
- 57:00remember again that during the nucleic
- 57:03acid polymerization process, DNA
- 57:06polymerase uses a single-stranded DNA
- 57:09strand as a template, but it cannot
- 57:12polymerize de novo; it can only extend
- 57:15a previously paired strand that has a
- 57:18free three-prime hydroxyl end. We see
- 57:24on the right side of the slide this
- 57:27three-prime hydroxyl end in greater
- 57:30molecular detail. We are seeing that
- 57:33the last nucleotide added to a chain at
- 57:36the three-prime position of the sugar
- 57:39has a hydroxyl group that will be used
- 57:42to form a new covalent bond with the
- 57:44incoming nucleotide, forming a
- 57:47phosphodiester bond and allowing the
- 57:49incorporation of the new nucleotide.
- 57:54Then, the one that will function as the
- 57:56next new terminal nucleotide of the
- 57:59chain will in turn have another
- 58:01three-prime hydroxyl end that will
- 58:03function in a similar way, that is,
- 58:06allowing the incorporation of the
- 58:08subsequent nucleotide. Taking this into
- 58:13account, Sanger devised using variants
- 58:18of nucleotides that would prevent the
- 58:20development of polymerization. Here we
- 58:25see a standard nucleotide that has the
- 58:29hydroxyl group associated with the 3-
- 58:32prime carbon of the sugar and a variant
- 58:35that was what he used, called
- 58:38dideoxyribonucleotide triphosphates,
- 58:41which lack this hydroxyl group at the 3
- 58:45-prime position. If one, in an in vitro
- 58:50polymerization reaction, adds these
- 58:52dideoxynucleotides, that is,
- 58:54nucleotides that lack this 3-prime
- 58:57hydroxyl end, once they are
- 58:59incorporated, they terminate
- 59:01polymerization because no nucleotide
- 59:03can be incorporated by binding to them,
- 59:06since they lack this functional group
- 59:09that allowed the incorporation of the
- 59:12new nucleotide. How then did this
- 59:16particular biochemistry allow Sanger to
- 59:19devise a methodology for the sequencing
- 59:21of nucleic acids? Let's think about a
- 59:27polymerization reaction that occurs in
- 59:31a laboratory tube. In it, the
- 59:36substrates for polymerization are
- 59:39introduced: a DNA polymerase, the dNTPs
- 59:42, that is, deoxyribonucleotides of
- 59:45adenine, guanine, cytosine, and thymine
- 59:49, and small amounts of
- 59:51dideoxynucleotides of the four classes.
- 59:55Hm. Furthermore, let's think that these
- 1:00:00dideoxynucleotides, which are in much
- 1:00:02smaller amounts than the rest of the
- 1:00:04natural nucleotides, will each be
- 1:00:06labeled with a different fluorophore.
- 1:00:11On the other hand, the template is
- 1:00:14added, which will be a DNA fragment
- 1:00:16whose base sequence one wants to
- 1:00:18determine. You might think that
- 1:00:23information about the sequence of this
- 1:00:27fragment will be necessary to design a
- 1:00:30primer needed to introduce as a
- 1:00:32condition of possibility for the DNA
- 1:00:35polymerase to function. A biochemical
- 1:00:42trick that was used in this context was
- 1:00:45to chemically add to the end of the DNA
- 1:00:48molecule that one wanted to sequence,
- 1:00:52that is, whose nucleotide sequence was
- 1:00:55still unknown, a DNA fragment that we
- 1:00:58are representing here with this green
- 1:01:01line of known sequence. When this known
- 1:01:06sequence fragment is chemically
- 1:01:08attached, one will be able to design a
- 1:01:11complementary primer to this known
- 1:01:13sequence, which will then allow it to
- 1:01:15act as an initiator for the
- 1:01:17polymerization of the region that is
- 1:01:19still unknown. I am not representing it
- 1:01:23here, but let's imagine then that a
- 1:01:25primer complementary to this known
- 1:01:27region, designed by the researcher, is
- 1:01:30also added to this polymerization
- 1:01:32reaction. Let's think about this
- 1:01:36reaction where, although we are
- 1:01:39representing the strand as
- 1:01:41single-stranded, it is originally
- 1:01:44double-stranded, but we will only use a
- 1:01:48single primer from the PCR that binds
- 1:01:52to the strand to be used as a template.
- 1:01:55Hm. Let's imagine as a second step, to
- 1:02:01make the representation of what is
- 1:02:03happening here more complex, that this
- 1:02:06reaction is occurring millions of times
- 1:02:08inside this tube in parallel. What does
- 1:02:11this mean? That there isn't just one
- 1:02:14DNA polymerase and one set of these
- 1:02:16nucleotides, but rather these
- 1:02:18components are there by the millions,
- 1:02:20and the template and primers are also
- 1:02:23there by the millions. So, let's
- 1:02:26imagine that the polymerization
- 1:02:31reaction begins, and then the primers
- 1:02:33—which we are not representing here
- 1:02:36—bind to the known region that was
- 1:02:39attached to one end of the molecule,
- 1:02:42and polymerization begins. Since most
- 1:02:46of the natural components are in excess
- 1:02:49, it is most likely that in the
- 1:02:50majority of the millions of molecules
- 1:02:52we have as a template, polymerization
- 1:02:54will occur completely. But in some of
- 1:02:58the molecules inside that tube, by
- 1:03:01chance, the DNA polymerase might place
- 1:03:04a dideoxy right at the first nucleotide
- 1:03:07to be polymerized. And for that
- 1:03:11molecule which was polymerized with
- 1:03:13that dideoxy, the polymerization of
- 1:03:15that molecule will end there. It will
- 1:03:17not be able to incorporate any more
- 1:03:18nucleotides. The same will happen with
- 1:03:22the second position. Most of the
- 1:03:25molecules continue their polymerization
- 1:03:27, but there will be a small subset that
- 1:03:30will, by chance, incorporate a dideoxy
- 1:03:33nucleotide at the second position of
- 1:03:35polymerization that is complementary to
- 1:03:38the template, and it will be
- 1:03:40polymerized, ending the polymerization.
- 1:03:45This will then occur consecutively at
- 1:03:47each of the polymerization positions,
- 1:03:50and when this reaction ends, we will
- 1:03:52have a complex mixture in that tube of
- 1:03:54fully polymerized molecules; that is,
- 1:03:57those that did not incorporate any
- 1:03:59dideoxy, but there will surely be a
- 1:04:01subset of molecules that have ended at
- 1:04:04every possible position within the
- 1:04:06polymerization. And this happens
- 1:04:10because there are millions and millions
- 1:04:12of molecules that participated in the
- 1:04:14polymerization. And so, statistically,
- 1:04:17it is most likely that at every
- 1:04:20position there is at least a subset of
- 1:04:23molecules represented. What does this
- 1:04:26mean? That all these molecules in which
- 1:04:30polymerization was interrupted at
- 1:04:32different parts will have different
- 1:04:34lengths. Hm. They will vary in length
- 1:04:38one by one. And we know that a
- 1:04:41technique to separate DNA molecules
- 1:04:44according to their size is
- 1:04:45electrophoresis. Hm. Electrophoresis in
- 1:04:49a porous matrix, for example, in an
- 1:04:51agarose gel or in a capillary that has
- 1:04:54a porous matrix. And by separating
- 1:04:59these fragments according to their size
- 1:05:02, each of the fragments, we will be
- 1:05:04able to identify which nucleotide was
- 1:05:06at the end, because each one of them
- 1:05:09emitted a fluorescent light depending
- 1:05:11on the fluorophore that was used,
- 1:05:13marking that so that, by analyzing them
- 1:05:18according to their size, the, uh, the
- 1:05:22light emitted by each of these
- 1:05:24fluorophores will be able to
- 1:05:27reconstruct the sequence of nucleotides
- 1:05:30that made up that fragment. This
- 1:05:34procedure is usually performed using
- 1:05:37not a common gel electrophoresis, but a
- 1:05:41capillary electrophoresis, just as we
- 1:05:44analyzed in the first part of the class
- 1:05:47. And the results are observed as these
- 1:05:52electropherograms, where we see a
- 1:05:56series of fluorescence peaks, and we
- 1:05:59now know that each position implies a
- 1:06:03fragment that differs by one nucleotide
- 1:06:06in length from the other. And the type
- 1:06:11of fluorescent light observed at each
- 1:06:14of these positions can be digitally
- 1:06:17decoded with the nature of one of the
- 1:06:21nucleotides, since the different
- 1:06:23fluorophores report each of the four
- 1:06:26DNA nucleotides. And in this way, we
- 1:06:30can digitally obtain, after this
- 1:06:33analysis of capillary electrophoresis,
- 1:06:36the nucleotide sequence, uh, of the
- 1:06:40fragment that was analyzed by Sanger
- 1:06:43sequencing. Let's observe this
- 1:06:47particularity. Let's think that we have
- 1:06:50obtained three different results
- 1:06:52through Sanger sequencing of a
- 1:06:54particular gene from three individuals.
- 1:06:59These three individuals have high
- 1:07:02similarities in this region, uh, of the
- 1:07:05particular gene that was analyzed, but
- 1:07:08the individual represented in the top
- 1:07:11electropherogram at this position has
- 1:07:14an adenine nucleotide. Hm. On the other
- 1:07:19hand, the individual represented in the
- 1:07:21bottom electropherogram, at this same
- 1:07:24position, has a guanine nucleotide.
- 1:07:26That is, this technique allows us, on
- 1:07:29one hand, to distinguish allelic
- 1:07:31variants that are present in different
- 1:07:33individuals, but I would like to point
- 1:07:36out the individual represented by the
- 1:07:38middle electropherogram, where two
- 1:07:40peaks appear simultaneously at this
- 1:07:43same position. One fluorophore
- 1:07:45reporting that there is adenine at that
- 1:07:47position and another fluorophore
- 1:07:49reporting that there is guanine at that
- 1:07:52same position. To correctly interpret
- 1:07:56this electropherogram, we have to
- 1:07:58remember that our genetic material is
- 1:08:01diploid. Therefore, it could happen
- 1:08:05that if the maternal and paternal
- 1:08:07allelic variants differ in their
- 1:08:10sequence, two different peaks will
- 1:08:12appear at the same position. In other
- 1:08:17words, this double peak being
- 1:08:19represented in the middle
- 1:08:21electropherogram is indicating an
- 1:08:24individual who has this allelic variant
- 1:08:27in heterozygosity. This Sanger
- 1:08:33sequencing method was widely used until
- 1:08:36the year 2000, but one of its
- 1:08:38particularities that can be thought of
- 1:08:40as a limitation is that for each Sanger
- 1:08:43experiment, only one type of molecule
- 1:08:45is sequenced. Right? For example, a
- 1:08:50particular exon of a gene, or an entire
- 1:08:52gene if it is small, can sometimes be
- 1:08:55sequenced. Hm. And while it has high
- 1:09:00sequencing fidelity, its costs are
- 1:09:03relatively high. Hm. This means it is
- 1:09:07only used if there is a clinical reason
- 1:09:11to suspect a relevant variant in a
- 1:09:14specific gene. Hm. Because only by
- 1:09:19being guided by clinical findings or
- 1:09:21this hypothesis will we go on to
- 1:09:23sequence that specific gene or exon. Hm
- 1:09:27. In other words, we would not apply
- 1:09:30the Sanger sequencing technique if we
- 1:09:32have a much more general hypothesis
- 1:09:34that there could be a pathogenic
- 1:09:36variant in the genome, but we do not
- 1:09:39know which gene is involved. For this
- 1:09:43other type of question, high-throughput
- 1:09:47sequencing methods might be more
- 1:09:49suitable than Sanger-based variants.
- 1:09:54What is the fundamental difference?
- 1:09:57Unlike Sanger, which sequenced a single
- 1:10:01type of molecule at a time, these
- 1:10:03methods sequence millions of molecules
- 1:10:07in parallel and could sequence, for
- 1:10:10example, an individual's entire genome
- 1:10:13at a relatively affordable cost and in
- 1:10:16much shorter times. Although we are not
- 1:10:23going to analyze the biochemistry of
- 1:10:26the high-throughput sequencing
- 1:10:29procedure, which is the goal of this
- 1:10:32course, it involves: extracting DNA
- 1:10:35from an individual, fragmenting these
- 1:10:38DNA molecules, sequencing the fragments
- 1:10:41in parallel, and then performing a
- 1:10:45bioinformatic analysis where the
- 1:10:48sequenced fragments are aligned with
- 1:10:51the complete reference genome available
- 1:10:55in public databases. This allows us to
- 1:10:58infer what differences were found in
- 1:11:01the individual from whom the DNA was
- 1:11:03extracted and subjected to this
- 1:11:05procedure, compared to the reference
- 1:11:08human genome. This methodology is
- 1:11:11increasingly being used in the context
- 1:11:14of medical genetics, but it has the
- 1:11:16difficulty that many variants are found
- 1:11:19with respect to the reference human
- 1:11:22genome. However, it is not always
- 1:11:24simple to determine if those variations
- 1:11:27found have a causal relationship with
- 1:11:30the pathology or if they are simply
- 1:11:33non-pathogenic variants that reflect
- 1:11:35the population's genetic diversity. Hm.
- 1:11:40I would finally like to mention
- 1:11:42regarding this technique that not only
- 1:11:44does it allow for whole-genome
- 1:11:46sequencing, but in some contexts,
- 1:11:48subsets of the genome can be sequenced.
- 1:11:51For example, we can focus only on the
- 1:11:54exonic regions of the genome, which are
- 1:11:57a much smaller percentage and allow us
- 1:12:00to focus on variations occurring in
- 1:12:02those coding regions of the genome. To
- 1:12:09finish this class, I would like to put
- 1:12:12into context what some situations might
- 1:12:16be where all the techniques we have
- 1:12:19analyzed today could be relevant. For
- 1:12:23example, let's consider we are
- 1:12:25performing a molecular diagnosis of
- 1:12:28monogenic entities. A first question is
- 1:12:31whether the clinical presentation, the
- 1:12:35signs and symptoms of the individual
- 1:12:38who consulted us, lead us to think
- 1:12:41there is a particular gene responsible
- 1:12:45for the signs and symptoms the
- 1:12:47individual presents. And based on prior
- 1:12:52history, because syndromes or clinical
- 1:12:55characteristics are already described
- 1:12:58that strongly suggest a pathology, we
- 1:13:01will be in certain conditions to make a
- 1:13:04molecular diagnosis that are different
- 1:13:07than if we do not have, for example, a
- 1:13:10strong hypothesis. If we don't have one
- 1:13:13, we probably have to use a
- 1:13:15hypothesis-free technique, right? An
- 1:13:18unbiased technique such as
- 1:13:20high-throughput sequencing. Hm. And
- 1:13:23there, we will eventually find, or not,
- 1:13:25variants in that individual that
- 1:13:28perhaps differ from the human reference
- 1:13:31genome. Hm. And we could develop a
- 1:13:34hypothesis regarding whether that
- 1:13:37variant is causal in producing the
- 1:13:40pathology. For example, we could
- 1:13:43corroborate that hypothesis if this
- 1:13:45same variant has been found in other
- 1:13:47individuals who have the same signs and
- 1:13:50symptoms, or if, in the context of
- 1:13:52basic research, some negative
- 1:13:54functional consequence of the
- 1:13:55appearance of that variant on protein
- 1:13:58function has been demonstrated; for
- 1:14:00example, if that variant is in a coding
- 1:14:02region. But let's imagine the other
- 1:14:07scenario where the signs and symptoms
- 1:14:10allow us to develop a hypothesis
- 1:14:12regarding which gene is involved in
- 1:14:15this disease. The next question is
- 1:14:20whether there are frequent mutations
- 1:14:22already described for that gene. And in
- 1:14:26the case where there are no frequently
- 1:14:28described mutations, one could perform
- 1:14:30Sanger sequencing of that gene and,
- 1:14:32again, see in that particular gene
- 1:14:34suggested by my clinical hypothesis if
- 1:14:36there are variants relative to the
- 1:14:38human reference genome or not, and if
- 1:14:40those variants have been reported as
- 1:14:43pathogenic or not. But in the case of
- 1:14:50having frequent mutations, or if it has
- 1:14:52already been found through other
- 1:14:54methodologies what mutations are
- 1:14:56present in a particular family, one
- 1:14:58could use more biased techniques. In
- 1:15:01other words, if I am looking for a
- 1:15:04particular mutation or allelic variant,
- 1:15:07I could use PCR-based techniques
- 1:15:09depending on the nature of that
- 1:15:11mutation. For example, we have seen
- 1:15:14that if the mutation we are looking for
- 1:15:17is an insertion or a deletion, we could
- 1:15:19perform conventional PCR. Given the
- 1:15:23difference in the size of the
- 1:15:24amplification products, the amplicons
- 1:15:26could report those allelic differences
- 1:15:28to us. We saw that in the case of
- 1:15:32substitution mutations, the appropriate
- 1:15:35variants of the PCR technique are, for
- 1:15:37example, allele-specific PCR or RFLP
- 1:15:40PCR. And in the particular case where
- 1:15:44our clinical context suggests a triplet
- 1:15:47expansion disease, we would use TPPCR,
- 1:15:50or Triple Prime PCR, since the other
- 1:15:53technique, uh, not based on PCR, the
- 1:15:56Southern blot technique, is one that,
- 1:15:59uh, while still used in some contexts,
- 1:16:02is increasingly falling into disuse due
- 1:16:06to its implementation difficulties. We
- 1:16:10also saw that uh or a variant of the
- 1:16:15PCR technique, specifically multiplex
- 1:16:17PCR, is used in the molecular context
- 1:16:20of forensic genetics. Here, there is no
- 1:16:24pathological entity involved; rather,
- 1:16:27the goal is to determine the profile of
- 1:16:30short tandem repeats, particularly
- 1:16:32microsatellites, in the context of
- 1:16:35parentage testing, for example, or
- 1:16:38other techniques associated with
- 1:16:40forensic genetics. With that, we finish
- 1:16:45today's class and we will see each
- 1:16:48other in a future session.
About this transcript
This page contains the full transcript of Seminario 11 Técnicas de Biología Molecular aplicadas a Entidades Monogénicas - Sebastián Giusti by Biología Molecular y Genética FMED - UBA, generated from the public captions YouTube serves with the video. The transcript has 9,023 words across 1,481 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.