YouTube2Text

Seminario 11 Técnicas de Biología Molecular aplicadas a Entidades Monogénicas - Sebastián Giusti — Transcript

by Biología Molecular y Genética FMED - UBA · 9,023 words · 1,481 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hello, how are you? In today's class,
  2. 0:04we are going to analyze various
  3. 0:06molecular biology techniques used in
  4. 0:09the context of medical genetics, which
  5. 0:12aim to detect different allelic
  6. 0:14variants present in our genome. In the
  7. 0:19first part of the class, we will focus
  8. 0:22on variants of the PCR technique to
  9. 0:25understand the foundation of this
  10. 0:27method and its variations in different
  11. 0:30scenarios that may arise in the context
  12. 0:33of medical genetics. And in the second
  13. 0:36part, we will analyze techniques based
  14. 0:39on DNA sequencing. In both cases, the
  15. 0:43goal will be to understand, on one hand
  16. 0:47, the molecular foundations of each of
  17. 0:50these techniques, and also the medical
  18. 0:54and clinical context where these
  19. 0:56techniques become necessary,
  20. 0:59understanding in each case what
  21. 1:02information they provide in each of
  22. 1:05these contexts. We will begin by
  23. 1:09reviewing what you have surely seen in
  24. 1:13previous classes, which are the
  25. 1:15fundamentals of the PCR technique,
  26. 1:18which is the acronym in English for
  27. 1:21polymerase chain reaction. This
  28. 1:25technique, which was developed in the
  29. 1:2880s by the biochemist Kary Mullis, who
  30. 1:31received the Nobel Prize in Chemistry
  31. 1:35in 1993 for it. 3 is a technique that
  32. 1:40allows for the amplification of a
  33. 1:43specific DNA segment starting from a
  34. 1:46complex mixture. Hm. It is, therefore,
  35. 1:50the reproduction of a chemical reaction
  36. 1:54where what occurs is the polymerization
  37. 1:58of nucleic acids. One aspect, an
  38. 2:04advantage of this technique, is that
  39. 2:07one can design the PCR technique to
  40. 2:09amplify a specific segment of the
  41. 2:12genome, even if the starting template
  42. 2:15used is total genomic DNA. Hm. And this
  43. 2:21technique can then, despite starting
  44. 2:24from an extensive template, total
  45. 2:27genomic DNA, for example, extracted
  46. 2:30from a peripheral blood sample or other
  47. 2:33biological material, specifically
  48. 2:36amplify a single DNA segment from that
  49. 2:39sample. In this polymerase chain
  50. 2:44reaction, the elements introduced to
  51. 2:47produce it are, as we already saw, the
  52. 2:50genomic DNA extracted from a patient
  53. 2:53that constitutes the template for the
  54. 2:56polymerization reaction. The DNA
  55. 3:01polymerase enzyme used is a DNA
  56. 3:03polymerase of bacterial origin, Taq
  57. 3:06polymerase, because it is
  58. 3:08heat-resistant. And for this enzyme,
  59. 3:13which is the enzyme that will
  60. 3:16polymerize the substrates in the
  61. 3:19amplification reactions, it needs dNTPs
  62. 3:22as substrates, that is,
  63. 3:24deoxynucleotides of adenine, thymine,
  64. 3:28cytosine, and guanine. And it will also
  65. 3:32need primers, which we will briefly
  66. 3:34look at and review to understand the
  67. 3:36role they play in this reaction. For
  68. 3:41DNA polymerase to function, this enzyme
  69. 3:44requires a buffer, a stabilizing
  70. 3:46solution that keeps the pH of the
  71. 3:49reaction stable, and divalent magnesium
  72. 3:52cations, which act as a cofactor
  73. 3:57preserving the active three-dimensional
  74. 4:00structure of the DNA polymerase. Let's
  75. 4:05focus once again on dNTPs and primers
  76. 4:09as molecular requirements for DNA
  77. 4:12polymerase activity. We just mentioned
  78. 4:18that these deoxyribonucleoside
  79. 4:20triphosphates of adenine, thymine,
  80. 4:23guanine, and cytosine are the
  81. 4:25substrates that DNA polymerase will
  82. 4:28polymerize, depending on the template
  83. 4:30used during replication. And these
  84. 4:35molecules are found in large excess in
  85. 4:38the reaction so that it proceeds
  86. 4:41quickly and effectively. Let's remember
  87. 4:46that for all DNA polymerases, including
  88. 4:49Taq polymerase, which is the polymerase
  89. 4:52of bacterial origin used here, they
  90. 4:55cannot synthesize de novo. Hm. On one
  91. 5:01hand, they need a single-stranded DNA
  92. 5:04template for polymerization, and on the
  93. 5:07other hand, since they do not
  94. 5:10synthesize de novo, this means they
  95. 5:13cannot start a new strand from scratch
  96. 5:16even with a template; rather, the
  97. 5:19enzymatic activity is restricted to
  98. 5:22extending a pre-existing strand
  99. 5:24previously paired to the template. In
  100. 5:29particular, the aspect—the chemical
  101. 5:32group necessary for this polymerase
  102. 5:35polymerization reaction—is that the
  103. 5:38last nucleotide incorporated into this
  104. 5:41pre-paired strand must have a free 3 '
  105. 5:44hydroxyl end, since this hydroxyl group
  106. 5:48at the 3' position is the necessary
  107. 5:51requirement for the next nucleotide to
  108. 5:55be incorporated in the polymerization
  109. 5:58reaction, and the new nucleotide will
  110. 6:01again have this free 3 'hydroxyl end
  111. 6:04that will allow the incorporation of
  112. 6:07the subsequent nucleotide. Precisely,
  113. 6:13primers fulfill this function: being a
  114. 6:17segment of DNA. Since the primers used
  115. 6:22during the polymerization that occurs
  116. 6:25in PCR are made of DNA to be more
  117. 6:29stable, they provide this previously
  118. 6:32paired strand that will then be the
  119. 6:36site for the polymerization reaction.
  120. 6:41But that is not the only function of
  121. 6:44primers; in addition to meeting that
  122. 6:46requirement that allows the
  123. 6:48functionality of the DNA polymerases in
  124. 6:51the reaction, primers fulfill another
  125. 6:54function: determining which region is
  126. 6:57going to be amplified. That is to say,
  127. 7:01they also have the function of
  128. 7:03conferring specificity as to which DNA
  129. 7:08segment will be amplified in that
  130. 7:10reaction. This is because in each PCR
  131. 7:15reaction, at least a pair of primers is
  132. 7:19used. Hm. One of those primers binds to
  133. 7:23one of the DNA strands, and the other
  134. 7:27primer binds to the complementary
  135. 7:30strand. Hm. When these primers are
  136. 7:35annealed during the polymerization
  137. 7:37reaction, they set the limit for the
  138. 7:40new strands, and therefore, when they
  139. 7:43are used as templates themselves in the
  140. 7:46next amplification cycle, they will
  141. 7:49provide a fixed endpoint for that
  142. 7:51amplification reaction. The PCR
  143. 7:58reaction consists of a series of
  144. 8:00amplification cycles where each of
  145. 8:03these cycles has three steps. Hm. The
  146. 8:06first step involves DNA denaturation,
  147. 8:09bringing the reaction to 95 ° for a
  148. 8:12short time. This allows genomic DNA,
  149. 8:17for example, and later all
  150. 8:21amplification products as the reactions
  151. 8:25progress, to denature. That is, the
  152. 8:28double-stranded DNA separates into
  153. 8:31single strands, since the template must
  154. 8:33be single-stranded to be a substrate
  155. 8:35for amplification by DNA polymerase. In
  156. 8:41the second step, the temperature drops
  157. 8:44to between 50 and 70 °, depending on
  158. 8:47the composition of the primers, because
  159. 8:50this lower temperature allows the
  160. 8:53primers to hybridize, to bind to the
  161. 8:56DNA strands. And then in the third step
  162. 9:01, which typically occurs at 72 degrees,
  163. 9:04which is the optimal polymerization
  164. 9:07temperature for Taq polymerase,
  165. 9:09elongation occurs, the effective
  166. 9:11polymerization of those fragments. Hm.
  167. 9:16When these cycles are repeated many
  168. 9:19times, what ends up being generated is
  169. 9:22an exponential amplification. And the
  170. 9:26exponentiality is due to the fact that
  171. 9:29in each of these cycles, the amount of
  172. 9:32template of that region to be
  173. 9:34specifically amplified doubles. So, in
  174. 9:37each cycle, you have double the amount
  175. 9:40of templates you had in the previous
  176. 9:42one. Hm. Since dNTPs, Taq polymerase
  177. 9:47molecules, and primers are placed in
  178. 9:50excess inside that tube, the reaction
  179. 9:54occurs, amplifying exponentially, and
  180. 9:57after 30 cycles, for example, there
  181. 10:01will be 1 billion specifically
  182. 10:03amplified DNA molecules, making the
  183. 10:07genomic DNA we initially added a
  184. 10:10minuscule contaminant. compared to the
  185. 10:14enormous amplification product
  186. 10:16throughout these cycles. Therefore, it
  187. 10:19is considered that by the end of the
  188. 10:23PCR reaction, we have a segment of
  189. 10:27amplified DNA with billions of copies
  190. 10:31that is virtually pure. How are these
  191. 10:36amplification products visualized?
  192. 10:41Usually, a technique called nucleic
  193. 10:43acid gel electrophoresis is used.
  194. 10:46Electrophoresis involves subjecting
  195. 10:48charged molecules to an electric field
  196. 10:51so that they migrate based on their
  197. 10:53charges. As you know, nucleic acid
  198. 10:56molecules are negatively charged
  199. 10:59because the phosphate groups that
  200. 11:02contain their ribonucleotides have a
  201. 11:06net negative electrical charge and will
  202. 11:09migrate toward the positive electrode
  203. 11:13in this reaction of molecule separation
  204. 11:18. This electrophoresis does not occur
  205. 11:21in a liquid medium, but in a gel. And
  206. 11:24this gel, if we were to see it
  207. 11:27microscopically, is built by a matrix
  208. 11:30of—it's a polymer that forms a matrix
  209. 11:33of pores, a porous matrix. And so, when
  210. 11:39one subjects nucleic acid molecules to
  211. 11:42migrate in a solution through this
  212. 11:46porous matrix, what happens is that the
  213. 11:49smaller fragments can pass through this
  214. 11:54network of pores more easily than the
  215. 11:57larger fragments, which will collide a
  216. 11:59greater number of times against the
  217. 12:01walls of this gel and will be slowed
  218. 12:03down in their speed. The final effect,
  219. 12:07then, is that electrophoresis allows
  220. 12:10for the separation of DNA fragments
  221. 12:12according to their size. To visualize
  222. 12:17the DNA fragments within the gel, a
  223. 12:21molecule like ethidium bromide is used,
  224. 12:25which is a base intercalator. It is a
  225. 12:29molecule of such a size that it can fit
  226. 12:32in between two successive DNA base
  227. 12:34pairs. Hm. And one of the properties of
  228. 12:38this ethidium bromide molecule is that
  229. 12:40when it is illuminated by ultraviolet
  230. 12:42light, it emits fluorescence. So, if we
  231. 12:46introduced molecules of a dye like
  232. 12:48ethidium bromide into the gel matrix
  233. 12:51and then developed the gel with
  234. 12:53ultraviolet light, we will see that
  235. 12:56they have separated; that is, we will
  236. 12:58see fluorescent marks where there is
  237. 13:01DNA. In this photo of this gel, we
  238. 13:05observe that different samples have
  239. 13:07been loaded into each of the lanes of
  240. 13:10this gel. For example, in the first
  241. 13:13lane and the last lane, molecular
  242. 13:18weight markers were loaded, that is, a
  243. 13:21pre-made mixture of different sizes of
  244. 13:23DNA molecules that then serve to
  245. 13:26estimate the size of the molecule that
  246. 13:28was the product of the PCR reaction.
  247. 13:31Let's imagine that we ran four
  248. 13:35reactions here: 1, 2, 3, and 4. We can
  249. 13:40see that in the first three there was
  250. 13:42an amplification product and we could
  251. 13:44estimate the size of that amplification
  252. 13:47product by comparing it with the
  253. 13:49molecular weight markers, while in the
  254. 13:52last reaction there was no
  255. 13:53amplification product at all. Sometimes
  256. 13:58, these images obtained from developing
  257. 14:01the gels are often digitally inverted,
  258. 14:04showing us dark bands on a light
  259. 14:07background, as this is usually more
  260. 14:10comfortable for our eyes for
  261. 14:12visualization; but both types of
  262. 14:15visualizations are common in research
  263. 14:18papers and textbooks. I would like to
  264. 14:24tell you that there are alternatives in
  265. 14:28electrophoresis to visualize
  266. 14:30amplification products, and
  267. 14:32electrophoresis can sometimes be
  268. 14:35performed in a capillary. Hm. That is,
  269. 14:39the system is equivalent; there is an
  270. 14:42electrophoresis in the sense that an
  271. 14:45electric field is applied and there is
  272. 14:48a negative pole and a positive pole,
  273. 14:51but now the gel is it is contained
  274. 14:56within a small, thin capillary tube. Hm
  275. 14:59. But it is also a porous matrix. And
  276. 15:04when DNA molecules are separated by
  277. 15:07size, using the same principle we just
  278. 15:10saw, visualization occurs because at a
  279. 15:13specific point in this capillary there
  280. 15:16is a detector that will provide a
  281. 15:19series of fluorescence peaks as the
  282. 15:22fluorescently labeled DNA molecules
  283. 15:24pass through, either by a base
  284. 15:27intercalator or another detection
  285. 15:30method. And the result of this
  286. 15:34capillary electrophoresis is usually
  287. 15:37observed digitally in the form of an
  288. 15:39electropherogram. Hm. But, analogously,
  289. 15:44these peaks represent what would be two
  290. 15:47bands in a gel. Hm. Where the height of
  291. 15:52the peaks implies the fluorescence
  292. 15:55intensity of the bands, and what is
  293. 15:58represented on the X-axis is the
  294. 16:00separation according to size. In other
  295. 16:04words, both types of information can be
  296. 16:07obtained by gel electrophoresis or by
  297. 16:09capillary electrophoresis. In the first
  298. 16:13case, it is the visualization of the
  299. 16:15gel and the interpretation of the
  300. 16:17banding pattern that is informative,
  301. 16:19and in the second case, it is the
  302. 16:20interpretation of the corresponding
  303. 16:22electropherogram. Here we see another
  304. 16:29example where we see at the top an
  305. 16:31electropherogram corresponding to a PCR
  306. 16:35reaction where there was a single
  307. 16:38amplification product, a single
  308. 16:40amplicon. That is the technical name we
  309. 16:43give it, and that allows us to explain
  310. 16:45this single peak. On the other hand, in
  311. 16:48a second example, let's imagine that in
  312. 16:50another PCR reaction two different
  313. 16:52amplification products were generated,
  314. 16:54and this is observed as a series of two
  315. 16:56peaks. Hm. Next, having seen the
  316. 17:04foundations of the PCR technique, the
  317. 17:07reagents a basic reaction of this type
  318. 17:10entails, and the modes of visualizing
  319. 17:13PCR products, we will try to understand
  320. 17:17the uses of this technique in the
  321. 17:20context of medical genetics. That is to
  322. 17:25say, thinking that we are going to use
  323. 17:27this technique to determine the
  324. 17:29presence or absence of specific allelic
  325. 17:32variants. Something important to
  326. 17:36consider is that PCR is a
  327. 17:38hypothesis-based technique, meaning
  328. 17:41that if we are using it in a medical or
  329. 17:45clinical context, it is because we are
  330. 17:48looking for a particular allelic
  331. 17:50variant. That is, the clinical context
  332. 17:54should suggest that we are thinking of
  333. 17:57looking for allelic variants of a
  334. 17:59specific gene, and even within the same
  335. 18:02gene, that we are looking for a certain
  336. 18:05mutation in particular. Hm. Since the
  337. 18:10design of the primers in this technique
  338. 18:13requires having hypothesized about the
  339. 18:17nature of these causes of our clinical
  340. 18:20picture. Let's see an example in the
  341. 18:25context of a monogenic entity such as
  342. 18:28cystic fibrosis. We won't get into the
  343. 18:32details of this pathology, uh, but let
  344. 18:35me tell you in advance that we already
  345. 18:38know there is a symptomatology in
  346. 18:41certain boys and girls that allows us
  347. 18:43to suspect this condition, which is
  348. 18:46characterized by mucus plugging of
  349. 18:48multiple ducts. And we also know that
  350. 18:53it originates from pathogenic variants
  351. 18:56in a chloride channel gene called CFTR,
  352. 19:02and that the most frequent mutation in
  353. 19:05the population, uh, that causes this
  354. 19:09pathology, is a mutation called
  355. 19:12phenylalanine 508, where a deletion of
  356. 19:16three nucleotides in the gene, in an
  357. 19:19exon, in a strictly coding region,
  358. 19:22correlates to the deletion of a codon
  359. 19:26that codes for the amino acid.
  360. 19:30phenylalanine, and that this absence of
  361. 19:32this amino acid has negative
  362. 19:34consequences for protein function. So,
  363. 19:39if there is a patient who has clinical
  364. 19:45signs compatible with cystic fibrosis,
  365. 19:48it could be useful, since it is known
  366. 19:51that in many patients, uh, this
  367. 19:54particular allelic variant that
  368. 19:56involves the deletion of the
  369. 19:59phenylalanine is the one present, for
  370. 20:02one to go and look if this
  371. 20:05phenylalanine deletion is present or
  372. 20:08absent at this specific position of the
  373. 20:11CFTR gene. So, in this context, I would
  374. 20:17like to show you how conventional PCR
  375. 20:20can be very useful for detecting
  376. 20:23insertions or deletions as in this case
  377. 20:26and guided by a clinical hypothesis.
  378. 20:32For that, since we know which gene we
  379. 20:37want to explore to see if it contains
  380. 20:40this insertion or deletion, we can
  381. 20:43design a pair of primers, because we
  382. 20:46know the human genome sequence and can
  383. 20:49then access it via databases and use it
  384. 20:53to, uh, design our primers; and we will
  385. 20:56design them in a region flanking the
  386. 20:59position where the deletion may or may
  387. 21:02not be. Here in this example, we are,
  388. 21:06uh, sketching at the top an allelic
  389. 21:10variant that does not contain the
  390. 21:13deletion in question and another that
  391. 21:16does. Note that since the primers are
  392. 21:20in a flanking region, in both cases, in
  393. 21:22both allelic variants, the primers can
  394. 21:25bind and will be able to generate
  395. 21:27amplicons, that is, amplification
  396. 21:29reactions in both cases. Hm. But the
  397. 21:34amplification products will be slightly
  398. 21:36different, since if we compare them
  399. 21:38with each other, the amplification
  400. 21:40product of the variant that does not
  401. 21:42contain the deletion will have three
  402. 21:45nucleotides more than the variant that
  403. 21:47does possess the deletion. Although
  404. 21:51this difference of three nucleotides is
  405. 21:54small, polyacrylamide gels, which are
  406. 21:59different from agarose gels, can be
  407. 22:01used to discriminate these differences.
  408. 22:04Hm. That is to say, in a polyacrylamide
  409. 22:08gel matrix, electrophoresis, which will
  410. 22:10separate the molecules according to
  411. 22:12their size, can occur in the same way.
  412. 22:16And let's analyze then Uh, the result
  413. 22:21of this PCR, imagining that we obtained
  414. 22:24biological samples, perhaps through
  415. 22:26blood, peripheral blood samples, from
  416. 22:29members of a family where some of them
  417. 22:32showed clinical signs of cystic
  418. 22:35fibrosis and we wanted to explore the
  419. 22:38presence or absence of this particular
  420. 22:41deletion in the CFTR gene. Let's focus,
  421. 22:46for example, on individual number four.
  422. 22:52Hm. Individual number four has a single
  423. 22:57amplification product. Hm. And if we
  424. 23:01compare it with individual number two,
  425. 23:04which also has a single, uh, amplicon
  426. 23:06fragment, an amplification product, if
  427. 23:08we compare them with each other, we can
  428. 23:11say that the amplification product of
  429. 23:13individual four is larger in base pairs
  430. 23:16than the amplification product of two;
  431. 23:18because, as we had said, the smaller
  432. 23:20fragments migrate further in the gel,
  433. 23:23they can pass through the porous matrix
  434. 23:26with less difficulty than larger
  435. 23:28fragments. So, considering, we would
  436. 23:33now have to make considerations
  437. 23:35regarding the structure of the human
  438. 23:37genome to know what has happened there.
  439. 23:42We have to consider, for example, the
  440. 23:44diploid nature of our genome, that we
  441. 23:46have two allelic variants, one of
  442. 23:48maternal origin and another of paternal
  443. 23:50origin for each of the genes. If we
  444. 23:53have, in both individual two and
  445. 23:55individual four, a single amplification
  446. 23:57product, we will say that the maternal
  447. 24:01and paternal variants are behaving in
  448. 24:03the same way, they are generating the
  449. 24:05same type of amplification product. Hm.
  450. 24:07Now then, how to understand a larger
  451. 24:11amplification product versus a smaller
  452. 24:15one? In the context of this PCR, where
  453. 24:18we use primers to amplify a certain
  454. 24:20segment of the CFTR gene, one could say
  455. 24:23that individual 4, since it has the
  456. 24:28larger sizes, its two allelic variants,
  457. 24:30the maternal and the paternal, were
  458. 24:32allelic variants without the
  459. 24:34phenylalanine deletion. 508, while
  460. 24:37individual 2 has both of its allelic
  461. 24:41variants with the deletion. That is why
  462. 24:44we see a single amplification product
  463. 24:46resulting from those two allelic
  464. 24:50variants that were used as a template
  465. 24:52during the PCR, but which generated the
  466. 24:54same amplification product. In contrast
  467. 24:58to these cases, individuals 1 and 3
  468. 25:01have two amplification products, one
  469. 25:04larger and one smaller. Hm. And so we
  470. 25:08can think that in those individuals,
  471. 25:10one of the allelic variants has the
  472. 25:13deletion and the other does not. Hm. In
  473. 25:15other words, the deletion would be in
  474. 25:18heterozygosis in these two individuals.
  475. 25:25With what we have seen so far, we could
  476. 25:28think that since insertions and
  477. 25:30deletions can generate amplicons, that
  478. 25:33is, amplification products of different
  479. 25:36sizes that will later be separated by
  480. 25:39gel electrophoresis, we can understand
  481. 25:42why PCR can be used to detect
  482. 25:45insertions and deletions. Now, what
  483. 25:48would happen if we want to detect other
  484. 25:51types of allelic variants? For example,
  485. 25:56those caused by a substitution mutation
  486. 25:59where there will be no size differences
  487. 26:03between the amplification products. To
  488. 26:09detect allelic variants caused by
  489. 26:12substitution mutations, you must use
  490. 26:16variations of the PCR technique
  491. 26:19specifically designed to detect these
  492. 26:22substitutions. And one of these
  493. 26:27variations is allele-specific PCR,
  494. 26:30which is specifically designed for this
  495. 26:34purpose. What is the design feature
  496. 26:39that allows this? Imagine we have two
  497. 26:42allelic variants that differ by a point
  498. 26:46mutation, meaning that in the same
  499. 26:49region where one allelic variant has an
  500. 26:52AT pair, another allelic variant at
  501. 26:56that same position of the same gene has
  502. 26:59a CG base pair. The key here is to
  503. 27:05design a pair of primers where one of
  504. 27:09them binds specifically at its 3'
  505. 27:11hydroxyl end, which must necessarily be
  506. 27:15base-paired to the template for
  507. 27:17amplification to occur. We have to
  508. 27:22design that primer to be
  509. 27:24allele-specific as well. In other words
  510. 27:28, if we design it so that it can only
  511. 27:30bind to one of the allelic variants,
  512. 27:32there will be a mismatch at this point
  513. 27:34in the other allelic variant. And this
  514. 27:38implies that even if the other primer
  515. 27:41is bound, one of the allelic variants
  516. 27:43will not be able to amplify during the
  517. 27:46PCR reaction, while the other will. Hmm
  518. 27:49. So, depending on whether there is an
  519. 27:53amplification product or not, one could
  520. 27:56infer which of the two allelic variants
  521. 27:58was present. For example, if we take
  522. 28:05samples from five patients to determine
  523. 28:08which of the two allelic variants they
  524. 28:11possessed and we have these results,
  525. 28:13here we see a lane as a result of the
  526. 28:16gel electrophoresis revealing a
  527. 28:18molecular weight marker; but the PCR of
  528. 28:21patient 1, and of patient 2 and 5, give
  529. 28:24us an amplification product. We can say
  530. 28:29that they had at least one CG-type
  531. 28:31allelic variant in their genome. Now,
  532. 28:36what we can also say is that
  533. 28:38individuals 3 and 4 apparently did not
  534. 28:40have any CG variant in their genome.
  535. 28:44Usually, these allele-specific PCRs
  536. 28:47need to perform another PCR to detect
  537. 28:50the other variant to have complete
  538. 28:53information. Hmm. For example, if we
  539. 28:58use primers with the same patients '
  540. 29:00genomic DNA to detect the other variant
  541. 29:03, that is, primers that can bind when
  542. 29:06the AT dinucleotide is present, but not
  543. 29:09the CG dinucleotide, we complete the
  544. 29:12information here. For instance, we see
  545. 29:16that individuals 3 and 4, which
  546. 29:17previously showed no amplification
  547. 29:19signal, now show one. Then we can say
  548. 29:23that the two allelic variants, both
  549. 29:25maternal and paternal in origin, of
  550. 29:28individuals 3 and 4 had an AT allelic
  551. 29:30variant and did not have any CG. But
  552. 29:35what can we say, for example, about
  553. 29:38individual 1, who has an amplification
  554. 29:41product in both one PCR and the other?
  555. 29:45In this case, this individual is
  556. 29:47probably heterozygous, meaning they
  557. 29:49have one allelic variant of parental
  558. 29:52origin of the CG type and another AT
  559. 29:54variant. Therefore, there was an
  560. 29:56amplification product in both cases. Hm
  561. 29:59. I leave it to you to think about what
  562. 30:03the genetic configuration will be in
  563. 30:04terms of these variants for individual
  564. 30:075, who presents amplification here and
  565. 30:11not here. Let me tell you that there is
  566. 30:18also another variant of the PCR
  567. 30:21technique called PCR-RFLP which is used
  568. 30:28to detect allelic variants with point
  569. 30:32substitution mutations, that is, that
  570. 30:35differ from one another by a
  571. 30:38substitution mutation. This acronym
  572. 30:44RFLP stands for restriction fragment
  573. 30:50length polymorphisms. And this refers
  574. 30:55to the use in the research context of
  575. 30:58endonucleases, called restriction
  576. 31:01enzymes, which are bacterial
  577. 31:03endonucleases that can cleave DNA at
  578. 31:06specific sequences. In this case, if
  579. 31:11one wants to differentiate two allelic
  580. 31:13variants, for example, let's see that
  581. 31:15this variant differs at this point. Hm.
  582. 31:18Here there is AT and here there is TA.
  583. 31:22That is to say, there is a substitution
  584. 31:24at a particular point. Whoever designs
  585. 31:29this technique has to infer if in the
  586. 31:32context where the variation one wants
  587. 31:34to detect is located, it contains any
  588. 31:37recognition sequence for a restriction
  589. 31:40enzyme. If one goes to the databases
  590. 31:44and compares the recognition sequences
  591. 31:47of restriction enzymes and finds that
  592. 31:50this variation affects one of these
  593. 31:52sites, one could say the following. For
  594. 31:56example, in this case there is a
  595. 31:57restriction enzyme which is D1 that can
  596. 32:02cut, that is, it has a recognition site
  597. 32:06in the allelic variant that had AT at
  598. 32:09this position, but when that nucleotide
  599. 32:12is substituted by its inverse TA, the
  600. 32:15restriction enzyme can no longer cut.
  601. 32:21Realizing this aspect allows us to
  602. 32:24design a PCR in combination with a
  603. 32:26cleavage reaction using this
  604. 32:28restriction enzyme, which will
  605. 32:30eventually allow us—and this is our
  606. 32:32goal—to distinguish between these
  607. 32:34allelic variants. In particular, to
  608. 32:38determine which allelic variants an
  609. 32:40individual has, um, that we are
  610. 32:44analyzing. For that, we simply design
  611. 32:49primers. that flank the restriction
  612. 32:52site where the variant we are looking
  613. 32:54for is located. And since these primers
  614. 32:57do not affect or bind to the
  615. 32:59potentially polymorphic region, they
  616. 33:02bind to both one allelic variant and
  617. 33:05the other. When amplifying by PCR, in
  618. 33:09both cases, we will obtain
  619. 33:11amplification products of the same size
  620. 33:15. And as a second characteristic step
  621. 33:18of this technique, we treat the PCR
  622. 33:20amplification products with the
  623. 33:22restriction enzyme. So, let's imagine
  624. 33:27what happens in the first step. The two
  625. 33:32allelic variants are amplified by PCR.
  626. 33:35We have millions and millions of copies
  627. 33:37of these amplification products in each
  628. 33:40PCR tube, which are equal in size. Now,
  629. 33:43when we treat them with the restriction
  630. 33:47enzyme, only one of the fragments can
  631. 33:50be cut and the other will remain uncut
  632. 33:53because the allelic variant precisely
  633. 33:56affected this enzyme's restriction site
  634. 33:59. And so now we have a polymorphism in
  635. 34:02the length of the fragments. Hm. Hence
  636. 34:06the name of the technique, and we know
  637. 34:09that DNA fragments of different sizes
  638. 34:11can be separated using the
  639. 34:13electrophoresis technique. Let's
  640. 34:18interpret once again five patients
  641. 34:20whose genetic makeup we have explored
  642. 34:23thanks to this technique. Let us recall
  643. 34:28, then, that the allelic variant we
  644. 34:30call AT here was the one that retained
  645. 34:33the restriction site and could be
  646. 34:35cleaved, generating these two fragments
  647. 34:38, while the TA variant was not going to
  648. 34:40be a substrate for this restriction
  649. 34:43reaction and would generate a longer
  650. 34:45fragment. For example, if we think of
  651. 34:50patient number four, who has a single
  652. 34:52large-sized amplification product, Hm.
  653. 34:56because it is the one that migrates the
  654. 34:57least. We can say that individual four
  655. 35:00had a maternal and paternal allelic
  656. 35:02variant with the TA-type variant, since
  657. 35:05the largest fragments are the ones that
  658. 35:08are not substrates for the restriction
  659. 35:10enzyme. Hm. On the other hand, if we
  660. 35:15think of individual one, where we have
  661. 35:18two restriction fragments and both are
  662. 35:21smaller in size than the one we had
  663. 35:23found in individual four, we can think
  664. 35:26that this individual only had AT-type
  665. 35:29allelic variants that generated these
  666. 35:32two fragments that are different from
  667. 35:35each other, but in both cases smaller
  668. 35:38than the other allelic variant. Hm. And
  669. 35:43the most complex case to interpret is
  670. 35:45found in individuals 3 and 5, where we
  671. 35:48see three amplification products on the
  672. 35:51gel. And in this case, we interpret
  673. 35:56that these individuals had a single TA
  674. 35:58allelic variant, which was not cut by
  675. 36:01the restriction enzyme and generated
  676. 36:03the larger fragment. And the other
  677. 36:07allelic variant was of the AT type,
  678. 36:09which was fragmented and generated the
  679. 36:12smaller fragments. I would now like to
  680. 36:20tell you about another application of
  681. 36:22PCR, which is multiplex PCR, in a very
  682. 36:27particular context: the context of
  683. 36:29forensic genetics, where, in this case,
  684. 36:32it is not a pathological context, but
  685. 36:35rather an attempt to determine genetic
  686. 36:37variants that allow for the
  687. 36:39establishment of biological
  688. 36:41relationships between individuals, for
  689. 36:44example, filiation relationships. Hm.
  690. 36:48And in this context, usually what is
  691. 36:52explored are highly polymorphic regions
  692. 36:56of our genome, which are short tandem
  693. 36:59repeats, or STRs, usually represented
  694. 37:03by elements called microsatellites that
  695. 37:07we analyzed in seminar number nine.
  696. 37:13These, uh, these microsatellites, these
  697. 37:17short tandem repeats, behave like
  698. 37:19allelic variants, that is to say, all
  699. 37:22human beings at the same position of
  700. 37:25certain chromosomes possess a variant
  701. 37:29of maternal origin and a variant of
  702. 37:31paternal origin of these repetitive
  703. 37:34regions. And since they are highly
  704. 37:38polymorphic, they differ from each
  705. 37:40other very frequently in the exact
  706. 37:42number of repeats that each one
  707. 37:44possesses. Hm. And since each one of us
  708. 37:50received these variants, one of
  709. 37:52maternal origin and one of paternal
  710. 37:54origin, if one explores, for example,
  711. 37:57for one, uh, for one of these
  712. 37:59microsatellites in particular, we
  713. 38:01should probably find one of those
  714. 38:04allelic variants in our biological
  715. 38:06mother and the other allelic variant we
  716. 38:09should find in our biological father.
  717. 38:15To increase the fidelity of these
  718. 38:18analyses, for example, of filiation, a
  719. 38:21single microsatellite is not analyzed,
  720. 38:25but rather, uh, in the year 1997, the
  721. 38:30FBI standardized a procedure that
  722. 38:35analyzes 13 sequences, uh, of mini
  723. 38:40microsatellites scattered throughout
  724. 38:43different human chromosomes. Therefore,
  725. 38:47the inferences will not only be based
  726. 38:50on a single microsatellite, but on
  727. 38:52multiple, uh, loci, multiple
  728. 38:54chromosomal positions that are on
  729. 38:56different chromosomes. These names that
  730. 39:01we see here, TPOX, D8S 1179, VWA, etc.,
  731. 39:05are the specific names of tandem repeat
  732. 39:12regions, which are highly polymorphic
  733. 39:15in the human population, but which all
  734. 39:18humans have at those positions on those
  735. 39:21chromosomes. So, I wanted to tell you
  736. 39:28that this multiplex PCR technique is
  737. 39:31applied to analyze the length of these
  738. 39:35regions, to determine the number of
  739. 39:39repetitions that each of these
  740. 39:42microsatellites has in a single
  741. 39:45reaction at the same time. What is the
  742. 39:51key? What specific aspect of multiplex
  743. 39:56PCR, which, unlike classical PCRs where
  744. 39:59we used a pair of primers to amplify
  745. 40:02those regions of the genome that we
  746. 40:05wanted to amplify. In this case, we are
  747. 40:10going to use many different pairs of
  748. 40:13primers, each one of them designed to
  749. 40:15amplify one of these microsatellites in
  750. 40:18particular, that is, the flanking
  751. 40:21region that allows amplifying that
  752. 40:23microsatellite which is on each of the
  753. 40:26chromosomes we analyze. Precisely
  754. 40:31because this reaction is complex,
  755. 40:34because amplicons from different genes,
  756. 40:37from different pairs of primers, are
  757. 40:40going to be mixed in the same tube,
  758. 40:43fluorophores are usually used to
  759. 40:45visualize the amplification products,
  760. 40:49molecules linked to the primers, that
  761. 40:52will emit fluorescence in different
  762. 40:55colors, thus allowing the different PCR
  763. 40:58products to be differentiated by color.
  764. 41:03Since this reaction is very complex,
  765. 41:05because we have at least 13 different
  766. 41:08amplification products, what is done is
  767. 41:11to use the same fluorophores for some
  768. 41:14genes that have amplicons of different
  769. 41:17sizes. So that by combining these size
  770. 41:21differences with differences in
  771. 41:23fluorophores, we can obtain the total
  772. 41:25information. Usually, this multiplex
  773. 41:30PCR is combined with capillary
  774. 41:32electrophoresis to then detect the
  775. 41:35different products as we saw. And in
  776. 41:40the detector, the fluorescences coming
  777. 41:42from the different colors that we used
  778. 41:45in the fluorophores of the different
  779. 41:47pairs of primers can be detected. I
  780. 41:51show you an example here in this
  781. 41:54electropherogram. As a result of having
  782. 41:59used multiple pairs of primers with
  783. 42:01different fluorophores, we obtain a
  784. 42:04series of peaks with different colors.
  785. 42:09Let's remember that on the vertical
  786. 42:11axis, the height of the peaks implies
  787. 42:13the fluorescence intensity of each of
  788. 42:15those amplification products, and on
  789. 42:17the horizontal axis, these peaks are
  790. 42:20separated based on the size of the
  791. 42:22amplicon. Hm. Digitally, one can
  792. 42:26separate these signals of different
  793. 42:27colors to perform a more detailed
  794. 42:29analysis. For example, here we see how
  795. 42:33we separate the channel of those
  796. 42:35amplification products that had a blue
  797. 42:38fluorophore in the upper part and in
  798. 42:40the lower part the channel of those
  799. 42:43green fluorophores. Note that because
  800. 42:49of how the primers were designed, one
  801. 42:50can know in which region we will obtain
  802. 42:55the amplification products of a
  803. 42:56particular microsatellite. For example,
  804. 42:59for this particular microsatellite we
  805. 43:01have two amplification products. We see
  806. 43:03two peaks. In this case, for this
  807. 43:06particular microsatellite, this
  808. 43:08individual we analyzed had a
  809. 43:10microsatellite with 12 repetitions and
  810. 43:12another with 13 of this one in
  811. 43:15particular. Conversely, for this other
  812. 43:18microsatellite from the same individual
  813. 43:21, we have a single peak. What does this
  814. 43:23mean? That the maternal and paternal
  815. 43:26allelic variants had the same number of
  816. 43:28repeats. Hm. So, by combining the
  817. 43:31information from these multiple primer
  818. 43:34pairs associated with different
  819. 43:36fluorophores, one can, uh, know the
  820. 43:41exact number of repeats an individual
  821. 43:44has for each of these tandem repeat
  822. 43:47regions and compare it, then, with
  823. 43:51potential, uh, likely subjects, uh,
  824. 43:54biological fathers or mothers, and
  825. 43:58determine the probability of whether or
  826. 44:01not a biological affiliation exists
  827. 44:04between them. Finally, I would like to
  828. 44:12tell you about this last variant of the
  829. 44:15PCR technique we are going to analyze
  830. 44:19today, which is TP-PCR, or Triplet
  831. 44:22Primed PCR, used in the genetic
  832. 44:24diagnosis of triplet repeat expansion
  833. 44:28diseases. This is a set of
  834. 44:31non-classical monogenic diseases that
  835. 44:34will be analyzed in depth. in an
  836. 44:37upcoming seminar, but they are
  837. 44:40characterized by the expansion of a
  838. 44:42three-nucleotide repeat within the
  839. 44:45affected gene. So, in the population,
  840. 44:49there are allelic variants that have a
  841. 44:52normal range of repeats, meaning they
  842. 44:55are non-pathogenic. And these, for
  843. 44:59example, in this particular case,
  844. 45:01unaffected individuals have between 5
  845. 45:03and 50 repeats of this trinucleotide in
  846. 45:05their maternal and/or paternal allelic
  847. 45:08variant. A characteristic of this
  848. 45:13disease is that some individuals have
  849. 45:15an allelic variant called a premutation
  850. 45:18with a higher number of repeats. And
  851. 45:21the premutation is characterized by the
  852. 45:24fact that those who carry this
  853. 45:27premutation do not show signs and
  854. 45:29symptoms of the pathological entity,
  855. 45:32but there is a high probability that in
  856. 45:35their gametes, an expansion of the
  857. 45:37mutation occurs, h, of a higher range
  858. 45:40that will then be associated, in the
  859. 45:43next generation, with the appearance of
  860. 45:46signs and symptoms. As you will see,
  861. 45:50this technique, what it tries to
  862. 45:52determine in the context where there is
  863. 45:55a hypothesis that there is a triplet
  864. 45:57expansion disease behind the signs and
  865. 46:00symptoms of some family members, is to
  866. 46:02determine if an individual is a carrier
  867. 46:05of allelic variants with expansions
  868. 46:07within the normal range, within the
  869. 46:09premutation range, or the pathogenic
  870. 46:12mutation range. And you might ask
  871. 46:16yourselves, why couldn't conventional
  872. 46:18PCR discriminate this? Since we have
  873. 46:21seen that insertions, for example, that
  874. 46:23variations in length could be detected
  875. 46:25by the PCR technique. But although from
  876. 46:31a theoretical point of view that is
  877. 46:33true, there are specific problems that
  878. 46:36arise for conventional PCR when it
  879. 46:39comes to detecting, uh, triplet
  880. 46:41expansions. In this case, we see the
  881. 46:46FMR1 gene as an example, which is one
  882. 46:50that can have this huge triplet
  883. 46:53expansion in its 5-prime UTR region. We
  884. 47:01mentioned that the limitation of
  885. 47:04conventional PCR when detecting
  886. 47:07amplification products of these
  887. 47:09expanded alleles is due to the fact
  888. 47:12that these regions containing multiple
  889. 47:15repeats are often a difficult context
  890. 47:18for the DNA polymerase used in these
  891. 47:21reactions to effectively amplify the
  892. 47:24entire fragment. And this causes a
  893. 47:29series of errors, a series of
  894. 47:31incomplete fragments of the potentially
  895. 47:33expanded fragment to be amplified,
  896. 47:36which results in the template not being
  897. 47:39able to be effectively duplicated from
  898. 47:41cycle to cycle, and towards the end of
  899. 47:44the PCR, we observe no amplification
  900. 47:47product if there was an expanded allele
  901. 47:50in the template. Simply put, the PCR
  902. 47:54fails and fails to amplify. When we
  903. 47:59only use a pair of primers flanking the
  904. 48:02potentially expanded region, it fails
  905. 48:05to amplify that fragment. Hm. So, if a
  906. 48:10diploid individual has an allelic
  907. 48:12variant from a parental origin without
  908. 48:15amplification, this will be amplified.
  909. 48:18But if it has an expanded allele as the
  910. 48:21other allelic variant, there will be no
  911. 48:24amplification product, and since only
  912. 48:27the amplification product of the
  913. 48:29non-expanded variant will be seen on
  914. 48:32the gel, there is a masked result.
  915. 48:36These problems were originally solved
  916. 48:39with a different technique, which was
  917. 48:41the Southern blot technique, which we
  918. 48:43will not cover here because it is a
  919. 48:46technique that is falling into disuse,
  920. 48:48but variations of the PCR technique
  921. 48:50allowed the problem to be solved. To
  922. 48:54avoid the problems of lack of
  923. 48:57amplification that occur with triplet
  924. 49:00expansion when we try to determine it
  925. 49:02with conventional PCR, this variant of
  926. 49:05PCR uses three primers. One of them,
  927. 49:10acting as a forward primer, that is, on
  928. 49:13one of the strands, is placed in the
  929. 49:15flanking region of the amplification.
  930. 49:19Then a second primer is used that will
  931. 49:25be complementary to the repeat region
  932. 49:27and a third primer that will function
  933. 49:30as a reverse primer, that is, in the
  934. 49:32opposite direction, which will bind to
  935. 49:35the complementary strand. Let us
  936. 49:39consider that the use of these three
  937. 49:41primers, in particular, the primer that
  938. 49:44binds to the repeat regions, has the
  939. 49:47following peculiarity. This primer can
  940. 49:51bind anywhere in the repeat region,
  941. 49:54generating diverse amplification
  942. 49:57products when it creates amplicons in
  943. 50:00combination with the primer facing in
  944. 50:03the opposite direction. So, let's think
  945. 50:08that in the PCR reaction we have
  946. 50:12millions of molecules acting as
  947. 50:14templates and many reactions occurring
  948. 50:17simultaneously. Consequently, one would
  949. 50:21expect from this configuration that,
  950. 50:24for example, amplification products are
  951. 50:27generated, uh, of different lengths, as
  952. 50:33we were just saying, when the repeat
  953. 50:35primer acts, and some of full length
  954. 50:37when both flanking primers act.
  955. 50:41Consequently, uh, this technique has
  956. 50:44the advantage that it will allow for
  957. 50:47estimating the presence or absence of
  958. 50:50these expanded allelic variants, uh, by
  959. 50:54two methodologies. Hm. On one hand, uh,
  960. 50:58we will try to see if there is
  961. 51:00amplification of the full expansion
  962. 51:03product, but we will have, uh, as an
  963. 51:06aid in this technique, the presence of
  964. 51:09a series of fragments of different
  965. 51:12sizes that span the different regions
  966. 51:15of the repeat. TP-PCR results are
  967. 51:20usually analyzed by capillary
  968. 51:23electrophoresis. Let us remember that,
  969. 51:28uh, these capillary electrophoresis
  970. 51:30results are usually visualized in these
  971. 51:33graphs called electropherograms, where
  972. 51:36on the vertical axis we observe that
  973. 51:39the height of the peaks, uh, reflects
  974. 51:41the level of signal fluorescence
  975. 51:44intensity, uh, the presence of
  976. 51:46amplification products and their
  977. 51:48intensity. And on the horizontal axis,
  978. 51:52what we see is the variation in the
  979. 51:55sizes of those amplification products.
  980. 51:58In this case, uh, we are observing
  981. 52:01different length ranges of those
  982. 52:03products that allow us to classify
  983. 52:06these allelic variants into the normal
  984. 52:10zone, premutation, or fully expanded
  985. 52:13mutation. Let us analyze how the
  986. 52:17results look in the case of an allelic
  987. 52:19variant that possesses trinucleotide
  988. 52:22repeats in the normal range. Here we
  989. 52:25see, on one hand, an intense
  990. 52:27amplification product of the allelic
  991. 52:30variant in the normal range, which is
  992. 52:33the amplification product of the two
  993. 52:35pairs of primers that are in the
  994. 52:38flanking regions. Hm. And we also see a
  995. 52:42series of peaks in the form of a ladder
  996. 52:44, which are those produced by the
  997. 52:46amplification products where the primer
  998. 52:49acts that can bind to the repeat region
  999. 52:51and, therefore, can yield different
  1000. 52:53amplification products. But this entire
  1001. 52:58ladder of peaks is also restricted to
  1002. 53:00the normal range because the range of
  1003. 53:03repeats that these alleles possess is
  1004. 53:05relatively limited. How is it observed?
  1005. 53:10the the result of the electropherogram
  1006. 53:14when there is an individual who
  1007. 53:16possesses a premutated variant, that is
  1008. 53:18, with an intermediate range of triplet
  1009. 53:21expansion. In the first place, we see
  1010. 53:24that the series of peaks, uh, that is
  1011. 53:27produced as a consequence of the
  1012. 53:29different fragments amplified from the
  1013. 53:32primer that binds to the repeat region
  1014. 53:34is wider. Yes, we have a larger range
  1015. 53:38of amplification of those products and,
  1016. 53:41in addition, we see a stronger signal,
  1017. 53:44uh, of the larger amplification product
  1018. 53:46, the one produced by those primers
  1019. 53:49that flank the repeat region, which is
  1020. 53:52centered at a higher molecular weight.
  1021. 53:56Hm. Here we don't see a single peak
  1022. 53:59because usually, the longer the
  1023. 54:01fragment to be proliferated, the
  1024. 54:03greater the number of errors that occur
  1025. 54:06in the PCR products. Hm. We also see
  1026. 54:11that the peak height has decreased
  1027. 54:13compared to the previous time,
  1028. 54:14precisely because that amplification
  1029. 54:16product is less efficient and the total
  1030. 54:18amount of product we will obtain will
  1031. 54:20be smaller, and its fluorescence
  1032. 54:21intensity, the peak height, will also
  1033. 54:23be lower. Finally, let's compare this
  1034. 54:28with a carrier of an allelic variant
  1035. 54:31with full expansion, where we see that
  1036. 54:34the larger the expanded region, the
  1037. 54:37region of peaks—hm—produced as a
  1038. 54:40product of the amplification of the
  1039. 54:42primer that bound to the repeat region
  1040. 54:45covers a much wider range of molecular
  1041. 54:48weights, and we see that the peak
  1042. 54:52corresponding to the amplification of
  1043. 54:55the full region has a lower height.
  1044. 54:58This means that the efficiency of its
  1045. 55:00production is lower. It was produced in
  1046. 55:03a smaller quantity, but the size of
  1047. 55:06that total product is larger than
  1048. 55:09before and is in the full mutation
  1049. 55:12region. In summary, this TP-PCR
  1050. 55:15technique is a specific adaptation of
  1051. 55:19the PCR technique to detect allelic
  1052. 55:22variants due to triplet expansion. In
  1053. 55:28this second part of the lecture, we are
  1054. 55:30going to study the techniques based on
  1055. 55:33DNA sequencing that are used in the
  1056. 55:35context of medical genetics. This set
  1057. 55:39of techniques focuses on determining
  1058. 55:42the order of base pairs that are
  1059. 55:45present in a given DNA fragment. And
  1060. 55:48within this set of methodologies, we
  1061. 55:51will analyze, first, the sequencing
  1062. 55:54method developed by Frederik Sanger,
  1063. 55:57who won the Nobel Prize in Chemistry in
  1064. 56:001980 thanks to this development, and
  1065. 56:03which, as we will see, was the
  1066. 56:05methodology used to produce the first
  1067. 56:08sequencing of the human genome in the
  1068. 56:11context of the Human Genome Project.
  1069. 56:16Then we will address subsequent
  1070. 56:18developments from the last 15 to 20
  1071. 56:20years that, based on Sanger sequencing
  1072. 56:22methods, allowed it to be expanded and
  1073. 56:25are collectively called high-throughput
  1074. 56:27sequencing methodologies, also known as
  1075. 56:30next-generation sequencing or
  1076. 56:32high-throughput sequencing. To
  1077. 56:36understand the Sanger sequencing
  1078. 56:39methodology, we have to re-analyze some
  1079. 56:44molecular mechanisms that are involved
  1080. 56:46in the polymerization of nucleic acids,
  1081. 56:49since it is the particular biochemistry
  1082. 56:52of polymerization that Sanger used as
  1083. 56:55the foundation for his technique. Let's
  1084. 57:00remember again that during the nucleic
  1085. 57:03acid polymerization process, DNA
  1086. 57:06polymerase uses a single-stranded DNA
  1087. 57:09strand as a template, but it cannot
  1088. 57:12polymerize de novo; it can only extend
  1089. 57:15a previously paired strand that has a
  1090. 57:18free three-prime hydroxyl end. We see
  1091. 57:24on the right side of the slide this
  1092. 57:27three-prime hydroxyl end in greater
  1093. 57:30molecular detail. We are seeing that
  1094. 57:33the last nucleotide added to a chain at
  1095. 57:36the three-prime position of the sugar
  1096. 57:39has a hydroxyl group that will be used
  1097. 57:42to form a new covalent bond with the
  1098. 57:44incoming nucleotide, forming a
  1099. 57:47phosphodiester bond and allowing the
  1100. 57:49incorporation of the new nucleotide.
  1101. 57:54Then, the one that will function as the
  1102. 57:56next new terminal nucleotide of the
  1103. 57:59chain will in turn have another
  1104. 58:01three-prime hydroxyl end that will
  1105. 58:03function in a similar way, that is,
  1106. 58:06allowing the incorporation of the
  1107. 58:08subsequent nucleotide. Taking this into
  1108. 58:13account, Sanger devised using variants
  1109. 58:18of nucleotides that would prevent the
  1110. 58:20development of polymerization. Here we
  1111. 58:25see a standard nucleotide that has the
  1112. 58:29hydroxyl group associated with the 3-
  1113. 58:32prime carbon of the sugar and a variant
  1114. 58:35that was what he used, called
  1115. 58:38dideoxyribonucleotide triphosphates,
  1116. 58:41which lack this hydroxyl group at the 3
  1117. 58:45-prime position. If one, in an in vitro
  1118. 58:50polymerization reaction, adds these
  1119. 58:52dideoxynucleotides, that is,
  1120. 58:54nucleotides that lack this 3-prime
  1121. 58:57hydroxyl end, once they are
  1122. 58:59incorporated, they terminate
  1123. 59:01polymerization because no nucleotide
  1124. 59:03can be incorporated by binding to them,
  1125. 59:06since they lack this functional group
  1126. 59:09that allowed the incorporation of the
  1127. 59:12new nucleotide. How then did this
  1128. 59:16particular biochemistry allow Sanger to
  1129. 59:19devise a methodology for the sequencing
  1130. 59:21of nucleic acids? Let's think about a
  1131. 59:27polymerization reaction that occurs in
  1132. 59:31a laboratory tube. In it, the
  1133. 59:36substrates for polymerization are
  1134. 59:39introduced: a DNA polymerase, the dNTPs
  1135. 59:42, that is, deoxyribonucleotides of
  1136. 59:45adenine, guanine, cytosine, and thymine
  1137. 59:49, and small amounts of
  1138. 59:51dideoxynucleotides of the four classes.
  1139. 59:55Hm. Furthermore, let's think that these
  1140. 1:00:00dideoxynucleotides, which are in much
  1141. 1:00:02smaller amounts than the rest of the
  1142. 1:00:04natural nucleotides, will each be
  1143. 1:00:06labeled with a different fluorophore.
  1144. 1:00:11On the other hand, the template is
  1145. 1:00:14added, which will be a DNA fragment
  1146. 1:00:16whose base sequence one wants to
  1147. 1:00:18determine. You might think that
  1148. 1:00:23information about the sequence of this
  1149. 1:00:27fragment will be necessary to design a
  1150. 1:00:30primer needed to introduce as a
  1151. 1:00:32condition of possibility for the DNA
  1152. 1:00:35polymerase to function. A biochemical
  1153. 1:00:42trick that was used in this context was
  1154. 1:00:45to chemically add to the end of the DNA
  1155. 1:00:48molecule that one wanted to sequence,
  1156. 1:00:52that is, whose nucleotide sequence was
  1157. 1:00:55still unknown, a DNA fragment that we
  1158. 1:00:58are representing here with this green
  1159. 1:01:01line of known sequence. When this known
  1160. 1:01:06sequence fragment is chemically
  1161. 1:01:08attached, one will be able to design a
  1162. 1:01:11complementary primer to this known
  1163. 1:01:13sequence, which will then allow it to
  1164. 1:01:15act as an initiator for the
  1165. 1:01:17polymerization of the region that is
  1166. 1:01:19still unknown. I am not representing it
  1167. 1:01:23here, but let's imagine then that a
  1168. 1:01:25primer complementary to this known
  1169. 1:01:27region, designed by the researcher, is
  1170. 1:01:30also added to this polymerization
  1171. 1:01:32reaction. Let's think about this
  1172. 1:01:36reaction where, although we are
  1173. 1:01:39representing the strand as
  1174. 1:01:41single-stranded, it is originally
  1175. 1:01:44double-stranded, but we will only use a
  1176. 1:01:48single primer from the PCR that binds
  1177. 1:01:52to the strand to be used as a template.
  1178. 1:01:55Hm. Let's imagine as a second step, to
  1179. 1:02:01make the representation of what is
  1180. 1:02:03happening here more complex, that this
  1181. 1:02:06reaction is occurring millions of times
  1182. 1:02:08inside this tube in parallel. What does
  1183. 1:02:11this mean? That there isn't just one
  1184. 1:02:14DNA polymerase and one set of these
  1185. 1:02:16nucleotides, but rather these
  1186. 1:02:18components are there by the millions,
  1187. 1:02:20and the template and primers are also
  1188. 1:02:23there by the millions. So, let's
  1189. 1:02:26imagine that the polymerization
  1190. 1:02:31reaction begins, and then the primers
  1191. 1:02:33—which we are not representing here
  1192. 1:02:36—bind to the known region that was
  1193. 1:02:39attached to one end of the molecule,
  1194. 1:02:42and polymerization begins. Since most
  1195. 1:02:46of the natural components are in excess
  1196. 1:02:49, it is most likely that in the
  1197. 1:02:50majority of the millions of molecules
  1198. 1:02:52we have as a template, polymerization
  1199. 1:02:54will occur completely. But in some of
  1200. 1:02:58the molecules inside that tube, by
  1201. 1:03:01chance, the DNA polymerase might place
  1202. 1:03:04a dideoxy right at the first nucleotide
  1203. 1:03:07to be polymerized. And for that
  1204. 1:03:11molecule which was polymerized with
  1205. 1:03:13that dideoxy, the polymerization of
  1206. 1:03:15that molecule will end there. It will
  1207. 1:03:17not be able to incorporate any more
  1208. 1:03:18nucleotides. The same will happen with
  1209. 1:03:22the second position. Most of the
  1210. 1:03:25molecules continue their polymerization
  1211. 1:03:27, but there will be a small subset that
  1212. 1:03:30will, by chance, incorporate a dideoxy
  1213. 1:03:33nucleotide at the second position of
  1214. 1:03:35polymerization that is complementary to
  1215. 1:03:38the template, and it will be
  1216. 1:03:40polymerized, ending the polymerization.
  1217. 1:03:45This will then occur consecutively at
  1218. 1:03:47each of the polymerization positions,
  1219. 1:03:50and when this reaction ends, we will
  1220. 1:03:52have a complex mixture in that tube of
  1221. 1:03:54fully polymerized molecules; that is,
  1222. 1:03:57those that did not incorporate any
  1223. 1:03:59dideoxy, but there will surely be a
  1224. 1:04:01subset of molecules that have ended at
  1225. 1:04:04every possible position within the
  1226. 1:04:06polymerization. And this happens
  1227. 1:04:10because there are millions and millions
  1228. 1:04:12of molecules that participated in the
  1229. 1:04:14polymerization. And so, statistically,
  1230. 1:04:17it is most likely that at every
  1231. 1:04:20position there is at least a subset of
  1232. 1:04:23molecules represented. What does this
  1233. 1:04:26mean? That all these molecules in which
  1234. 1:04:30polymerization was interrupted at
  1235. 1:04:32different parts will have different
  1236. 1:04:34lengths. Hm. They will vary in length
  1237. 1:04:38one by one. And we know that a
  1238. 1:04:41technique to separate DNA molecules
  1239. 1:04:44according to their size is
  1240. 1:04:45electrophoresis. Hm. Electrophoresis in
  1241. 1:04:49a porous matrix, for example, in an
  1242. 1:04:51agarose gel or in a capillary that has
  1243. 1:04:54a porous matrix. And by separating
  1244. 1:04:59these fragments according to their size
  1245. 1:05:02, each of the fragments, we will be
  1246. 1:05:04able to identify which nucleotide was
  1247. 1:05:06at the end, because each one of them
  1248. 1:05:09emitted a fluorescent light depending
  1249. 1:05:11on the fluorophore that was used,
  1250. 1:05:13marking that so that, by analyzing them
  1251. 1:05:18according to their size, the, uh, the
  1252. 1:05:22light emitted by each of these
  1253. 1:05:24fluorophores will be able to
  1254. 1:05:27reconstruct the sequence of nucleotides
  1255. 1:05:30that made up that fragment. This
  1256. 1:05:34procedure is usually performed using
  1257. 1:05:37not a common gel electrophoresis, but a
  1258. 1:05:41capillary electrophoresis, just as we
  1259. 1:05:44analyzed in the first part of the class
  1260. 1:05:47. And the results are observed as these
  1261. 1:05:52electropherograms, where we see a
  1262. 1:05:56series of fluorescence peaks, and we
  1263. 1:05:59now know that each position implies a
  1264. 1:06:03fragment that differs by one nucleotide
  1265. 1:06:06in length from the other. And the type
  1266. 1:06:11of fluorescent light observed at each
  1267. 1:06:14of these positions can be digitally
  1268. 1:06:17decoded with the nature of one of the
  1269. 1:06:21nucleotides, since the different
  1270. 1:06:23fluorophores report each of the four
  1271. 1:06:26DNA nucleotides. And in this way, we
  1272. 1:06:30can digitally obtain, after this
  1273. 1:06:33analysis of capillary electrophoresis,
  1274. 1:06:36the nucleotide sequence, uh, of the
  1275. 1:06:40fragment that was analyzed by Sanger
  1276. 1:06:43sequencing. Let's observe this
  1277. 1:06:47particularity. Let's think that we have
  1278. 1:06:50obtained three different results
  1279. 1:06:52through Sanger sequencing of a
  1280. 1:06:54particular gene from three individuals.
  1281. 1:06:59These three individuals have high
  1282. 1:07:02similarities in this region, uh, of the
  1283. 1:07:05particular gene that was analyzed, but
  1284. 1:07:08the individual represented in the top
  1285. 1:07:11electropherogram at this position has
  1286. 1:07:14an adenine nucleotide. Hm. On the other
  1287. 1:07:19hand, the individual represented in the
  1288. 1:07:21bottom electropherogram, at this same
  1289. 1:07:24position, has a guanine nucleotide.
  1290. 1:07:26That is, this technique allows us, on
  1291. 1:07:29one hand, to distinguish allelic
  1292. 1:07:31variants that are present in different
  1293. 1:07:33individuals, but I would like to point
  1294. 1:07:36out the individual represented by the
  1295. 1:07:38middle electropherogram, where two
  1296. 1:07:40peaks appear simultaneously at this
  1297. 1:07:43same position. One fluorophore
  1298. 1:07:45reporting that there is adenine at that
  1299. 1:07:47position and another fluorophore
  1300. 1:07:49reporting that there is guanine at that
  1301. 1:07:52same position. To correctly interpret
  1302. 1:07:56this electropherogram, we have to
  1303. 1:07:58remember that our genetic material is
  1304. 1:08:01diploid. Therefore, it could happen
  1305. 1:08:05that if the maternal and paternal
  1306. 1:08:07allelic variants differ in their
  1307. 1:08:10sequence, two different peaks will
  1308. 1:08:12appear at the same position. In other
  1309. 1:08:17words, this double peak being
  1310. 1:08:19represented in the middle
  1311. 1:08:21electropherogram is indicating an
  1312. 1:08:24individual who has this allelic variant
  1313. 1:08:27in heterozygosity. This Sanger
  1314. 1:08:33sequencing method was widely used until
  1315. 1:08:36the year 2000, but one of its
  1316. 1:08:38particularities that can be thought of
  1317. 1:08:40as a limitation is that for each Sanger
  1318. 1:08:43experiment, only one type of molecule
  1319. 1:08:45is sequenced. Right? For example, a
  1320. 1:08:50particular exon of a gene, or an entire
  1321. 1:08:52gene if it is small, can sometimes be
  1322. 1:08:55sequenced. Hm. And while it has high
  1323. 1:09:00sequencing fidelity, its costs are
  1324. 1:09:03relatively high. Hm. This means it is
  1325. 1:09:07only used if there is a clinical reason
  1326. 1:09:11to suspect a relevant variant in a
  1327. 1:09:14specific gene. Hm. Because only by
  1328. 1:09:19being guided by clinical findings or
  1329. 1:09:21this hypothesis will we go on to
  1330. 1:09:23sequence that specific gene or exon. Hm
  1331. 1:09:27. In other words, we would not apply
  1332. 1:09:30the Sanger sequencing technique if we
  1333. 1:09:32have a much more general hypothesis
  1334. 1:09:34that there could be a pathogenic
  1335. 1:09:36variant in the genome, but we do not
  1336. 1:09:39know which gene is involved. For this
  1337. 1:09:43other type of question, high-throughput
  1338. 1:09:47sequencing methods might be more
  1339. 1:09:49suitable than Sanger-based variants.
  1340. 1:09:54What is the fundamental difference?
  1341. 1:09:57Unlike Sanger, which sequenced a single
  1342. 1:10:01type of molecule at a time, these
  1343. 1:10:03methods sequence millions of molecules
  1344. 1:10:07in parallel and could sequence, for
  1345. 1:10:10example, an individual's entire genome
  1346. 1:10:13at a relatively affordable cost and in
  1347. 1:10:16much shorter times. Although we are not
  1348. 1:10:23going to analyze the biochemistry of
  1349. 1:10:26the high-throughput sequencing
  1350. 1:10:29procedure, which is the goal of this
  1351. 1:10:32course, it involves: extracting DNA
  1352. 1:10:35from an individual, fragmenting these
  1353. 1:10:38DNA molecules, sequencing the fragments
  1354. 1:10:41in parallel, and then performing a
  1355. 1:10:45bioinformatic analysis where the
  1356. 1:10:48sequenced fragments are aligned with
  1357. 1:10:51the complete reference genome available
  1358. 1:10:55in public databases. This allows us to
  1359. 1:10:58infer what differences were found in
  1360. 1:11:01the individual from whom the DNA was
  1361. 1:11:03extracted and subjected to this
  1362. 1:11:05procedure, compared to the reference
  1363. 1:11:08human genome. This methodology is
  1364. 1:11:11increasingly being used in the context
  1365. 1:11:14of medical genetics, but it has the
  1366. 1:11:16difficulty that many variants are found
  1367. 1:11:19with respect to the reference human
  1368. 1:11:22genome. However, it is not always
  1369. 1:11:24simple to determine if those variations
  1370. 1:11:27found have a causal relationship with
  1371. 1:11:30the pathology or if they are simply
  1372. 1:11:33non-pathogenic variants that reflect
  1373. 1:11:35the population's genetic diversity. Hm.
  1374. 1:11:40I would finally like to mention
  1375. 1:11:42regarding this technique that not only
  1376. 1:11:44does it allow for whole-genome
  1377. 1:11:46sequencing, but in some contexts,
  1378. 1:11:48subsets of the genome can be sequenced.
  1379. 1:11:51For example, we can focus only on the
  1380. 1:11:54exonic regions of the genome, which are
  1381. 1:11:57a much smaller percentage and allow us
  1382. 1:12:00to focus on variations occurring in
  1383. 1:12:02those coding regions of the genome. To
  1384. 1:12:09finish this class, I would like to put
  1385. 1:12:12into context what some situations might
  1386. 1:12:16be where all the techniques we have
  1387. 1:12:19analyzed today could be relevant. For
  1388. 1:12:23example, let's consider we are
  1389. 1:12:25performing a molecular diagnosis of
  1390. 1:12:28monogenic entities. A first question is
  1391. 1:12:31whether the clinical presentation, the
  1392. 1:12:35signs and symptoms of the individual
  1393. 1:12:38who consulted us, lead us to think
  1394. 1:12:41there is a particular gene responsible
  1395. 1:12:45for the signs and symptoms the
  1396. 1:12:47individual presents. And based on prior
  1397. 1:12:52history, because syndromes or clinical
  1398. 1:12:55characteristics are already described
  1399. 1:12:58that strongly suggest a pathology, we
  1400. 1:13:01will be in certain conditions to make a
  1401. 1:13:04molecular diagnosis that are different
  1402. 1:13:07than if we do not have, for example, a
  1403. 1:13:10strong hypothesis. If we don't have one
  1404. 1:13:13, we probably have to use a
  1405. 1:13:15hypothesis-free technique, right? An
  1406. 1:13:18unbiased technique such as
  1407. 1:13:20high-throughput sequencing. Hm. And
  1408. 1:13:23there, we will eventually find, or not,
  1409. 1:13:25variants in that individual that
  1410. 1:13:28perhaps differ from the human reference
  1411. 1:13:31genome. Hm. And we could develop a
  1412. 1:13:34hypothesis regarding whether that
  1413. 1:13:37variant is causal in producing the
  1414. 1:13:40pathology. For example, we could
  1415. 1:13:43corroborate that hypothesis if this
  1416. 1:13:45same variant has been found in other
  1417. 1:13:47individuals who have the same signs and
  1418. 1:13:50symptoms, or if, in the context of
  1419. 1:13:52basic research, some negative
  1420. 1:13:54functional consequence of the
  1421. 1:13:55appearance of that variant on protein
  1422. 1:13:58function has been demonstrated; for
  1423. 1:14:00example, if that variant is in a coding
  1424. 1:14:02region. But let's imagine the other
  1425. 1:14:07scenario where the signs and symptoms
  1426. 1:14:10allow us to develop a hypothesis
  1427. 1:14:12regarding which gene is involved in
  1428. 1:14:15this disease. The next question is
  1429. 1:14:20whether there are frequent mutations
  1430. 1:14:22already described for that gene. And in
  1431. 1:14:26the case where there are no frequently
  1432. 1:14:28described mutations, one could perform
  1433. 1:14:30Sanger sequencing of that gene and,
  1434. 1:14:32again, see in that particular gene
  1435. 1:14:34suggested by my clinical hypothesis if
  1436. 1:14:36there are variants relative to the
  1437. 1:14:38human reference genome or not, and if
  1438. 1:14:40those variants have been reported as
  1439. 1:14:43pathogenic or not. But in the case of
  1440. 1:14:50having frequent mutations, or if it has
  1441. 1:14:52already been found through other
  1442. 1:14:54methodologies what mutations are
  1443. 1:14:56present in a particular family, one
  1444. 1:14:58could use more biased techniques. In
  1445. 1:15:01other words, if I am looking for a
  1446. 1:15:04particular mutation or allelic variant,
  1447. 1:15:07I could use PCR-based techniques
  1448. 1:15:09depending on the nature of that
  1449. 1:15:11mutation. For example, we have seen
  1450. 1:15:14that if the mutation we are looking for
  1451. 1:15:17is an insertion or a deletion, we could
  1452. 1:15:19perform conventional PCR. Given the
  1453. 1:15:23difference in the size of the
  1454. 1:15:24amplification products, the amplicons
  1455. 1:15:26could report those allelic differences
  1456. 1:15:28to us. We saw that in the case of
  1457. 1:15:32substitution mutations, the appropriate
  1458. 1:15:35variants of the PCR technique are, for
  1459. 1:15:37example, allele-specific PCR or RFLP
  1460. 1:15:40PCR. And in the particular case where
  1461. 1:15:44our clinical context suggests a triplet
  1462. 1:15:47expansion disease, we would use TPPCR,
  1463. 1:15:50or Triple Prime PCR, since the other
  1464. 1:15:53technique, uh, not based on PCR, the
  1465. 1:15:56Southern blot technique, is one that,
  1466. 1:15:59uh, while still used in some contexts,
  1467. 1:16:02is increasingly falling into disuse due
  1468. 1:16:06to its implementation difficulties. We
  1469. 1:16:10also saw that uh or a variant of the
  1470. 1:16:15PCR technique, specifically multiplex
  1471. 1:16:17PCR, is used in the molecular context
  1472. 1:16:20of forensic genetics. Here, there is no
  1473. 1:16:24pathological entity involved; rather,
  1474. 1:16:27the goal is to determine the profile of
  1475. 1:16:30short tandem repeats, particularly
  1476. 1:16:32microsatellites, in the context of
  1477. 1:16:35parentage testing, for example, or
  1478. 1:16:38other techniques associated with
  1479. 1:16:40forensic genetics. With that, we finish
  1480. 1:16:45today's class and we will see each
  1481. 1:16:48other in a future session.

About this transcript

This page contains the full transcript of Seminario 11 Técnicas de Biología Molecular aplicadas a Entidades Monogénicas - Sebastián Giusti by Biología Molecular y Genética FMED - UBA, generated from the public captions YouTube serves with the video. The transcript has 9,023 words across 1,481 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.