YouTube2Text

MPG Primer: Genomic Privacy: Key Issues and Emerging Solutions (2024) — Transcript

by Broad Institute · 8,064 words · 1,251 segments · language en · Watch on YouTube

Full transcript

  1. 0:00great thank you all for being here today
  2. 0:02for our MPG primer how could you
  3. 0:06forget we are very excited to have Dr
  4. 0:09hun Cho here who's an assistant
  5. 0:10professor of biomedical information and
  6. 0:12data science at Yale University he was
  7. 0:15previously here at bro as a Schmid
  8. 0:17fellow and uh principal investigator he
  9. 0:19received his PhD in electrical
  10. 0:21engineering and computer science at MIT
  11. 0:23in 2009 um and before that had been at
  12. 0:25Stanford where he got both a bachelor's
  13. 0:27and master's degree uh in computer
  14. 0:29science with honors he's been named a
  15. 0:31recipient of the NIH director's early
  16. 0:33Independence award
  17. 0:35congratulations um and uh the Cho lab
  18. 0:37now works to create algorithmic
  19. 0:39solutions to tackle a variety of
  20. 0:40computational challenges uh introduced
  21. 0:42by the presence now of these large scale
  22. 0:45heterogeneous um data sets that have uh
  23. 0:48important privacy concerns and so Dr Cho
  24. 0:51is very happy to take your questions
  25. 0:53during the talk today um and so please
  26. 0:55feel free to come up to the microphone
  27. 0:57uh or type in the Q&A on Zoom um and
  28. 1:00thank you again for joining us all the
  29. 1:01way from you yeah thank you Sarah thanks
  30. 1:03everyone for making here in person
  31. 1:05despite the rain I appreciate people on
  32. 1:06Zoom as well um as Sarah mentioned I was
  33. 1:10a fellow here for four years before I
  34. 1:12moved to Yale last September um so I'm
  35. 1:14very familiar with this space it's
  36. 1:15always good to be
  37. 1:17back so yeah in this primary I'll be
  38. 1:19talking about genomic privacy and um
  39. 1:22given my background in computer science
  40. 1:24I'll be mostly sharing my perspectives
  41. 1:26on algorithmic and computational side of
  42. 1:28things so just wanted to mention that up
  43. 1:29front
  44. 1:30uh in the first part of this talk I'll
  45. 1:33be giving you an overview of privacy
  46. 1:34risks that are known to be associated
  47. 1:36with genomic data and then in the second
  48. 1:38part I'll talk about some of these uh
  49. 1:40emerging Solutions based on what are
  50. 1:42called privacy enhancing Technologies
  51. 1:44for addressing these risks and
  52. 1:46throughout the talk I'll try to both uh
  53. 1:48you know Define the the key Concepts in
  54. 1:50these areas and also give you some
  55. 1:53vignettes of our recent work on these
  56. 1:55areas as well so I want to first start
  57. 1:58by talking about what we mean when we
  58. 2:00say privacy and distinguish it from a
  59. 2:03another closely related term security so
  60. 2:06privacy in general is a right to be let
  61. 2:08alone that's a common way to define it
  62. 2:11and it also means that an individual has
  63. 2:12control over how information about them
  64. 2:16uh is collected shared or used in
  65. 2:19different settings on the other hand
  66. 2:21security is generally about a property
  67. 2:23of a system that provides protection
  68. 2:25against unauthorized attempts at
  69. 2:27accessing the data or tampering with it
  70. 2:30and also tampering with the uh the
  71. 2:31functionality of these systems so these
  72. 2:33are closely related as you can tell um
  73. 2:37but one does not necessarily imply the
  74. 2:39other and we're generally interested in
  75. 2:41developing systems that are both secure
  76. 2:43and provide protection uh for the
  77. 2:44privacy of the underlying
  78. 2:47individuals and as you know privacy is
  79. 2:49an important topic in biom medicine and
  80. 2:51genomics uh in general and this is
  81. 2:54mostly due to the fact that biomedical
  82. 2:57information it contains something that
  83. 2:59very that's very
  84. 3:00personal to us uh about our biology and
  85. 3:04health and the thing that makes this
  86. 3:06problem so complex is that there are all
  87. 3:08these different types of biomedical data
  88. 3:10types that uh reveal different things
  89. 3:13about our biology so including things
  90. 3:15like electronic health records medical
  91. 3:17images personal genomes and more
  92. 3:19recently data from wearable devices
  93. 3:21there's also these variety of omix data
  94. 3:24types transcriptomics proteomics and
  95. 3:26also metagenomic data from microbiomes
  96. 3:29so all these different types tell us
  97. 3:31different things about um human health
  98. 3:33but at the same time they increase the
  99. 3:34stes uh the stakes for keeping all this
  100. 3:37information safe and protect it from
  101. 3:39potential
  102. 3:41misuse and this is not just hypothetical
  103. 3:44concerns there um many real world
  104. 3:46examples of privacy threats you may have
  105. 3:48seen in the news uh a recent privacy
  106. 3:50breach data breach at 23 andme last year
  107. 3:53which affected millions of users the
  108. 3:55attackers were able to access raw
  109. 3:57genotype data as well as health related
  110. 3:59information like disease predisposition
  111. 4:01status as well as genetic relative
  112. 4:03information for these users is quite
  113. 4:06concerning and it's also even more
  114. 4:08concerning that the attackers were able
  115. 4:10to create these curated lists targeting
  116. 4:12specific groups of individuals and they
  117. 4:15were s selling these data on the dark
  118. 4:18web here's another more recent example
  119. 4:20from just a few weeks ago I think the
  120. 4:22situation is still unfolding there was a
  121. 4:24large uh ransomware attack at change
  122. 4:27Healthcare which is a prescription
  123. 4:29processing arm of this insurance company
  124. 4:32called United
  125. 4:34Health it was reported that around 94%
  126. 4:38of hospitals in the United States were
  127. 4:39affected financially due to this attack
  128. 4:42and there's still an ongoing
  129. 4:43investigation into what kinds of data
  130. 4:46had been accessed uh through this
  131. 4:49attack so there are real threats and I
  132. 4:52want to make the point that this is not
  133. 4:53just a concern in uh the in the
  134. 4:55commercial
  135. 4:56domain those of us uh who are involved
  136. 4:59in academic research and genomics
  137. 5:00genomics should also care about privacy
  138. 5:02too and here are some reasons so behind
  139. 5:04every data set there are real people who
  140. 5:06could be affected by a potential privacy
  141. 5:08breach so we have a certain amount of
  142. 5:11responsibility to uh look for safer ways
  143. 5:13to conduct research and second these
  144. 5:16privacy concerns and related regulations
  145. 5:18they make it difficult to for
  146. 5:21researchers across different
  147. 5:23institutions to collaborate and share
  148. 5:24data so having better ways to protect
  149. 5:26privacy can help us uh break down these
  150. 5:29barriers
  151. 5:30and acceler
  152. 5:32Science and lastly if we think about
  153. 5:35converting biomedical insights from from
  154. 5:37uh Research into public health impact
  155. 5:39for example for supporting health
  156. 5:41related decision- making for private
  157. 5:42individuals we need reliable and
  158. 5:45trustworthy systems that can process
  159. 5:47private information so having better
  160. 5:49safeguards can help us in that regard as
  161. 5:52well so now let's talk about what's
  162. 5:55known about genomic privacy in more
  163. 5:57detail um so there are Key Properties of
  164. 5:59genetic data uh that make it unique and
  165. 6:03different from other types of personal
  166. 6:05information like passwords or online
  167. 6:07search history uh and what have you
  168. 6:10these are six of these properties that
  169. 6:12were summarized by uh this review
  170. 6:14article by naita at
  171. 6:16all I'll briefly go through them so uh
  172. 6:19genetic data it contains information
  173. 6:21about our health and behavior as I
  174. 6:22mentioned this is static kind of
  175. 6:24information it's mostly determined at
  176. 6:26Birth and it's hard to change it's also
  177. 6:28very unique to each IND idual so only
  178. 6:30tens of snips let's say would be enough
  179. 6:32to single out an individual from the
  180. 6:34global human population Mystique is
  181. 6:36referring to the fact that there's still
  182. 6:38a lot that we don't know about what can
  183. 6:39be inferred from genetic information
  184. 6:41these data have intrinsic value because
  185. 6:44they can be used for biomedical research
  186. 6:45or clinical purposes and lastly kinship
  187. 6:48means that when a privacy breach happens
  188. 6:50it reveals information not just about
  189. 6:52the affected individuals but also about
  190. 6:54their biological relatives so when we're
  191. 6:56talking about risks and safeguards it's
  192. 6:58important that we consider these uh
  193. 7:00properties of the
  194. 7:01data and there's been a lot of
  195. 7:03literature on different ways of
  196. 7:05breaching genomic privacy and this is a
  197. 7:08categorization that was provided in this
  198. 7:09other review article in nature genetics
  199. 7:11from a few years ago and they broadly
  200. 7:14classify the the attacks into two
  201. 7:16categories identification and phenotype
  202. 7:18inference uh I think this is fairly
  203. 7:20self-explanatory but just to Define it
  204. 7:22uh for clarity identification means that
  205. 7:24you start with a supposedly Anonymous
  206. 7:26data instance and we're uh uncovering
  207. 7:29the real IND idual behind that data
  208. 7:30instance pype inference means we start
  209. 7:33with some piece of data about an IND
  210. 7:35individual and we infer additional
  211. 7:37attributes that are not included in the
  212. 7:38data set so including their health
  213. 7:39status disease predisposition things
  214. 7:41like that so in the next few slides I'll
  215. 7:44go over some of the seminal works that
  216. 7:46explore different types of privacy
  217. 7:48attacks in this
  218. 7:49domain one of the earliest examples was
  219. 7:52from 2018 uh when next Generation
  220. 7:55sequencing first came about this attack
  221. 7:57demonstrated the difficulties of asking
  222. 7:59phenotypic information and genetic data
  223. 8:02so uh the first personal genome that was
  224. 8:05released was uh from Dr James Watson's
  225. 8:08um sample and in this release data set
  226. 8:12variance in the apoe gene which is known
  227. 8:14to be linked to alzheimer's disease were
  228. 8:16masked due to personal requests from Dr
  229. 8:19Watson but just a few months later
  230. 8:21another group of researchers show that
  231. 8:23these were actually not really Mass uh
  232. 8:24you could infer these variants using
  233. 8:27nearby variants using linkage dis
  234. 8:29equilibrium patterns in other words we
  235. 8:31could impute these
  236. 8:32genotypes so this led to a larger window
  237. 8:35around this Gene of two uh meab base
  238. 8:38pairs uh to be excluded from the data
  239. 8:41set and note that this is just for
  240. 8:43hiding one Gene and so it's going to be
  241. 8:46really difficult if you want to hide
  242. 8:47information about much larger array of
  243. 8:50disease Rel genes that we know uh that
  244. 8:52are scattered across the
  245. 8:55genome and from the same year there was
  246. 8:57also this seminal work by Homer at
  247. 9:00where they showed that even aggregate
  248. 9:02level information like minor Al
  249. 9:04frequencies and gwos summary statistics
  250. 9:07can reveal private information as well
  251. 9:09and what they were able to show is that
  252. 9:10given a Target subject's genome you
  253. 9:12could compare it against these released
  254. 9:14statistics to figure out whether the
  255. 9:16person was in the underlying cohort or
  256. 9:19not and this is also called membership
  257. 9:21inference attack and while this
  258. 9:23information about membership might not
  259. 9:25be too concerning by itself uh sometimes
  260. 9:27you study course or associate with
  261. 9:29sensitive attributes like rare medical
  262. 9:32conditions so that could lead to a
  263. 9:33potential privacy breach and uh which is
  264. 9:36a major concern
  265. 9:38here and uh so the intuition here behind
  266. 9:41this attack is that if if the subject
  267. 9:44has some rare variants that are uh not
  268. 9:48really found in the general population
  269. 9:49at high frequency then if these variants
  270. 9:51are represented highly in the released
  271. 9:54data then it makes it more likely that
  272. 9:55the subject wasn't the in the original
  273. 9:57data set that's the main idea
  274. 10:00in a more recent work people have shown
  275. 10:03that it's also possible to demonstrate
  276. 10:05this type of leakage uh with genomic
  277. 10:07Beacon Services where uh we answer
  278. 10:09questions like uh is is there a person
  279. 10:12with this particular mutation in the
  280. 10:14underlying data set and you can think of
  281. 10:16this as a different type of aggregate
  282. 10:19level information that's focusing on
  283. 10:21individual mutations and the idea here
  284. 10:24is the Same by combining information
  285. 10:26across different queries we can try to
  286. 10:27single out a uh an
  287. 10:31individual another type of attack that's
  288. 10:34been um receiving a lot of attention in
  289. 10:37the media uh a few years ago is
  290. 10:40identification attack based on public
  291. 10:42genealogy databases uh like GED match so
  292. 10:45these Services include genetic profiles
  293. 10:47from millions of users and their
  294. 10:50genealogical
  295. 10:51relationships and this is a great
  296. 10:53resource for recreational purposes for
  297. 10:55finding one's Roots but they can also be
  298. 10:56used for other purposes uh like police
  299. 10:58officers to catch um criminals based on
  300. 11:01the sample that they get from the crime
  301. 11:03scenes and the the main idea here is
  302. 11:05that you again you start with a genetic
  303. 11:07profile from a Target subject and then
  304. 11:09we create against these databases to
  305. 11:11find out their uh close or distant
  306. 11:13relatives in these databases and we know
  307. 11:15who they are so now we can try to
  308. 11:17triangulate and figure out who the
  309. 11:18person was uh sometimes additionally
  310. 11:20using information like their um their
  311. 11:23demographic information like age or
  312. 11:26rough geographic area where the person
  313. 11:28is
  314. 11:29and there's been several papers that
  315. 11:32explore different aspects of this uh
  316. 11:34threat Factor so this first paper by
  317. 11:36jimar at all um I think Melissa came
  318. 11:39here to give it talk a few months ago
  319. 11:42last year and uh so this work showed
  320. 11:45that just using short tandem repeats on
  321. 11:47chromosome y um You can predict the
  322. 11:50surname of the target individual later
  323. 11:54early Cal show that uh the majority of
  324. 11:57individuals of European descent are
  325. 11:59susceptible to an identity infer attack
  326. 12:02like the one that I explained and this
  327. 12:03is indeed how um the Golden State killer
  328. 12:05in San Francisco was caught if you're
  329. 12:07familiar with that
  330. 12:08story and more recently NL showed that
  331. 12:12uh it's actually not just the the target
  332. 12:14subject whose ident identity could be
  333. 12:16revealed the individuals in the database
  334. 12:19itself could also have their genomes
  335. 12:21leaked U through what's being revealed
  336. 12:22by these services and uh they also
  337. 12:25demonstrated that one can insert these
  338. 12:28uh artificial samples into these
  339. 12:30databases to Mis direct uh police
  340. 12:32investigators who trying to find the the
  341. 12:35the
  342. 12:37perer so going Beyond genotype profiles
  343. 12:41there's also leakage in other functional
  344. 12:43genomic data that we need to think about
  345. 12:45uh so I'm referring to gene expression
  346. 12:47data protein expression things like that
  347. 12:50the idea here is that even though these
  348. 12:52other types of molecular measurements
  349. 12:54tend to be a lot noisier they contain
  350. 12:56enough identifying information so that
  351. 12:58we can link uh samples from the same
  352. 13:01individual across different data sources
  353. 13:04and this is called a linkage
  354. 13:05attack and going beyond that we can also
  355. 13:08try to predict genotypes from these data
  356. 13:11through what are called quantitative
  357. 13:12trait low size or
  358. 13:14qtls these are genetic variants that are
  359. 13:16correlated with other molecular
  360. 13:19measurements and by predicting genotypes
  361. 13:22of the person based on their omix data
  362. 13:23we can also create a linkage between
  363. 13:25their genetic profile data in one data
  364. 13:27set against another omix data in a
  365. 13:30different data
  366. 13:31set and as I mentioned uh these
  367. 13:34possibilities have been demonstrated for
  368. 13:36different types of uh data so including
  369. 13:38transatomic proteomics and microbiome
  370. 13:41data and next I want to describe a
  371. 13:43little bit about our recent work from
  372. 13:44last year in genome research where we
  373. 13:46try to uh study this question of um this
  374. 13:51topic of transatomic data privacy a
  375. 13:52little
  376. 13:54better so again we're interested in
  377. 13:56looking at these quantitative tra low
  378. 13:58size uh but for gene expression so these
  379. 14:00are called
  380. 14:01eqtls and the concern is that a an
  381. 14:05attacker a potential attacker can start
  382. 14:07with the gene expression level that
  383. 14:09they've obtained and try to guess the
  384. 14:11genotypes of that person and there are
  385. 14:13many of these eils in the genome uh so I
  386. 14:16think it's known that the majority of
  387. 14:17human genes have some utl associated
  388. 14:20with them so you can imagine extracting
  389. 14:23a substantial portion of one's uh DNA
  390. 14:27through uh this approach
  391. 14:30the main issue is that the existing
  392. 14:32Works um on this
  393. 14:34idea were focusing on demonstrating this
  394. 14:37possibility and not really fully
  395. 14:39characterizing the extent of leakage so
  396. 14:42there were limitations in their
  397. 14:43statistical models that allow them to
  398. 14:45only analyze a smaller subset of eqls
  399. 14:48that are statistically independent so
  400. 14:50this is the problem that we try to
  401. 14:51address and we introduce this model
  402. 14:53that's working at the sequence level uh
  403. 14:55so we're able to look at all the DLS
  404. 14:58accounting for their corations as well
  405. 15:00for those of you who are familiar with
  406. 15:01this we're starting with this Le and
  407. 15:02Stevens hden Markoff model which um
  408. 15:05basically puts a distribution over an
  409. 15:07entire genetic sequence accounting for
  410. 15:09Link Link this equilibrium and we're
  411. 15:12adding another layer on top of this
  412. 15:14model or on the bottom here uh that's
  413. 15:17corresponding to gene expression
  414. 15:18variables and we create these links that
  415. 15:20capture statistical correlation between
  416. 15:22genotypes and gene expression and note
  417. 15:24that this is again including all theuts
  418. 15:26that are known and we train this model
  419. 15:29in an end to-end fashion so that the
  420. 15:31probabilities are calibrated to account
  421. 15:33for uh correlation between nearby
  422. 15:37genotypes and the idea is that this
  423. 15:39using this model we can produce a match
  424. 15:41score for a given pair of gene
  425. 15:43expression and genotype profile and this
  426. 15:46could be used to try to link individuals
  427. 15:48across data sets uh which can
  428. 15:50potentially lead to lead to
  429. 15:52reidentification
  430. 15:54so in terms of evaluating this model we
  431. 15:57took a data set we had pairs of these
  432. 15:59samples gen expression genotype samples
  433. 16:02uh from a set of individuals around 300
  434. 16:05individuals here in this case we
  435. 16:06included additional genotype profiles
  436. 16:08from a larger uh um data set in this
  437. 16:12case it was hype reference Consortium
  438. 16:14data and this was a much larger set so
  439. 16:1622,000 individuals in the um in the
  440. 16:19additional cohort we're trying to figure
  441. 16:21out for a given gen expression profile
  442. 16:24uh which genotype profile it corresponds
  443. 16:26to and what I'm showing you here is the
  444. 16:28fraction of these individuals who were
  445. 16:30able to correctly link between these two
  446. 16:32samples as we add more of the additional
  447. 16:35candidate um genotype profiles to the uh
  448. 16:39to the data
  449. 16:40set and I'm comparing our method DSM to
  450. 16:44two other Baseline methods ebl and gmbb
  451. 16:46these are the the previous approaches
  452. 16:48that I briefly mentioned that uses an
  453. 16:50independent set of
  454. 16:53etls the main thing to note here is that
  455. 16:55even after including this large number
  456. 16:58of individuals so uh 22,000 individuals
  457. 17:01we're still able to link around 90% of
  458. 17:03more of the original individuals based
  459. 17:05on gen expression information
  460. 17:08only and uh our method is able to link a
  461. 17:12greater fraction compared to the
  462. 17:13Baseline and this Gap becomes more
  463. 17:16pronounced if we consider a a setting
  464. 17:20that's using a smaller set of eils as
  465. 17:22well so just uh based on Chism 20 um the
  466. 17:26overall numbers go down a lot but the
  467. 17:28gap our method and the previous
  468. 17:29approaches uh becomes
  469. 17:33clearer the other thing that we can do
  470. 17:35here is that we can measure the
  471. 17:37confidence of our prediction of the
  472. 17:39links and this is a metric that we
  473. 17:42developed uh it's basically a P value in
  474. 17:45some sense that measures the gap between
  475. 17:46the best match and the second best
  476. 17:48match and if we compare this between our
  477. 17:51method and the previous methods you can
  478. 17:52also see here that uh our method is on
  479. 17:55the y- axis the previous ones are on the
  480. 17:57xaxis the the the confidence that we
  481. 17:59have over the links that we produce tend
  482. 18:01to be
  483. 18:03greater and what this leads to is that
  484. 18:06if we're working in a setting where we
  485. 18:08don't know if the matching individual is
  486. 18:10in the candidate pool or not we have to
  487. 18:12draw a threshold somewhere right and
  488. 18:14have some uh number of false positives
  489. 18:17and uh because of this higher confidence
  490. 18:20uh we also get predictions that are um
  491. 18:24that minimizes the number of mistakes
  492. 18:27compared to the the through mattress
  493. 18:29that we're able to find as a trade-off
  494. 18:30fing Precision
  495. 18:33recall and now I want to describe
  496. 18:35another recent work which is
  497. 18:37illustrating a different type of um uh
  498. 18:40privacy risk and I'm bringing this up
  499. 18:42because I think this is a th Vector that
  500. 18:44not a lot of us uh you know have been
  501. 18:48aware of so it's kind of a neat idea I
  502. 18:51think that we should all be aware of and
  503. 18:54so here we're looking at the security of
  504. 18:56genotype imputation servers I think many
  505. 18:59of you know about this and so this is
  506. 19:01referring to services like topmed and
  507. 19:03Michigan imputation servers what they
  508. 19:05generally do is we provide a uh a
  509. 19:08partially observed genotype profile from
  510. 19:11array uh microarray genotyping platforms
  511. 19:14for example and then we obtain an
  512. 19:17imputed genotypes that fills in all
  513. 19:18these missing positions and the way
  514. 19:21we're able to do this is by using a
  515. 19:23large reference panel of high quality
  516. 19:26genomes and using linkage dis
  517. 19:28equilibrium pattern uh between nearby
  518. 19:30Snips and the question that we asked
  519. 19:33was this process is using some
  520. 19:36information from the reference panel to
  521. 19:37be able to fill in the missing data but
  522. 19:39how much information is actually being
  523. 19:41revealed uh throughout this process and
  524. 19:43included in the output and in fact can
  525. 19:45we actually try to reconstruct some of
  526. 19:46these sequences in the reference panel
  527. 19:48by interacting with these servers in
  528. 19:50certain
  529. 19:52ways and just to avoid any doubt we were
  530. 19:55we took great care to do this work um
  531. 19:58respon II L and ethically so none of
  532. 20:00this work was performed on these servers
  533. 20:01it was all done in a local simulation
  534. 20:03environment we also informed the
  535. 20:05stakeholders well in advance actually
  536. 20:07almost a year in advance and uh we're
  537. 20:10aware that additional security measures
  538. 20:12have been introduced after we published
  539. 20:14our work and all the information that
  540. 20:15I'm sharing today is public and it's
  541. 20:17published last
  542. 20:19year okay so the the key idea that led
  543. 20:22to this work is actually the fact that
  544. 20:24if you plug in an input sequence that
  545. 20:27uniquely matches with a single reference
  546. 20:28sequence it will extract the the
  547. 20:31matching person's genome um almost
  548. 20:34exactly so this is this uh this figure
  549. 20:38is showing that idea so uh in this
  550. 20:40bottom left corner you're seeing an
  551. 20:42example output from the imputation tool
  552. 20:45um we're showing it as predict the
  553. 20:47dosages which is a value between zero
  554. 20:49and one and you can see this value
  555. 20:51across the the genomic window that's
  556. 20:52being imputed on the
  557. 20:54xaxis if there's a single match the
  558. 20:56predictions are very confident it's near
  559. 20:59zero near one and this corresponds to
  560. 21:01the actual genotypes of the matching
  561. 21:03person what's even more surprising is
  562. 21:05that if you have a match um have a query
  563. 21:08that matches a small number of reference
  564. 21:10samples you can actually see this unique
  565. 21:14pattern that corresponds to that too so
  566. 21:15if it matches three samples then you'll
  567. 21:17see these predicted dosages that are
  568. 21:19aggregated around multiples of 1/3s so
  569. 21:22what this means is that you can look at
  570. 21:23the imputation output and figure out how
  571. 21:24many mattress there were in the
  572. 21:26reference panel and what does these to
  573. 21:28is this type of um potential attack
  574. 21:31pipeline where we start with a lot of
  575. 21:33random queries impute them and see what
  576. 21:36we get in terms of these dosage outputs
  577. 21:39and figure out which one of them uh was
  578. 21:41unique which ones had multiple matches
  579. 21:43and for those with multiple matches we
  580. 21:45can try to extend the query to make it
  581. 21:46more unique and in the end we pull
  582. 21:49together all of the unique matches and
  583. 21:50this corresponds to a set of um path
  584. 21:53type sequences from the reference
  585. 21:57panel and using 1,000 genomes data as
  586. 22:01the reference panel uh as an example we
  587. 22:03were able to demonstrate that this type
  588. 22:04of attack can actually extract a fair
  589. 22:06amount of sequences from the reference
  590. 22:08panel so for example if we impute 1
  591. 22:12million um input queries that we
  592. 22:14constructed in some way we're able to
  593. 22:17extract almost 80% of the uh the
  594. 22:20sequences in the reference panel there's
  595. 22:22some amount of error that's happening
  596. 22:24here because of the uh Mis
  597. 22:26classifications um in the pipeline but
  598. 22:29this error rate tends to be fairly
  599. 22:32low and here's a slightly different
  600. 22:35setting where you know to try to prevent
  601. 22:38this attack what we could do is instead
  602. 22:40of releasing the predicted probabilities
  603. 22:44we can uh release just the discrete
  604. 22:46genotype predictions so either zero or
  605. 22:49one nothing in between but even in the
  606. 22:51setting there is a workr that we that
  607. 22:53allows us to extract unique matches so
  608. 22:56the general idea is that you can take
  609. 22:59the output um sequence introduce
  610. 23:02additional
  611. 23:04changes or or keep a subset of positions
  612. 23:07uh mask again some of the other
  613. 23:09positions imputed again if the results
  614. 23:12stays the same then it's more likely
  615. 23:14that the original sequence that was
  616. 23:15output was corresponding to a unique
  617. 23:17individual so using this type of idea we
  618. 23:19can still extract sequences um at
  619. 23:21slightly lower Effectiveness as you can
  620. 23:25see one big caveat here is that
  621. 23:28sequences that we're getting here are
  622. 23:30chunked from the private uh genomes it's
  623. 23:33not the entire genome and this is
  624. 23:35because imputation pipeline is run
  625. 23:36independently on uh smaller genomic
  626. 23:39Windows typically around 20 megabase
  627. 23:42pairs but it turns out that you can
  628. 23:44start from these fragments of the genome
  629. 23:47and Link them across different Windows
  630. 23:49using relatedness patterns and uh this
  631. 23:51is just a visualization I won't go into
  632. 23:53details here that if you compute the
  633. 23:55kinship coefficient between these
  634. 23:57individual pieces with um relative of a
  635. 24:00certain degree you can see that the
  636. 24:02kinship is uh separated from is far
  637. 24:06enough from zero that we can distinguish
  638. 24:08it from unrelated
  639. 24:10samples and what this means is that we
  640. 24:12can imagine a setting where we have
  641. 24:15access to a secondary data set which
  642. 24:17might potentially include relatives of
  643. 24:20the individuals who are in the reference
  644. 24:22panel and then we could compute these
  645. 24:25kinship coefficients uh it's written as
  646. 24:27semi kinship because it's a modified
  647. 24:28version that's working with half types
  648. 24:30instead of typo types so we can compute
  649. 24:33these K coefficients between Pairs of
  650. 24:35samples between t between these two
  651. 24:38sources and then we feed this into our
  652. 24:40prisic linking algorithm that's taking
  653. 24:42this Matrix as input and then outputs uh
  654. 24:44groups of sequences that are likely to
  655. 24:46have come from the same individual so
  656. 24:48that's the overall outline of the attack
  657. 24:53factor and uh we were able to show that
  658. 24:57um given a data set with um real
  659. 24:59relatives that are known to be included
  660. 25:01in the relatives that we were able to
  661. 25:03link a substantial portion of the parget
  662. 25:06genome um up to third degree relative so
  663. 25:10beyond that the signal becomes too weak
  664. 25:12so we can't like link too many uh
  665. 25:16samples and based on this data we can
  666. 25:19extrapolate to get a sense of what
  667. 25:21fraction of the genome could be linked
  668. 25:23for different proportions of the
  669. 25:24reference panel individuals given
  670. 25:27different sizes of this uh relative set
  671. 25:29that I was referring to um relative to
  672. 25:32the overall size of the population that
  673. 25:33the reference panel came from so for
  674. 25:35example if we're given a relative set
  675. 25:37that includes 0.5% of the underlying
  676. 25:40population uh we
  677. 25:42could expect to link 17% of the genome
  678. 25:46for at least 5% of the individuals in
  679. 25:48the reference panel so that's the one of
  680. 25:50the estimates that we were able to get
  681. 25:52and this is a realistic number given
  682. 25:54that 05% is comparable to size of the UK
  683. 25:57biank for example
  684. 26:00so just to summarize this portion
  685. 26:03um I've shown you different types of
  686. 26:05genomic pry risks we looked at
  687. 26:08reidentification attacks membership
  688. 26:09phenotype inference data linkage as well
  689. 26:12as data
  690. 26:13reconstruction and the overall takeaway
  691. 26:16is that our understanding is still
  692. 26:19rapidly shifting and this needs to
  693. 26:21continue to change as we gain access to
  694. 26:23different types of data different models
  695. 26:25uh more advanced Ai and ml tools as well
  696. 26:28as well as other types of data sharing
  697. 26:30systems that are emerging and given all
  698. 26:33these changes it's important to have
  699. 26:34these rigorous approaches to analyze
  700. 26:36these risks and assess them uh using new
  701. 26:40models and come up with new ways to
  702. 26:42address these
  703. 26:43risks so on that point in the second
  704. 26:46part I'll give you some ideas of some of
  705. 26:48the emerging Technical Solutions for
  706. 26:49addressing use
  707. 26:54risks I think you're all familiar with
  708. 26:56these existing policy Frameworks like
  709. 26:58common role privacy uh Hippa privacy
  710. 27:01role gdpr they Define these Notions that
  711. 27:04we all know and love like IRB review
  712. 27:06processes informed consent the the idea
  713. 27:09of protected health
  714. 27:12information and uh there's no question
  715. 27:15that these Frameworks provide an
  716. 27:16important uh tool for us to safeguard
  717. 27:19the use of biomedical data but it's also
  718. 27:21important to note some of the inherent
  719. 27:23limitations to these regulatory
  720. 27:25approaches so the first thing is that
  721. 27:27designing and updating these uh
  722. 27:29regulations is generally a slow and
  723. 27:31long-term process for a good reason um
  724. 27:33one example is that genetic data was
  725. 27:35classified as protected health
  726. 27:37information fora only in 2013 several
  727. 27:41years after genetic data have become
  728. 27:43more widely available there are also
  729. 27:45these key ambiguities in some of the
  730. 27:47terms uh that we use in these
  731. 27:49regulations right so when we say
  732. 27:50deidentified or anonymized data what
  733. 27:53does that actually mean and how does it
  734. 27:54apply to the data sets that we're
  735. 27:56actually dealing with and
  736. 27:58it you know as I've demonstrated
  737. 28:00hopefully a lot of the bical data types
  738. 28:03it's not possible to completely de
  739. 28:05identify it there's always going to be
  740. 28:06some amount of ident identifying
  741. 28:08information sorry uh so we need to draw
  742. 28:11a line somewhere when we think about um
  743. 28:13laws and regulations and perhaps the
  744. 28:16most important uh Pitfall is that these
  745. 28:19approaches don't resolve the core
  746. 28:21conflict between the need to share data
  747. 28:23and the need to uh protect privacy and
  748. 28:25this is because uh they're typically
  749. 28:27about limiting the scope scope of um
  750. 28:29data sharing and usage depending on the
  751. 28:32context uh in order to provide PR
  752. 28:36privacy so it turns out that uh there's
  753. 28:39a collection of techniques from the
  754. 28:41computer science literature which are
  755. 28:43referred to as privacy enhancing
  756. 28:44technologies that can help us resolve
  757. 28:46this conflict and at a high level these
  758. 28:49provide us with mathematical techniques
  759. 28:52uh that can help us use and share
  760. 28:54private biomedical data while at the
  761. 28:56same time protecting privacy so this is
  762. 28:58what I'll tell you about briefly
  763. 29:01next um I'm going to highlight five of
  764. 29:04these technologies that belong to this
  765. 29:05category I'm going to start with homor
  766. 29:07for encryption which is a type of
  767. 29:09encryption
  768. 29:11technique and it's a special form where
  769. 29:14the uh the structure of the encryption
  770. 29:16allows us to perform operations directly
  771. 29:19on the encrypted data without having to
  772. 29:21decrypted
  773. 29:23First initially the uh these Frameworks
  774. 29:28were limited to performing uh limited
  775. 29:30types or limited numbers of operations
  776. 29:32on the private data but there's been
  777. 29:34remarkable progress over the years and
  778. 29:35now we have these practical systems that
  779. 29:37can help us Implement uh fairly complex
  780. 29:40analytic
  781. 29:41tests but they do come with a a
  782. 29:43computational overhead because now we're
  783. 29:45operating with encrypted data and that's
  784. 29:47a lot more expensive than working with
  785. 29:49non-encrypted
  786. 29:51data a slightly different approach as
  787. 29:53SEC multiparty computation it has two
  788. 29:55acronyms NPC or SMC depending on who you
  789. 29:58ask and the core of this framework is
  790. 30:03not encryption but we're going to use
  791. 30:05this technique called secret sharing and
  792. 30:06what it means is for each private number
  793. 30:08we um divide it into a set of random
  794. 30:11numbers that somehow all add up to the
  795. 30:14private number and we're going to split
  796. 30:16them up and then distribute it to
  797. 30:18multiple parties so they all together
  798. 30:20know uh well the shares collectively
  799. 30:23encode information about the private
  800. 30:25number but if you look at each party's
  801. 30:28information they don't know what the
  802. 30:29private number is and the idea is that
  803. 30:31we can uh design these interactive
  804. 30:33protocols where these parties work
  805. 30:35together to perform computation on the
  806. 30:37underlying secret without revealing
  807. 30:39anything throughout this process so at
  808. 30:42the end of a protocol like this we get
  809. 30:43the the um output of the analysis
  810. 30:45results as secret shares as well which
  811. 30:47we can combine to reveal the
  812. 30:50output so this is not using encryption
  813. 30:52so the computational operations tend to
  814. 30:55be cheaper than home or encryption but
  815. 30:59at the cost of uh greater communication
  816. 31:02burden so that's the
  817. 31:06tradeoff trusted execution environment
  818. 31:08or te it's a hardware approach to secure
  819. 31:12computation uh the most well-known
  820. 31:13examples are Intel sgx and amdv
  821. 31:16Technologies and what they provide as a
  822. 31:18way to uh create these isolated
  823. 31:20environments in a computer that protects
  824. 31:24the data inside it confidential uh keeps
  825. 31:27it uh confidential even from the person
  826. 31:29who's operating on these
  827. 31:32machines and the idea is that we can use
  828. 31:34this to send private data into this
  829. 31:36environment isolated environment of
  830. 31:38secure Enclave do some processing on
  831. 31:40them and return the results to the user
  832. 31:43uh through encrypted channels so that no
  833. 31:45private information is leaked during the
  834. 31:48process the key feature of these
  835. 31:50Technologies is what's called remote
  836. 31:52adastation and this is the use of
  837. 31:54cryptographic signatures that actually
  838. 31:56allows us to verify that um the The
  839. 32:00Enclave environment has been created
  840. 32:02faithfully and also the fact that the
  841. 32:04program that's running inside it is what
  842. 32:06we initially uh assigned to it to
  843. 32:11run so in this environment we're working
  844. 32:14with non- encrypted data so there's no
  845. 32:16computational overhead here or very
  846. 32:17minimal computational overhead but this
  847. 32:20comes at the cost of uh relaxed privacy
  848. 32:22notion since we're relying on a secure
  849. 32:24Hardware
  850. 32:26component
  851. 32:28you've probably heard of Federated
  852. 32:30learning this is popularized by Google
  853. 32:33several years ago initially designed for
  854. 32:36uh this use case of training machine
  855. 32:38learning models across millions of
  856. 32:40mobile devices the idea is fairly simple
  857. 32:42and that you know we have a model and
  858. 32:44then we compute these local model
  859. 32:46updates on each device aggregate them
  860. 32:49aggregate them across devices and apply
  861. 32:52the uh the global update to the the
  862. 32:54shared model and redistribute them and
  863. 32:57we're repeating this process to improve
  864. 32:58the accuracy of the overall
  865. 33:00model this has um you know immediate
  866. 33:04extension to the biomedical setting
  867. 33:06where we have a smaller group of
  868. 33:07collaborators we can also try to jointly
  869. 33:09train machine learning models this
  870. 33:12way the focus so far has been mostly on
  871. 33:15uh machine learning models or deep
  872. 33:16learning models that can be trained with
  873. 33:18gradient based optimization so the
  874. 33:19relevance for uh you know some of the
  875. 33:23genomic analysis tests may be limited
  876. 33:27because we're using different types of
  877. 33:28statistical models but it has found
  878. 33:31really useful application in in medical
  879. 33:33image analysis domain where we do need
  880. 33:36these deep learning models uh that are
  881. 33:38trained with
  882. 33:40gradients and this aggregation of local
  883. 33:42Updates this can reveal some private
  884. 33:45information from each party that's
  885. 33:47participating in this workflow so this
  886. 33:49needs to be carefully considered and
  887. 33:51typically these methods are combined
  888. 33:53with other techniques like encryption or
  889. 33:56differential privacy which is what I'll
  890. 33:57mention in the next slide to increase
  891. 33:59the security level of this
  892. 34:02solution so uh differential privacy it's
  893. 34:05addressing a slightly a different
  894. 34:06problem where we're concerned with the
  895. 34:08leakage of private information and the
  896. 34:10data that is released at the end of the
  897. 34:12computation so we can think about chwa
  898. 34:14statistics for
  899. 34:15example the key concept here is this
  900. 34:18notion of neighboring databases so these
  901. 34:20are two hypothetical uh databases that
  902. 34:23differ in exactly single individual and
  903. 34:26the idea is that we want to introduce
  904. 34:28some noise into the system or the into
  905. 34:30the computation so that only based on
  906. 34:33the release data we can't really
  907. 34:35distinguish between these two
  908. 34:36neighboring databases for any two
  909. 34:38neighboring
  910. 34:40databases and this privacy guarantee is
  911. 34:43typically controlled by this privacy
  912. 34:44parameter
  913. 34:46Epsilon and in this case values of
  914. 34:48Epsilon that are closer to zero means
  915. 34:50higher privacy and greater values
  916. 34:53um mean lower privacy but at the same
  917. 34:56time it allows us to roduce a smaller
  918. 34:58amount of noise so the results tend to
  919. 35:00be more accurate so this framework has
  920. 35:03been applied to Jo statistics release
  921. 35:05and other types of biomedical database
  922. 35:07queries like looking for corts that
  923. 35:09match certain
  924. 35:11criteria um but for many of these
  925. 35:15existing studies uh they were limited to
  926. 35:19uh specific simpler analysis settings uh
  927. 35:22for for example for the the case of
  928. 35:25chios we're instead of releasing entire
  929. 35:28genome wide Vector we can only release
  930. 35:30say like the top K most significant hits
  931. 35:32or something like that that's been the
  932. 35:34typical setting that it's been study and
  933. 35:37the the main reason is that the the
  934. 35:39actual noise that we need to introduce
  935. 35:41to achieve differential privacy
  936. 35:42guarantee tends to be very high uh in
  937. 35:45high dimensional
  938. 35:48cases so these techniques are useful in
  939. 35:52you know a range of different settings
  940. 35:53in biomedical research I've listed out
  941. 35:55some examples here that we've worked on
  942. 35:56over the years uh just to mention at a
  943. 35:58high level uh we can use these tools to
  944. 36:01build um uh softwares that can help
  945. 36:05researchers across institutions run
  946. 36:08collaborative studies in a secure way
  947. 36:10can also build analytic services that
  948. 36:12take users private data process them um
  949. 36:16and returns statistical insights back to
  950. 36:18them all in a privacy preserving way
  951. 36:21there's also this private data release
  952. 36:22setting similar to what I just mentioned
  953. 36:24with uh jwas where we want to release
  954. 36:27some privatized version of a bi medical
  955. 36:29data or analysis results with some
  956. 36:32meaningful guarantees of privacy for the
  957. 36:33underlying individuals in the data
  958. 36:35set and we're not the only ones working
  959. 36:38on this there's a growing community of
  960. 36:39researchers developing these tools I
  961. 36:41just wanted to mention a couple of
  962. 36:43examples of these uh real world
  963. 36:45competitions where we bring people uh
  964. 36:48together from different communities to
  965. 36:49create solutions for these tasks so idh
  966. 36:51is a popular one in this domain uh last
  967. 36:54year there was also a governmental
  968. 36:56collaboration between us and UK on
  969. 36:58developing tools based on these
  970. 37:00Technologies uh one of the tasks was
  971. 37:02creating a model for pandemic
  972. 37:05forecasting and uh some of us at bro
  973. 37:08entered these competitions and had
  974. 37:10winning solutions for for both of
  975. 37:13them so now I'll briefly go through some
  976. 37:16of our recent work in this domain as
  977. 37:19well uh so in our recent work we
  978. 37:22introduced this tool called secure
  979. 37:23Federated chws or sfj this is a
  980. 37:25cryptographic approach to performing Jos
  981. 37:28across multiple institutions without
  982. 37:30sharing any private information between
  983. 37:32them the key technical Insight that we
  984. 37:34introduce in this work is that as
  985. 37:36opposed to using a single technology um
  986. 37:39one of the ones that I mentioned earlier
  987. 37:40we can actually use a combination of
  988. 37:42them to gain computational speedups
  989. 37:44which tend to be quite important in
  990. 37:46practice so in this case we're combining
  991. 37:48NPC and homeric encryption which allowed
  992. 37:52us to design these Federated uh secure
  993. 37:55computation systems and what that means
  994. 37:58is that instead of encrypting the entire
  995. 38:00input data set and sharing it between uh
  996. 38:02the parties we can keep all the local
  997. 38:04input data sets local and only share
  998. 38:06encrypted versions of intermediate
  999. 38:08analysis results to carry out the global
  1000. 38:10uh computation this reduces
  1001. 38:13communication and also speeds up
  1002. 38:14computation at the site level because
  1003. 38:17now we're able to use unencrypted
  1004. 38:19data and we were able to show that this
  1005. 38:22um results in an order of magnitude
  1006. 38:25improvement over the prior art that was
  1007. 38:27only using MPC
  1008. 38:30technology across these different uh
  1009. 38:32data
  1010. 38:34bases and this scalability Improvement
  1011. 38:37allowed us to take these tools and then
  1012. 38:40apply to really large biank scale data
  1013. 38:42sets here I'm showing you two examples
  1014. 38:43emerg in UK byy bank so the latter in
  1015. 38:46our experiment included around 276,000
  1016. 38:49individuals we're splitting these
  1017. 38:51resources across uh six or seven
  1018. 38:53simulated Federated units to demonstrate
  1019. 38:57the ability to do um perform joint
  1020. 38:59analysis without sharing private data
  1021. 39:01and I'm showing you the results of chos
  1022. 39:03that matches accurately with centralized
  1023. 39:05analysis um where all the data is pulled
  1024. 39:08together into a single
  1025. 39:10location and the overall runtime for
  1026. 39:13these pipelines uh tend to be around a
  1027. 39:16few days for uh databases of this scale
  1028. 39:20so this is not a you know Che cheap
  1029. 39:23pipeline to run but uh the pract the run
  1030. 39:26times are still practically feasible and
  1031. 39:28these can be brought down further by
  1032. 39:30introducing additional Computing
  1033. 39:33resources and I want to point out that
  1034. 39:36the um the experiment that we've run
  1035. 39:39that we ran on UK bank is actually 2,000
  1036. 39:41times larger than the previous Benchmark
  1037. 39:43that we test tested in a previous
  1038. 39:46study and also want to highlight that uh
  1039. 39:50in this initial pipeline that we
  1040. 39:51demonstrated we're looking at PCA based
  1041. 39:54workflow where we run principal
  1042. 39:55component analysis to look for these
  1043. 39:57ancestry karious include as fixed effect
  1044. 40:00term in the gwas model we're but it's
  1045. 40:03also common to replace this with a
  1046. 40:04random effect model uh just known as
  1047. 40:07linear mix model approach to gas and
  1048. 40:11this is generally regarded as being more
  1049. 40:14accurate for jwas I I don't need to tell
  1050. 40:16you this but it also comes at a greater
  1051. 40:19computational cost even without any
  1052. 40:21encryption so we weren't able to do this
  1053. 40:23using secure computation techniques for
  1054. 40:25a long time but recently there was this
  1055. 40:27method that was introduced called
  1056. 40:29regini and it's doing an efficient
  1057. 40:32approximation of this random effect term
  1058. 40:34using a whole gen regression model and
  1059. 40:36what this allowed us to do is uh
  1060. 40:38distribute this computational workflow
  1061. 40:41across the parties using our
  1062. 40:42cryptographic tools and also obtain a
  1063. 40:44practical pipeline for uh running LM
  1064. 40:47based
  1065. 40:48Jos and um the distributed technique
  1066. 40:51that I alluded to turned out to be
  1067. 40:54fairly significant so we were able to uh
  1068. 40:57remove dependence on the size of the
  1069. 40:59cohort that goes into these stall
  1070. 41:01studies as you can see in these runtime
  1071. 41:04plots that stay nearly um constant
  1072. 41:07across different datas as sizes we're
  1073. 41:09able to show that the output of our lmm
  1074. 41:13analysis workflow closely matches with a
  1075. 41:15centralized execution of this regini
  1076. 41:18tool on the pool data set so now we're
  1077. 41:21able to run these different types of Jos
  1078. 41:23workflows in a secure and Federated
  1079. 41:26manner
  1080. 41:27we had this related work uh more
  1081. 41:30recently which was about finding genetic
  1082. 41:33relatives across different data sets
  1083. 41:35that cannot be pulled together due to
  1084. 41:37privacy concerns and this is a key
  1085. 41:40pre-processing pre-processing step in
  1086. 41:42many studies or Federated databases like
  1087. 41:45Nomad where we um where the presence of
  1088. 41:48these relatives could skew the
  1089. 41:50distribution of the statistics that we
  1090. 41:52want to compute and release to the
  1091. 41:54public so we have the cryptographic
  1092. 41:56tools to be able to comp compute this um
  1093. 41:58this relatedness coefficient privately
  1094. 42:01but the the main challenge here is the
  1095. 42:03fact that we need to perform all pawise
  1096. 42:06comparisons between these two data
  1097. 42:08sources and that turns out to be very
  1098. 42:10expensive um so the key idea that we had
  1099. 42:12in this work is that we
  1100. 42:15could hash and bucket the individuals so
  1101. 42:18that we can perform comparisons only
  1102. 42:21between individuals that belong to the
  1103. 42:23same
  1104. 42:23bucket and the idea is that due to
  1105. 42:26identity by this
  1106. 42:27uh sharing between close relatives uh
  1107. 42:32the close relatives have a higher
  1108. 42:34probability of getting assign to the
  1109. 42:35same bucket so this still retains high
  1110. 42:37accuracy we're able to demonstrate that
  1111. 42:40on UK by bank and all of us data you can
  1112. 42:42see that uh this overall
  1113. 42:45recall for detecting up to third degree
  1114. 42:48relatives uh is very high it's um 97% or
  1115. 42:53higher for in all three cases the
  1116. 42:55Precision is also High higher than 98 %
  1117. 42:58and some of these false positive
  1118. 43:00findings are um arguably not false
  1119. 43:03positives what they are is that they
  1120. 43:05have a relatedness coefficient that's
  1121. 43:07really close to threshold and due to
  1122. 43:08numerical Precision it's Crossing that
  1123. 43:12boundary and the overall runtime was
  1124. 43:15also practical less than a day for all
  1125. 43:16these data sets so where we head it next
  1126. 43:19I'll try to conclude the talk in the
  1127. 43:21next few minutes um so I think these
  1128. 43:24Technologies are maturing enough that
  1129. 43:27now we're able to see these uh inklings
  1130. 43:29of practical tools for different genomic
  1131. 43:32analysis tasks and I think over the next
  1132. 43:34few years we'll see a lot more of these
  1133. 43:36uh emerging for different analysis
  1134. 43:39workflows including things like fine
  1135. 43:41mapping or predicting disease risks
  1136. 43:43based on personal genome so we're
  1137. 43:45actively working on some of these ideas
  1138. 43:47and I think these would be valuable
  1139. 43:49tools for people doing genomic analysis
  1140. 43:51across different sites uh to help with
  1141. 43:55collaboration we are currently working
  1142. 43:57on a deployment study where we're trying
  1143. 43:59to run some of these joint analyses
  1144. 44:01between million V million veteran
  1145. 44:04program and all of us research program
  1146. 44:05so as you know these are two of the
  1147. 44:07largest fire banks in the US they're
  1148. 44:09also known for um their strict security
  1149. 44:12guidelines that prevent the data from
  1150. 44:14being shared externally so we're trying
  1151. 44:16to use our tools to create a link
  1152. 44:18between them to um enable joint
  1153. 44:21studies last year we launched our SF kit
  1154. 44:24web server this is Joint work with grow
  1155. 44:27data Sciences platform and this is a
  1156. 44:30webbased service that allows people to
  1157. 44:31run secure Federated tools the
  1158. 44:34cryptography based ones that I mentioned
  1159. 44:36uh in a push button kind of way uh and
  1160. 44:40they can bring their own data set run a
  1161. 44:42collaborative analysis with their
  1162. 44:43collaborators uh using the service all
  1163. 44:45the components that we have here are
  1164. 44:47open source and this is actually the
  1165. 44:49service that we're using for the ongoing
  1166. 44:50pilot
  1167. 44:51study and currently we're also working
  1168. 44:54on integrating this uh set of tools into
  1169. 44:58existing cloud-based analysis platforms
  1170. 45:00like Tera which has also developed at
  1171. 45:02Road as well as all of us researcher
  1172. 45:04workbench uh which is based on Tera so
  1173. 45:07these are some of the things that are
  1174. 45:08coming up down the road and lastly I'll
  1175. 45:10just mention that there's an upcoming
  1176. 45:11conference that would be of interest to
  1177. 45:13all of you so recom is happening it's
  1178. 45:15being organized at MIT how
  1179. 45:18convenient and there's actually this
  1180. 45:20inaugural satellite conference on
  1181. 45:22biomedical data privacy and equity which
  1182. 45:24is happening one day before the main
  1183. 45:26conference and we have several talks in
  1184. 45:28both of these venues if you're
  1185. 45:30interested please check out the programs
  1186. 45:32and you can also reach out to me if you
  1187. 45:33want to discuss more and with that I
  1188. 45:35want to thank all of the uh former and
  1189. 45:38Cur current labp members of my group and
  1190. 45:41my collaborators and mentors uh in
  1191. 45:44different institutions for all their
  1192. 45:46input and support without them we
  1193. 45:48couldn't have done you know a lot of
  1194. 45:50these
  1195. 45:50works thank you for your attention I'll
  1196. 45:53take any questions great thank you so
  1197. 45:55much Dr Cho I think we have time for one
  1198. 45:58or two questions I might start them off
  1199. 46:02um so if we think about how like each of
  1200. 46:05us we're all required at publication to
  1201. 46:07deposit all of our transcript omix right
  1202. 46:09in a public data frame and I try to
  1203. 46:12think of how we implement this does it
  1204. 46:15mean that like the gene on our gene
  1205. 46:17expression
  1206. 46:18Omnibus data like should it be at the
  1207. 46:22repository level that these things are
  1208. 46:23implemented so researchers could still
  1209. 46:25upload their kind of
  1210. 46:27naive data frames as we do currently and
  1211. 46:30then the the implementation of privacy
  1212. 46:33happens at the repository or do you
  1213. 46:35think this is actually something that
  1214. 46:36where even that step is not secure
  1215. 46:38enough and it should be uh the Privacy
  1216. 46:40should be implemented prior to upload so
  1217. 46:43I I can imagine a repository like Goo uh
  1218. 46:48being run like something like DB Gap
  1219. 46:50with the genotype data with stricter
  1220. 46:52access uh control mechanisms I'm not
  1221. 46:56necessarily proposing that needs to
  1222. 46:58happen but we do have some concerns
  1223. 47:00about the way currently transic data is
  1224. 47:03being shared because they do Le
  1225. 47:04genotypic information and it's not just
  1226. 47:06our concern there are other people um
  1227. 47:09who's been pushing on this as well and I
  1228. 47:11know that there's an ongoing
  1229. 47:12conversation with NIH on this topic so
  1230. 47:15uh yeah there might be some changes in
  1231. 47:17the future but we're not sure yet I I do
  1232. 47:20want to make a note that the the risks
  1233. 47:23that I showed about that data is kind of
  1234. 47:25in a controlled environment right where
  1235. 47:28we have this matching ancestry between
  1236. 47:30like the the target individuals and like
  1237. 47:32the other individuals in the candidate
  1238. 47:34pool so there are like other
  1239. 47:35considerations like that that needs to
  1240. 47:37go in when we actually think about the
  1241. 47:39the risk and the Practical
  1242. 47:42setting Mak
  1243. 47:44sense great well thank you so much for a
  1244. 47:48broad sweeping talk that really
  1245. 47:50highlighted I think so many different
  1246. 47:51ways in which privacy impacts the
  1247. 47:53research that we all do every day um and
  1248. 47:55then please feel free to go grab second
  1249. 47:57breakfast and then we'll start uh the
  1250. 47:59next session at
  1251. 48:079:30

About this transcript

This page contains the full transcript of MPG Primer: Genomic Privacy: Key Issues and Emerging Solutions (2024) by Broad Institute, generated from the public captions YouTube serves with the video. The transcript has 8,064 words across 1,251 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.