MPG Primer: Genomic Privacy: Key Issues and Emerging Solutions (2024) — Transcript
Full transcript
- 0:00great thank you all for being here today
- 0:02for our MPG primer how could you
- 0:06forget we are very excited to have Dr
- 0:09hun Cho here who's an assistant
- 0:10professor of biomedical information and
- 0:12data science at Yale University he was
- 0:15previously here at bro as a Schmid
- 0:17fellow and uh principal investigator he
- 0:19received his PhD in electrical
- 0:21engineering and computer science at MIT
- 0:23in 2009 um and before that had been at
- 0:25Stanford where he got both a bachelor's
- 0:27and master's degree uh in computer
- 0:29science with honors he's been named a
- 0:31recipient of the NIH director's early
- 0:33Independence award
- 0:35congratulations um and uh the Cho lab
- 0:37now works to create algorithmic
- 0:39solutions to tackle a variety of
- 0:40computational challenges uh introduced
- 0:42by the presence now of these large scale
- 0:45heterogeneous um data sets that have uh
- 0:48important privacy concerns and so Dr Cho
- 0:51is very happy to take your questions
- 0:53during the talk today um and so please
- 0:55feel free to come up to the microphone
- 0:57uh or type in the Q&A on Zoom um and
- 1:00thank you again for joining us all the
- 1:01way from you yeah thank you Sarah thanks
- 1:03everyone for making here in person
- 1:05despite the rain I appreciate people on
- 1:06Zoom as well um as Sarah mentioned I was
- 1:10a fellow here for four years before I
- 1:12moved to Yale last September um so I'm
- 1:14very familiar with this space it's
- 1:15always good to be
- 1:17back so yeah in this primary I'll be
- 1:19talking about genomic privacy and um
- 1:22given my background in computer science
- 1:24I'll be mostly sharing my perspectives
- 1:26on algorithmic and computational side of
- 1:28things so just wanted to mention that up
- 1:29front
- 1:30uh in the first part of this talk I'll
- 1:33be giving you an overview of privacy
- 1:34risks that are known to be associated
- 1:36with genomic data and then in the second
- 1:38part I'll talk about some of these uh
- 1:40emerging Solutions based on what are
- 1:42called privacy enhancing Technologies
- 1:44for addressing these risks and
- 1:46throughout the talk I'll try to both uh
- 1:48you know Define the the key Concepts in
- 1:50these areas and also give you some
- 1:53vignettes of our recent work on these
- 1:55areas as well so I want to first start
- 1:58by talking about what we mean when we
- 2:00say privacy and distinguish it from a
- 2:03another closely related term security so
- 2:06privacy in general is a right to be let
- 2:08alone that's a common way to define it
- 2:11and it also means that an individual has
- 2:12control over how information about them
- 2:16uh is collected shared or used in
- 2:19different settings on the other hand
- 2:21security is generally about a property
- 2:23of a system that provides protection
- 2:25against unauthorized attempts at
- 2:27accessing the data or tampering with it
- 2:30and also tampering with the uh the
- 2:31functionality of these systems so these
- 2:33are closely related as you can tell um
- 2:37but one does not necessarily imply the
- 2:39other and we're generally interested in
- 2:41developing systems that are both secure
- 2:43and provide protection uh for the
- 2:44privacy of the underlying
- 2:47individuals and as you know privacy is
- 2:49an important topic in biom medicine and
- 2:51genomics uh in general and this is
- 2:54mostly due to the fact that biomedical
- 2:57information it contains something that
- 2:59very that's very
- 3:00personal to us uh about our biology and
- 3:04health and the thing that makes this
- 3:06problem so complex is that there are all
- 3:08these different types of biomedical data
- 3:10types that uh reveal different things
- 3:13about our biology so including things
- 3:15like electronic health records medical
- 3:17images personal genomes and more
- 3:19recently data from wearable devices
- 3:21there's also these variety of omix data
- 3:24types transcriptomics proteomics and
- 3:26also metagenomic data from microbiomes
- 3:29so all these different types tell us
- 3:31different things about um human health
- 3:33but at the same time they increase the
- 3:34stes uh the stakes for keeping all this
- 3:37information safe and protect it from
- 3:39potential
- 3:41misuse and this is not just hypothetical
- 3:44concerns there um many real world
- 3:46examples of privacy threats you may have
- 3:48seen in the news uh a recent privacy
- 3:50breach data breach at 23 andme last year
- 3:53which affected millions of users the
- 3:55attackers were able to access raw
- 3:57genotype data as well as health related
- 3:59information like disease predisposition
- 4:01status as well as genetic relative
- 4:03information for these users is quite
- 4:06concerning and it's also even more
- 4:08concerning that the attackers were able
- 4:10to create these curated lists targeting
- 4:12specific groups of individuals and they
- 4:15were s selling these data on the dark
- 4:18web here's another more recent example
- 4:20from just a few weeks ago I think the
- 4:22situation is still unfolding there was a
- 4:24large uh ransomware attack at change
- 4:27Healthcare which is a prescription
- 4:29processing arm of this insurance company
- 4:32called United
- 4:34Health it was reported that around 94%
- 4:38of hospitals in the United States were
- 4:39affected financially due to this attack
- 4:42and there's still an ongoing
- 4:43investigation into what kinds of data
- 4:46had been accessed uh through this
- 4:49attack so there are real threats and I
- 4:52want to make the point that this is not
- 4:53just a concern in uh the in the
- 4:55commercial
- 4:56domain those of us uh who are involved
- 4:59in academic research and genomics
- 5:00genomics should also care about privacy
- 5:02too and here are some reasons so behind
- 5:04every data set there are real people who
- 5:06could be affected by a potential privacy
- 5:08breach so we have a certain amount of
- 5:11responsibility to uh look for safer ways
- 5:13to conduct research and second these
- 5:16privacy concerns and related regulations
- 5:18they make it difficult to for
- 5:21researchers across different
- 5:23institutions to collaborate and share
- 5:24data so having better ways to protect
- 5:26privacy can help us uh break down these
- 5:29barriers
- 5:30and acceler
- 5:32Science and lastly if we think about
- 5:35converting biomedical insights from from
- 5:37uh Research into public health impact
- 5:39for example for supporting health
- 5:41related decision- making for private
- 5:42individuals we need reliable and
- 5:45trustworthy systems that can process
- 5:47private information so having better
- 5:49safeguards can help us in that regard as
- 5:52well so now let's talk about what's
- 5:55known about genomic privacy in more
- 5:57detail um so there are Key Properties of
- 5:59genetic data uh that make it unique and
- 6:03different from other types of personal
- 6:05information like passwords or online
- 6:07search history uh and what have you
- 6:10these are six of these properties that
- 6:12were summarized by uh this review
- 6:14article by naita at
- 6:16all I'll briefly go through them so uh
- 6:19genetic data it contains information
- 6:21about our health and behavior as I
- 6:22mentioned this is static kind of
- 6:24information it's mostly determined at
- 6:26Birth and it's hard to change it's also
- 6:28very unique to each IND idual so only
- 6:30tens of snips let's say would be enough
- 6:32to single out an individual from the
- 6:34global human population Mystique is
- 6:36referring to the fact that there's still
- 6:38a lot that we don't know about what can
- 6:39be inferred from genetic information
- 6:41these data have intrinsic value because
- 6:44they can be used for biomedical research
- 6:45or clinical purposes and lastly kinship
- 6:48means that when a privacy breach happens
- 6:50it reveals information not just about
- 6:52the affected individuals but also about
- 6:54their biological relatives so when we're
- 6:56talking about risks and safeguards it's
- 6:58important that we consider these uh
- 7:00properties of the
- 7:01data and there's been a lot of
- 7:03literature on different ways of
- 7:05breaching genomic privacy and this is a
- 7:08categorization that was provided in this
- 7:09other review article in nature genetics
- 7:11from a few years ago and they broadly
- 7:14classify the the attacks into two
- 7:16categories identification and phenotype
- 7:18inference uh I think this is fairly
- 7:20self-explanatory but just to Define it
- 7:22uh for clarity identification means that
- 7:24you start with a supposedly Anonymous
- 7:26data instance and we're uh uncovering
- 7:29the real IND idual behind that data
- 7:30instance pype inference means we start
- 7:33with some piece of data about an IND
- 7:35individual and we infer additional
- 7:37attributes that are not included in the
- 7:38data set so including their health
- 7:39status disease predisposition things
- 7:41like that so in the next few slides I'll
- 7:44go over some of the seminal works that
- 7:46explore different types of privacy
- 7:48attacks in this
- 7:49domain one of the earliest examples was
- 7:52from 2018 uh when next Generation
- 7:55sequencing first came about this attack
- 7:57demonstrated the difficulties of asking
- 7:59phenotypic information and genetic data
- 8:02so uh the first personal genome that was
- 8:05released was uh from Dr James Watson's
- 8:08um sample and in this release data set
- 8:12variance in the apoe gene which is known
- 8:14to be linked to alzheimer's disease were
- 8:16masked due to personal requests from Dr
- 8:19Watson but just a few months later
- 8:21another group of researchers show that
- 8:23these were actually not really Mass uh
- 8:24you could infer these variants using
- 8:27nearby variants using linkage dis
- 8:29equilibrium patterns in other words we
- 8:31could impute these
- 8:32genotypes so this led to a larger window
- 8:35around this Gene of two uh meab base
- 8:38pairs uh to be excluded from the data
- 8:41set and note that this is just for
- 8:43hiding one Gene and so it's going to be
- 8:46really difficult if you want to hide
- 8:47information about much larger array of
- 8:50disease Rel genes that we know uh that
- 8:52are scattered across the
- 8:55genome and from the same year there was
- 8:57also this seminal work by Homer at
- 9:00where they showed that even aggregate
- 9:02level information like minor Al
- 9:04frequencies and gwos summary statistics
- 9:07can reveal private information as well
- 9:09and what they were able to show is that
- 9:10given a Target subject's genome you
- 9:12could compare it against these released
- 9:14statistics to figure out whether the
- 9:16person was in the underlying cohort or
- 9:19not and this is also called membership
- 9:21inference attack and while this
- 9:23information about membership might not
- 9:25be too concerning by itself uh sometimes
- 9:27you study course or associate with
- 9:29sensitive attributes like rare medical
- 9:32conditions so that could lead to a
- 9:33potential privacy breach and uh which is
- 9:36a major concern
- 9:38here and uh so the intuition here behind
- 9:41this attack is that if if the subject
- 9:44has some rare variants that are uh not
- 9:48really found in the general population
- 9:49at high frequency then if these variants
- 9:51are represented highly in the released
- 9:54data then it makes it more likely that
- 9:55the subject wasn't the in the original
- 9:57data set that's the main idea
- 10:00in a more recent work people have shown
- 10:03that it's also possible to demonstrate
- 10:05this type of leakage uh with genomic
- 10:07Beacon Services where uh we answer
- 10:09questions like uh is is there a person
- 10:12with this particular mutation in the
- 10:14underlying data set and you can think of
- 10:16this as a different type of aggregate
- 10:19level information that's focusing on
- 10:21individual mutations and the idea here
- 10:24is the Same by combining information
- 10:26across different queries we can try to
- 10:27single out a uh an
- 10:31individual another type of attack that's
- 10:34been um receiving a lot of attention in
- 10:37the media uh a few years ago is
- 10:40identification attack based on public
- 10:42genealogy databases uh like GED match so
- 10:45these Services include genetic profiles
- 10:47from millions of users and their
- 10:50genealogical
- 10:51relationships and this is a great
- 10:53resource for recreational purposes for
- 10:55finding one's Roots but they can also be
- 10:56used for other purposes uh like police
- 10:58officers to catch um criminals based on
- 11:01the sample that they get from the crime
- 11:03scenes and the the main idea here is
- 11:05that you again you start with a genetic
- 11:07profile from a Target subject and then
- 11:09we create against these databases to
- 11:11find out their uh close or distant
- 11:13relatives in these databases and we know
- 11:15who they are so now we can try to
- 11:17triangulate and figure out who the
- 11:18person was uh sometimes additionally
- 11:20using information like their um their
- 11:23demographic information like age or
- 11:26rough geographic area where the person
- 11:28is
- 11:29and there's been several papers that
- 11:32explore different aspects of this uh
- 11:34threat Factor so this first paper by
- 11:36jimar at all um I think Melissa came
- 11:39here to give it talk a few months ago
- 11:42last year and uh so this work showed
- 11:45that just using short tandem repeats on
- 11:47chromosome y um You can predict the
- 11:50surname of the target individual later
- 11:54early Cal show that uh the majority of
- 11:57individuals of European descent are
- 11:59susceptible to an identity infer attack
- 12:02like the one that I explained and this
- 12:03is indeed how um the Golden State killer
- 12:05in San Francisco was caught if you're
- 12:07familiar with that
- 12:08story and more recently NL showed that
- 12:12uh it's actually not just the the target
- 12:14subject whose ident identity could be
- 12:16revealed the individuals in the database
- 12:19itself could also have their genomes
- 12:21leaked U through what's being revealed
- 12:22by these services and uh they also
- 12:25demonstrated that one can insert these
- 12:28uh artificial samples into these
- 12:30databases to Mis direct uh police
- 12:32investigators who trying to find the the
- 12:35the
- 12:37perer so going Beyond genotype profiles
- 12:41there's also leakage in other functional
- 12:43genomic data that we need to think about
- 12:45uh so I'm referring to gene expression
- 12:47data protein expression things like that
- 12:50the idea here is that even though these
- 12:52other types of molecular measurements
- 12:54tend to be a lot noisier they contain
- 12:56enough identifying information so that
- 12:58we can link uh samples from the same
- 13:01individual across different data sources
- 13:04and this is called a linkage
- 13:05attack and going beyond that we can also
- 13:08try to predict genotypes from these data
- 13:11through what are called quantitative
- 13:12trait low size or
- 13:14qtls these are genetic variants that are
- 13:16correlated with other molecular
- 13:19measurements and by predicting genotypes
- 13:22of the person based on their omix data
- 13:23we can also create a linkage between
- 13:25their genetic profile data in one data
- 13:27set against another omix data in a
- 13:30different data
- 13:31set and as I mentioned uh these
- 13:34possibilities have been demonstrated for
- 13:36different types of uh data so including
- 13:38transatomic proteomics and microbiome
- 13:41data and next I want to describe a
- 13:43little bit about our recent work from
- 13:44last year in genome research where we
- 13:46try to uh study this question of um this
- 13:51topic of transatomic data privacy a
- 13:52little
- 13:54better so again we're interested in
- 13:56looking at these quantitative tra low
- 13:58size uh but for gene expression so these
- 14:00are called
- 14:01eqtls and the concern is that a an
- 14:05attacker a potential attacker can start
- 14:07with the gene expression level that
- 14:09they've obtained and try to guess the
- 14:11genotypes of that person and there are
- 14:13many of these eils in the genome uh so I
- 14:16think it's known that the majority of
- 14:17human genes have some utl associated
- 14:20with them so you can imagine extracting
- 14:23a substantial portion of one's uh DNA
- 14:27through uh this approach
- 14:30the main issue is that the existing
- 14:32Works um on this
- 14:34idea were focusing on demonstrating this
- 14:37possibility and not really fully
- 14:39characterizing the extent of leakage so
- 14:42there were limitations in their
- 14:43statistical models that allow them to
- 14:45only analyze a smaller subset of eqls
- 14:48that are statistically independent so
- 14:50this is the problem that we try to
- 14:51address and we introduce this model
- 14:53that's working at the sequence level uh
- 14:55so we're able to look at all the DLS
- 14:58accounting for their corations as well
- 15:00for those of you who are familiar with
- 15:01this we're starting with this Le and
- 15:02Stevens hden Markoff model which um
- 15:05basically puts a distribution over an
- 15:07entire genetic sequence accounting for
- 15:09Link Link this equilibrium and we're
- 15:12adding another layer on top of this
- 15:14model or on the bottom here uh that's
- 15:17corresponding to gene expression
- 15:18variables and we create these links that
- 15:20capture statistical correlation between
- 15:22genotypes and gene expression and note
- 15:24that this is again including all theuts
- 15:26that are known and we train this model
- 15:29in an end to-end fashion so that the
- 15:31probabilities are calibrated to account
- 15:33for uh correlation between nearby
- 15:37genotypes and the idea is that this
- 15:39using this model we can produce a match
- 15:41score for a given pair of gene
- 15:43expression and genotype profile and this
- 15:46could be used to try to link individuals
- 15:48across data sets uh which can
- 15:50potentially lead to lead to
- 15:52reidentification
- 15:54so in terms of evaluating this model we
- 15:57took a data set we had pairs of these
- 15:59samples gen expression genotype samples
- 16:02uh from a set of individuals around 300
- 16:05individuals here in this case we
- 16:06included additional genotype profiles
- 16:08from a larger uh um data set in this
- 16:12case it was hype reference Consortium
- 16:14data and this was a much larger set so
- 16:1622,000 individuals in the um in the
- 16:19additional cohort we're trying to figure
- 16:21out for a given gen expression profile
- 16:24uh which genotype profile it corresponds
- 16:26to and what I'm showing you here is the
- 16:28fraction of these individuals who were
- 16:30able to correctly link between these two
- 16:32samples as we add more of the additional
- 16:35candidate um genotype profiles to the uh
- 16:39to the data
- 16:40set and I'm comparing our method DSM to
- 16:44two other Baseline methods ebl and gmbb
- 16:46these are the the previous approaches
- 16:48that I briefly mentioned that uses an
- 16:50independent set of
- 16:53etls the main thing to note here is that
- 16:55even after including this large number
- 16:58of individuals so uh 22,000 individuals
- 17:01we're still able to link around 90% of
- 17:03more of the original individuals based
- 17:05on gen expression information
- 17:08only and uh our method is able to link a
- 17:12greater fraction compared to the
- 17:13Baseline and this Gap becomes more
- 17:16pronounced if we consider a a setting
- 17:20that's using a smaller set of eils as
- 17:22well so just uh based on Chism 20 um the
- 17:26overall numbers go down a lot but the
- 17:28gap our method and the previous
- 17:29approaches uh becomes
- 17:33clearer the other thing that we can do
- 17:35here is that we can measure the
- 17:37confidence of our prediction of the
- 17:39links and this is a metric that we
- 17:42developed uh it's basically a P value in
- 17:45some sense that measures the gap between
- 17:46the best match and the second best
- 17:48match and if we compare this between our
- 17:51method and the previous methods you can
- 17:52also see here that uh our method is on
- 17:55the y- axis the previous ones are on the
- 17:57xaxis the the the confidence that we
- 17:59have over the links that we produce tend
- 18:01to be
- 18:03greater and what this leads to is that
- 18:06if we're working in a setting where we
- 18:08don't know if the matching individual is
- 18:10in the candidate pool or not we have to
- 18:12draw a threshold somewhere right and
- 18:14have some uh number of false positives
- 18:17and uh because of this higher confidence
- 18:20uh we also get predictions that are um
- 18:24that minimizes the number of mistakes
- 18:27compared to the the through mattress
- 18:29that we're able to find as a trade-off
- 18:30fing Precision
- 18:33recall and now I want to describe
- 18:35another recent work which is
- 18:37illustrating a different type of um uh
- 18:40privacy risk and I'm bringing this up
- 18:42because I think this is a th Vector that
- 18:44not a lot of us uh you know have been
- 18:48aware of so it's kind of a neat idea I
- 18:51think that we should all be aware of and
- 18:54so here we're looking at the security of
- 18:56genotype imputation servers I think many
- 18:59of you know about this and so this is
- 19:01referring to services like topmed and
- 19:03Michigan imputation servers what they
- 19:05generally do is we provide a uh a
- 19:08partially observed genotype profile from
- 19:11array uh microarray genotyping platforms
- 19:14for example and then we obtain an
- 19:17imputed genotypes that fills in all
- 19:18these missing positions and the way
- 19:21we're able to do this is by using a
- 19:23large reference panel of high quality
- 19:26genomes and using linkage dis
- 19:28equilibrium pattern uh between nearby
- 19:30Snips and the question that we asked
- 19:33was this process is using some
- 19:36information from the reference panel to
- 19:37be able to fill in the missing data but
- 19:39how much information is actually being
- 19:41revealed uh throughout this process and
- 19:43included in the output and in fact can
- 19:45we actually try to reconstruct some of
- 19:46these sequences in the reference panel
- 19:48by interacting with these servers in
- 19:50certain
- 19:52ways and just to avoid any doubt we were
- 19:55we took great care to do this work um
- 19:58respon II L and ethically so none of
- 20:00this work was performed on these servers
- 20:01it was all done in a local simulation
- 20:03environment we also informed the
- 20:05stakeholders well in advance actually
- 20:07almost a year in advance and uh we're
- 20:10aware that additional security measures
- 20:12have been introduced after we published
- 20:14our work and all the information that
- 20:15I'm sharing today is public and it's
- 20:17published last
- 20:19year okay so the the key idea that led
- 20:22to this work is actually the fact that
- 20:24if you plug in an input sequence that
- 20:27uniquely matches with a single reference
- 20:28sequence it will extract the the
- 20:31matching person's genome um almost
- 20:34exactly so this is this uh this figure
- 20:38is showing that idea so uh in this
- 20:40bottom left corner you're seeing an
- 20:42example output from the imputation tool
- 20:45um we're showing it as predict the
- 20:47dosages which is a value between zero
- 20:49and one and you can see this value
- 20:51across the the genomic window that's
- 20:52being imputed on the
- 20:54xaxis if there's a single match the
- 20:56predictions are very confident it's near
- 20:59zero near one and this corresponds to
- 21:01the actual genotypes of the matching
- 21:03person what's even more surprising is
- 21:05that if you have a match um have a query
- 21:08that matches a small number of reference
- 21:10samples you can actually see this unique
- 21:14pattern that corresponds to that too so
- 21:15if it matches three samples then you'll
- 21:17see these predicted dosages that are
- 21:19aggregated around multiples of 1/3s so
- 21:22what this means is that you can look at
- 21:23the imputation output and figure out how
- 21:24many mattress there were in the
- 21:26reference panel and what does these to
- 21:28is this type of um potential attack
- 21:31pipeline where we start with a lot of
- 21:33random queries impute them and see what
- 21:36we get in terms of these dosage outputs
- 21:39and figure out which one of them uh was
- 21:41unique which ones had multiple matches
- 21:43and for those with multiple matches we
- 21:45can try to extend the query to make it
- 21:46more unique and in the end we pull
- 21:49together all of the unique matches and
- 21:50this corresponds to a set of um path
- 21:53type sequences from the reference
- 21:57panel and using 1,000 genomes data as
- 22:01the reference panel uh as an example we
- 22:03were able to demonstrate that this type
- 22:04of attack can actually extract a fair
- 22:06amount of sequences from the reference
- 22:08panel so for example if we impute 1
- 22:12million um input queries that we
- 22:14constructed in some way we're able to
- 22:17extract almost 80% of the uh the
- 22:20sequences in the reference panel there's
- 22:22some amount of error that's happening
- 22:24here because of the uh Mis
- 22:26classifications um in the pipeline but
- 22:29this error rate tends to be fairly
- 22:32low and here's a slightly different
- 22:35setting where you know to try to prevent
- 22:38this attack what we could do is instead
- 22:40of releasing the predicted probabilities
- 22:44we can uh release just the discrete
- 22:46genotype predictions so either zero or
- 22:49one nothing in between but even in the
- 22:51setting there is a workr that we that
- 22:53allows us to extract unique matches so
- 22:56the general idea is that you can take
- 22:59the output um sequence introduce
- 23:02additional
- 23:04changes or or keep a subset of positions
- 23:07uh mask again some of the other
- 23:09positions imputed again if the results
- 23:12stays the same then it's more likely
- 23:14that the original sequence that was
- 23:15output was corresponding to a unique
- 23:17individual so using this type of idea we
- 23:19can still extract sequences um at
- 23:21slightly lower Effectiveness as you can
- 23:25see one big caveat here is that
- 23:28sequences that we're getting here are
- 23:30chunked from the private uh genomes it's
- 23:33not the entire genome and this is
- 23:35because imputation pipeline is run
- 23:36independently on uh smaller genomic
- 23:39Windows typically around 20 megabase
- 23:42pairs but it turns out that you can
- 23:44start from these fragments of the genome
- 23:47and Link them across different Windows
- 23:49using relatedness patterns and uh this
- 23:51is just a visualization I won't go into
- 23:53details here that if you compute the
- 23:55kinship coefficient between these
- 23:57individual pieces with um relative of a
- 24:00certain degree you can see that the
- 24:02kinship is uh separated from is far
- 24:06enough from zero that we can distinguish
- 24:08it from unrelated
- 24:10samples and what this means is that we
- 24:12can imagine a setting where we have
- 24:15access to a secondary data set which
- 24:17might potentially include relatives of
- 24:20the individuals who are in the reference
- 24:22panel and then we could compute these
- 24:25kinship coefficients uh it's written as
- 24:27semi kinship because it's a modified
- 24:28version that's working with half types
- 24:30instead of typo types so we can compute
- 24:33these K coefficients between Pairs of
- 24:35samples between t between these two
- 24:38sources and then we feed this into our
- 24:40prisic linking algorithm that's taking
- 24:42this Matrix as input and then outputs uh
- 24:44groups of sequences that are likely to
- 24:46have come from the same individual so
- 24:48that's the overall outline of the attack
- 24:53factor and uh we were able to show that
- 24:57um given a data set with um real
- 24:59relatives that are known to be included
- 25:01in the relatives that we were able to
- 25:03link a substantial portion of the parget
- 25:06genome um up to third degree relative so
- 25:10beyond that the signal becomes too weak
- 25:12so we can't like link too many uh
- 25:16samples and based on this data we can
- 25:19extrapolate to get a sense of what
- 25:21fraction of the genome could be linked
- 25:23for different proportions of the
- 25:24reference panel individuals given
- 25:27different sizes of this uh relative set
- 25:29that I was referring to um relative to
- 25:32the overall size of the population that
- 25:33the reference panel came from so for
- 25:35example if we're given a relative set
- 25:37that includes 0.5% of the underlying
- 25:40population uh we
- 25:42could expect to link 17% of the genome
- 25:46for at least 5% of the individuals in
- 25:48the reference panel so that's the one of
- 25:50the estimates that we were able to get
- 25:52and this is a realistic number given
- 25:54that 05% is comparable to size of the UK
- 25:57biank for example
- 26:00so just to summarize this portion
- 26:03um I've shown you different types of
- 26:05genomic pry risks we looked at
- 26:08reidentification attacks membership
- 26:09phenotype inference data linkage as well
- 26:12as data
- 26:13reconstruction and the overall takeaway
- 26:16is that our understanding is still
- 26:19rapidly shifting and this needs to
- 26:21continue to change as we gain access to
- 26:23different types of data different models
- 26:25uh more advanced Ai and ml tools as well
- 26:28as well as other types of data sharing
- 26:30systems that are emerging and given all
- 26:33these changes it's important to have
- 26:34these rigorous approaches to analyze
- 26:36these risks and assess them uh using new
- 26:40models and come up with new ways to
- 26:42address these
- 26:43risks so on that point in the second
- 26:46part I'll give you some ideas of some of
- 26:48the emerging Technical Solutions for
- 26:49addressing use
- 26:54risks I think you're all familiar with
- 26:56these existing policy Frameworks like
- 26:58common role privacy uh Hippa privacy
- 27:01role gdpr they Define these Notions that
- 27:04we all know and love like IRB review
- 27:06processes informed consent the the idea
- 27:09of protected health
- 27:12information and uh there's no question
- 27:15that these Frameworks provide an
- 27:16important uh tool for us to safeguard
- 27:19the use of biomedical data but it's also
- 27:21important to note some of the inherent
- 27:23limitations to these regulatory
- 27:25approaches so the first thing is that
- 27:27designing and updating these uh
- 27:29regulations is generally a slow and
- 27:31long-term process for a good reason um
- 27:33one example is that genetic data was
- 27:35classified as protected health
- 27:37information fora only in 2013 several
- 27:41years after genetic data have become
- 27:43more widely available there are also
- 27:45these key ambiguities in some of the
- 27:47terms uh that we use in these
- 27:49regulations right so when we say
- 27:50deidentified or anonymized data what
- 27:53does that actually mean and how does it
- 27:54apply to the data sets that we're
- 27:56actually dealing with and
- 27:58it you know as I've demonstrated
- 28:00hopefully a lot of the bical data types
- 28:03it's not possible to completely de
- 28:05identify it there's always going to be
- 28:06some amount of ident identifying
- 28:08information sorry uh so we need to draw
- 28:11a line somewhere when we think about um
- 28:13laws and regulations and perhaps the
- 28:16most important uh Pitfall is that these
- 28:19approaches don't resolve the core
- 28:21conflict between the need to share data
- 28:23and the need to uh protect privacy and
- 28:25this is because uh they're typically
- 28:27about limiting the scope scope of um
- 28:29data sharing and usage depending on the
- 28:32context uh in order to provide PR
- 28:36privacy so it turns out that uh there's
- 28:39a collection of techniques from the
- 28:41computer science literature which are
- 28:43referred to as privacy enhancing
- 28:44technologies that can help us resolve
- 28:46this conflict and at a high level these
- 28:49provide us with mathematical techniques
- 28:52uh that can help us use and share
- 28:54private biomedical data while at the
- 28:56same time protecting privacy so this is
- 28:58what I'll tell you about briefly
- 29:01next um I'm going to highlight five of
- 29:04these technologies that belong to this
- 29:05category I'm going to start with homor
- 29:07for encryption which is a type of
- 29:09encryption
- 29:11technique and it's a special form where
- 29:14the uh the structure of the encryption
- 29:16allows us to perform operations directly
- 29:19on the encrypted data without having to
- 29:21decrypted
- 29:23First initially the uh these Frameworks
- 29:28were limited to performing uh limited
- 29:30types or limited numbers of operations
- 29:32on the private data but there's been
- 29:34remarkable progress over the years and
- 29:35now we have these practical systems that
- 29:37can help us Implement uh fairly complex
- 29:40analytic
- 29:41tests but they do come with a a
- 29:43computational overhead because now we're
- 29:45operating with encrypted data and that's
- 29:47a lot more expensive than working with
- 29:49non-encrypted
- 29:51data a slightly different approach as
- 29:53SEC multiparty computation it has two
- 29:55acronyms NPC or SMC depending on who you
- 29:58ask and the core of this framework is
- 30:03not encryption but we're going to use
- 30:05this technique called secret sharing and
- 30:06what it means is for each private number
- 30:08we um divide it into a set of random
- 30:11numbers that somehow all add up to the
- 30:14private number and we're going to split
- 30:16them up and then distribute it to
- 30:18multiple parties so they all together
- 30:20know uh well the shares collectively
- 30:23encode information about the private
- 30:25number but if you look at each party's
- 30:28information they don't know what the
- 30:29private number is and the idea is that
- 30:31we can uh design these interactive
- 30:33protocols where these parties work
- 30:35together to perform computation on the
- 30:37underlying secret without revealing
- 30:39anything throughout this process so at
- 30:42the end of a protocol like this we get
- 30:43the the um output of the analysis
- 30:45results as secret shares as well which
- 30:47we can combine to reveal the
- 30:50output so this is not using encryption
- 30:52so the computational operations tend to
- 30:55be cheaper than home or encryption but
- 30:59at the cost of uh greater communication
- 31:02burden so that's the
- 31:06tradeoff trusted execution environment
- 31:08or te it's a hardware approach to secure
- 31:12computation uh the most well-known
- 31:13examples are Intel sgx and amdv
- 31:16Technologies and what they provide as a
- 31:18way to uh create these isolated
- 31:20environments in a computer that protects
- 31:24the data inside it confidential uh keeps
- 31:27it uh confidential even from the person
- 31:29who's operating on these
- 31:32machines and the idea is that we can use
- 31:34this to send private data into this
- 31:36environment isolated environment of
- 31:38secure Enclave do some processing on
- 31:40them and return the results to the user
- 31:43uh through encrypted channels so that no
- 31:45private information is leaked during the
- 31:48process the key feature of these
- 31:50Technologies is what's called remote
- 31:52adastation and this is the use of
- 31:54cryptographic signatures that actually
- 31:56allows us to verify that um the The
- 32:00Enclave environment has been created
- 32:02faithfully and also the fact that the
- 32:04program that's running inside it is what
- 32:06we initially uh assigned to it to
- 32:11run so in this environment we're working
- 32:14with non- encrypted data so there's no
- 32:16computational overhead here or very
- 32:17minimal computational overhead but this
- 32:20comes at the cost of uh relaxed privacy
- 32:22notion since we're relying on a secure
- 32:24Hardware
- 32:26component
- 32:28you've probably heard of Federated
- 32:30learning this is popularized by Google
- 32:33several years ago initially designed for
- 32:36uh this use case of training machine
- 32:38learning models across millions of
- 32:40mobile devices the idea is fairly simple
- 32:42and that you know we have a model and
- 32:44then we compute these local model
- 32:46updates on each device aggregate them
- 32:49aggregate them across devices and apply
- 32:52the uh the global update to the the
- 32:54shared model and redistribute them and
- 32:57we're repeating this process to improve
- 32:58the accuracy of the overall
- 33:00model this has um you know immediate
- 33:04extension to the biomedical setting
- 33:06where we have a smaller group of
- 33:07collaborators we can also try to jointly
- 33:09train machine learning models this
- 33:12way the focus so far has been mostly on
- 33:15uh machine learning models or deep
- 33:16learning models that can be trained with
- 33:18gradient based optimization so the
- 33:19relevance for uh you know some of the
- 33:23genomic analysis tests may be limited
- 33:27because we're using different types of
- 33:28statistical models but it has found
- 33:31really useful application in in medical
- 33:33image analysis domain where we do need
- 33:36these deep learning models uh that are
- 33:38trained with
- 33:40gradients and this aggregation of local
- 33:42Updates this can reveal some private
- 33:45information from each party that's
- 33:47participating in this workflow so this
- 33:49needs to be carefully considered and
- 33:51typically these methods are combined
- 33:53with other techniques like encryption or
- 33:56differential privacy which is what I'll
- 33:57mention in the next slide to increase
- 33:59the security level of this
- 34:02solution so uh differential privacy it's
- 34:05addressing a slightly a different
- 34:06problem where we're concerned with the
- 34:08leakage of private information and the
- 34:10data that is released at the end of the
- 34:12computation so we can think about chwa
- 34:14statistics for
- 34:15example the key concept here is this
- 34:18notion of neighboring databases so these
- 34:20are two hypothetical uh databases that
- 34:23differ in exactly single individual and
- 34:26the idea is that we want to introduce
- 34:28some noise into the system or the into
- 34:30the computation so that only based on
- 34:33the release data we can't really
- 34:35distinguish between these two
- 34:36neighboring databases for any two
- 34:38neighboring
- 34:40databases and this privacy guarantee is
- 34:43typically controlled by this privacy
- 34:44parameter
- 34:46Epsilon and in this case values of
- 34:48Epsilon that are closer to zero means
- 34:50higher privacy and greater values
- 34:53um mean lower privacy but at the same
- 34:56time it allows us to roduce a smaller
- 34:58amount of noise so the results tend to
- 35:00be more accurate so this framework has
- 35:03been applied to Jo statistics release
- 35:05and other types of biomedical database
- 35:07queries like looking for corts that
- 35:09match certain
- 35:11criteria um but for many of these
- 35:15existing studies uh they were limited to
- 35:19uh specific simpler analysis settings uh
- 35:22for for example for the the case of
- 35:25chios we're instead of releasing entire
- 35:28genome wide Vector we can only release
- 35:30say like the top K most significant hits
- 35:32or something like that that's been the
- 35:34typical setting that it's been study and
- 35:37the the main reason is that the the
- 35:39actual noise that we need to introduce
- 35:41to achieve differential privacy
- 35:42guarantee tends to be very high uh in
- 35:45high dimensional
- 35:48cases so these techniques are useful in
- 35:52you know a range of different settings
- 35:53in biomedical research I've listed out
- 35:55some examples here that we've worked on
- 35:56over the years uh just to mention at a
- 35:58high level uh we can use these tools to
- 36:01build um uh softwares that can help
- 36:05researchers across institutions run
- 36:08collaborative studies in a secure way
- 36:10can also build analytic services that
- 36:12take users private data process them um
- 36:16and returns statistical insights back to
- 36:18them all in a privacy preserving way
- 36:21there's also this private data release
- 36:22setting similar to what I just mentioned
- 36:24with uh jwas where we want to release
- 36:27some privatized version of a bi medical
- 36:29data or analysis results with some
- 36:32meaningful guarantees of privacy for the
- 36:33underlying individuals in the data
- 36:35set and we're not the only ones working
- 36:38on this there's a growing community of
- 36:39researchers developing these tools I
- 36:41just wanted to mention a couple of
- 36:43examples of these uh real world
- 36:45competitions where we bring people uh
- 36:48together from different communities to
- 36:49create solutions for these tasks so idh
- 36:51is a popular one in this domain uh last
- 36:54year there was also a governmental
- 36:56collaboration between us and UK on
- 36:58developing tools based on these
- 37:00Technologies uh one of the tasks was
- 37:02creating a model for pandemic
- 37:05forecasting and uh some of us at bro
- 37:08entered these competitions and had
- 37:10winning solutions for for both of
- 37:13them so now I'll briefly go through some
- 37:16of our recent work in this domain as
- 37:19well uh so in our recent work we
- 37:22introduced this tool called secure
- 37:23Federated chws or sfj this is a
- 37:25cryptographic approach to performing Jos
- 37:28across multiple institutions without
- 37:30sharing any private information between
- 37:32them the key technical Insight that we
- 37:34introduce in this work is that as
- 37:36opposed to using a single technology um
- 37:39one of the ones that I mentioned earlier
- 37:40we can actually use a combination of
- 37:42them to gain computational speedups
- 37:44which tend to be quite important in
- 37:46practice so in this case we're combining
- 37:48NPC and homeric encryption which allowed
- 37:52us to design these Federated uh secure
- 37:55computation systems and what that means
- 37:58is that instead of encrypting the entire
- 38:00input data set and sharing it between uh
- 38:02the parties we can keep all the local
- 38:04input data sets local and only share
- 38:06encrypted versions of intermediate
- 38:08analysis results to carry out the global
- 38:10uh computation this reduces
- 38:13communication and also speeds up
- 38:14computation at the site level because
- 38:17now we're able to use unencrypted
- 38:19data and we were able to show that this
- 38:22um results in an order of magnitude
- 38:25improvement over the prior art that was
- 38:27only using MPC
- 38:30technology across these different uh
- 38:32data
- 38:34bases and this scalability Improvement
- 38:37allowed us to take these tools and then
- 38:40apply to really large biank scale data
- 38:42sets here I'm showing you two examples
- 38:43emerg in UK byy bank so the latter in
- 38:46our experiment included around 276,000
- 38:49individuals we're splitting these
- 38:51resources across uh six or seven
- 38:53simulated Federated units to demonstrate
- 38:57the ability to do um perform joint
- 38:59analysis without sharing private data
- 39:01and I'm showing you the results of chos
- 39:03that matches accurately with centralized
- 39:05analysis um where all the data is pulled
- 39:08together into a single
- 39:10location and the overall runtime for
- 39:13these pipelines uh tend to be around a
- 39:16few days for uh databases of this scale
- 39:20so this is not a you know Che cheap
- 39:23pipeline to run but uh the pract the run
- 39:26times are still practically feasible and
- 39:28these can be brought down further by
- 39:30introducing additional Computing
- 39:33resources and I want to point out that
- 39:36the um the experiment that we've run
- 39:39that we ran on UK bank is actually 2,000
- 39:41times larger than the previous Benchmark
- 39:43that we test tested in a previous
- 39:46study and also want to highlight that uh
- 39:50in this initial pipeline that we
- 39:51demonstrated we're looking at PCA based
- 39:54workflow where we run principal
- 39:55component analysis to look for these
- 39:57ancestry karious include as fixed effect
- 40:00term in the gwas model we're but it's
- 40:03also common to replace this with a
- 40:04random effect model uh just known as
- 40:07linear mix model approach to gas and
- 40:11this is generally regarded as being more
- 40:14accurate for jwas I I don't need to tell
- 40:16you this but it also comes at a greater
- 40:19computational cost even without any
- 40:21encryption so we weren't able to do this
- 40:23using secure computation techniques for
- 40:25a long time but recently there was this
- 40:27method that was introduced called
- 40:29regini and it's doing an efficient
- 40:32approximation of this random effect term
- 40:34using a whole gen regression model and
- 40:36what this allowed us to do is uh
- 40:38distribute this computational workflow
- 40:41across the parties using our
- 40:42cryptographic tools and also obtain a
- 40:44practical pipeline for uh running LM
- 40:47based
- 40:48Jos and um the distributed technique
- 40:51that I alluded to turned out to be
- 40:54fairly significant so we were able to uh
- 40:57remove dependence on the size of the
- 40:59cohort that goes into these stall
- 41:01studies as you can see in these runtime
- 41:04plots that stay nearly um constant
- 41:07across different datas as sizes we're
- 41:09able to show that the output of our lmm
- 41:13analysis workflow closely matches with a
- 41:15centralized execution of this regini
- 41:18tool on the pool data set so now we're
- 41:21able to run these different types of Jos
- 41:23workflows in a secure and Federated
- 41:26manner
- 41:27we had this related work uh more
- 41:30recently which was about finding genetic
- 41:33relatives across different data sets
- 41:35that cannot be pulled together due to
- 41:37privacy concerns and this is a key
- 41:40pre-processing pre-processing step in
- 41:42many studies or Federated databases like
- 41:45Nomad where we um where the presence of
- 41:48these relatives could skew the
- 41:50distribution of the statistics that we
- 41:52want to compute and release to the
- 41:54public so we have the cryptographic
- 41:56tools to be able to comp compute this um
- 41:58this relatedness coefficient privately
- 42:01but the the main challenge here is the
- 42:03fact that we need to perform all pawise
- 42:06comparisons between these two data
- 42:08sources and that turns out to be very
- 42:10expensive um so the key idea that we had
- 42:12in this work is that we
- 42:15could hash and bucket the individuals so
- 42:18that we can perform comparisons only
- 42:21between individuals that belong to the
- 42:23same
- 42:23bucket and the idea is that due to
- 42:26identity by this
- 42:27uh sharing between close relatives uh
- 42:32the close relatives have a higher
- 42:34probability of getting assign to the
- 42:35same bucket so this still retains high
- 42:37accuracy we're able to demonstrate that
- 42:40on UK by bank and all of us data you can
- 42:42see that uh this overall
- 42:45recall for detecting up to third degree
- 42:48relatives uh is very high it's um 97% or
- 42:53higher for in all three cases the
- 42:55Precision is also High higher than 98 %
- 42:58and some of these false positive
- 43:00findings are um arguably not false
- 43:03positives what they are is that they
- 43:05have a relatedness coefficient that's
- 43:07really close to threshold and due to
- 43:08numerical Precision it's Crossing that
- 43:12boundary and the overall runtime was
- 43:15also practical less than a day for all
- 43:16these data sets so where we head it next
- 43:19I'll try to conclude the talk in the
- 43:21next few minutes um so I think these
- 43:24Technologies are maturing enough that
- 43:27now we're able to see these uh inklings
- 43:29of practical tools for different genomic
- 43:32analysis tasks and I think over the next
- 43:34few years we'll see a lot more of these
- 43:36uh emerging for different analysis
- 43:39workflows including things like fine
- 43:41mapping or predicting disease risks
- 43:43based on personal genome so we're
- 43:45actively working on some of these ideas
- 43:47and I think these would be valuable
- 43:49tools for people doing genomic analysis
- 43:51across different sites uh to help with
- 43:55collaboration we are currently working
- 43:57on a deployment study where we're trying
- 43:59to run some of these joint analyses
- 44:01between million V million veteran
- 44:04program and all of us research program
- 44:05so as you know these are two of the
- 44:07largest fire banks in the US they're
- 44:09also known for um their strict security
- 44:12guidelines that prevent the data from
- 44:14being shared externally so we're trying
- 44:16to use our tools to create a link
- 44:18between them to um enable joint
- 44:21studies last year we launched our SF kit
- 44:24web server this is Joint work with grow
- 44:27data Sciences platform and this is a
- 44:30webbased service that allows people to
- 44:31run secure Federated tools the
- 44:34cryptography based ones that I mentioned
- 44:36uh in a push button kind of way uh and
- 44:40they can bring their own data set run a
- 44:42collaborative analysis with their
- 44:43collaborators uh using the service all
- 44:45the components that we have here are
- 44:47open source and this is actually the
- 44:49service that we're using for the ongoing
- 44:50pilot
- 44:51study and currently we're also working
- 44:54on integrating this uh set of tools into
- 44:58existing cloud-based analysis platforms
- 45:00like Tera which has also developed at
- 45:02Road as well as all of us researcher
- 45:04workbench uh which is based on Tera so
- 45:07these are some of the things that are
- 45:08coming up down the road and lastly I'll
- 45:10just mention that there's an upcoming
- 45:11conference that would be of interest to
- 45:13all of you so recom is happening it's
- 45:15being organized at MIT how
- 45:18convenient and there's actually this
- 45:20inaugural satellite conference on
- 45:22biomedical data privacy and equity which
- 45:24is happening one day before the main
- 45:26conference and we have several talks in
- 45:28both of these venues if you're
- 45:30interested please check out the programs
- 45:32and you can also reach out to me if you
- 45:33want to discuss more and with that I
- 45:35want to thank all of the uh former and
- 45:38Cur current labp members of my group and
- 45:41my collaborators and mentors uh in
- 45:44different institutions for all their
- 45:46input and support without them we
- 45:48couldn't have done you know a lot of
- 45:50these
- 45:50works thank you for your attention I'll
- 45:53take any questions great thank you so
- 45:55much Dr Cho I think we have time for one
- 45:58or two questions I might start them off
- 46:02um so if we think about how like each of
- 46:05us we're all required at publication to
- 46:07deposit all of our transcript omix right
- 46:09in a public data frame and I try to
- 46:12think of how we implement this does it
- 46:15mean that like the gene on our gene
- 46:17expression
- 46:18Omnibus data like should it be at the
- 46:22repository level that these things are
- 46:23implemented so researchers could still
- 46:25upload their kind of
- 46:27naive data frames as we do currently and
- 46:30then the the implementation of privacy
- 46:33happens at the repository or do you
- 46:35think this is actually something that
- 46:36where even that step is not secure
- 46:38enough and it should be uh the Privacy
- 46:40should be implemented prior to upload so
- 46:43I I can imagine a repository like Goo uh
- 46:48being run like something like DB Gap
- 46:50with the genotype data with stricter
- 46:52access uh control mechanisms I'm not
- 46:56necessarily proposing that needs to
- 46:58happen but we do have some concerns
- 47:00about the way currently transic data is
- 47:03being shared because they do Le
- 47:04genotypic information and it's not just
- 47:06our concern there are other people um
- 47:09who's been pushing on this as well and I
- 47:11know that there's an ongoing
- 47:12conversation with NIH on this topic so
- 47:15uh yeah there might be some changes in
- 47:17the future but we're not sure yet I I do
- 47:20want to make a note that the the risks
- 47:23that I showed about that data is kind of
- 47:25in a controlled environment right where
- 47:28we have this matching ancestry between
- 47:30like the the target individuals and like
- 47:32the other individuals in the candidate
- 47:34pool so there are like other
- 47:35considerations like that that needs to
- 47:37go in when we actually think about the
- 47:39the risk and the Practical
- 47:42setting Mak
- 47:44sense great well thank you so much for a
- 47:48broad sweeping talk that really
- 47:50highlighted I think so many different
- 47:51ways in which privacy impacts the
- 47:53research that we all do every day um and
- 47:55then please feel free to go grab second
- 47:57breakfast and then we'll start uh the
- 47:59next session at
- 48:079:30
About this transcript
This page contains the full transcript of MPG Primer: Genomic Privacy: Key Issues and Emerging Solutions (2024) by Broad Institute, generated from the public captions YouTube serves with the video. The transcript has 8,064 words across 1,251 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.