Foundation Models for Geodata Processing by Dr. Ashutosh Kumar Jha — Transcript
Full transcript
- 20:34Okay, good evening. Uh,
- 20:37today we'll be having a discussion on
- 20:40trying to understand the foundation
- 20:41models for Joda processing. So these are
- 20:44basically
- 20:45uh certain recent development which has
- 20:48happened in uh geospatial domain. So
- 20:51we'll be talking about uh those
- 20:53development. So we will not be going
- 20:54much in details just trying to
- 20:56understand what is going on uh currently
- 20:59in uh especially in the foundation
- 21:02models which is related to uh G data
- 21:04processings.
- 21:06So basically uh we'll just browse
- 21:09through the different uh evolution of AI
- 21:11vision models that are being used
- 21:14currently and there are different stages
- 21:16which has led to uh certain development
- 21:19and that has uh resulted in a situations
- 21:22where now we are seeing that lot of
- 21:25vision AI vision based models are having
- 21:27an accuracies which is at par to the
- 21:30supervised level of uh accuracy or
- 21:33classifications techniques or may be
- 21:36related to the task which is a human can
- 21:38do that. So the current systems are at
- 21:40least uh approaching to that kind of uh
- 21:43uh performance. Then we'll see a certain
- 21:46uh basic understanding uh of uh which is
- 21:50needed to uh look into different kind of
- 21:52learning approaches which are being used
- 21:55by these uh different vision models and
- 21:59we'll also focus uh slightly more on uh
- 22:02CNN and as well as the vidbased models.
- 22:04So I will not be touching more detail in
- 22:06CNN but rather than on just uh giving
- 22:10details on visual uh vision uh level
- 22:12transformations and how it is uh
- 22:15different from the CNN and so on. Then
- 22:17I'll just extend that to a
- 22:18self-supervised learning approach which
- 22:20is being used and then we'll be having a
- 22:22few geospatial foundation models which
- 22:25will be look into that. So if you see uh
- 22:28in in in basically in the vision based
- 22:32systems which is mostly related to a
- 22:34neural based approach uh the development
- 22:36started quite early in 1958 and probably
- 22:40might be have seen the seifur data sets
- 22:43uh some of you might have seen that. So
- 22:45digital identifications were the first
- 22:47task which was being attempted and
- 22:49accordingly there was a lot of
- 22:50development has happened. So perceptron
- 22:52was basically a kind of computational
- 22:54models uh which is uh based on how a
- 22:59neuron u works. So it's basically
- 23:02mathematical mapping of that is being
- 23:04done in term of the weighted uh uh uh
- 23:08summations of the input and accordingly
- 23:10the the mapping has been done in term of
- 23:12nonlinear uh approach using the
- 23:14exponential function has been used. So
- 23:16that's the basic principle of
- 23:17perceptron. Then there has been a
- 23:20certain development uh which is related
- 23:22to CNN. So CNN basically doesn't works
- 23:25on a pixels or in in term of image we
- 23:28can say pixel level point data sets
- 23:31rather than it takes into the local
- 23:33context. So it basically takes a window
- 23:35around surrounding the central pixels
- 23:36and surrounding that there will be a few
- 23:38pixels and that will be applied to uh is
- 23:42being applied for CNN cases. So CNN has
- 23:44an advantage that it gets a kind of
- 23:47pixel informations along with the
- 23:49context and that's how the uh the
- 23:52classifications
- 23:54uh is being uh uh basically is being
- 23:57improved compared to pixel level
- 24:00perceptron based systems. Then we had uh
- 24:03recurrent neural networks or LSTM. So
- 24:05these are basically using the first
- 24:08level of output of uh one CNN or one
- 24:11neural network based systems giving the
- 24:13output and then feeding it to the next
- 24:16uh level. So generally it is being used
- 24:17in a temporal uh cases where in the CNN
- 24:21and uh perception related uh neural
- 24:24network you don't have any memory. So
- 24:26whatever has happened earlier in
- 24:28instances that is not being taken into
- 24:30consideration but LSS and RS RSTM
- 24:33basically takes into those uh into the
- 24:35considerations. Then there was a big
- 24:37leap in uh 2012 and the imageet
- 24:41classifications was been put by
- 24:45Google and that was a kind of remarkable
- 24:48achievement uh especially the using the
- 24:50deep neural network. So whatever
- 24:52convolutional neural network which was
- 24:54being used earlier that was restricted
- 24:56to maybe two or three layers but later
- 24:58on in the image net they have gone into
- 25:01quite deeper 18 20 layers of their depth
- 25:04has been there. So due to that what has
- 25:06happened is there has been a lot of
- 25:08contextual information also uh
- 25:11percolating in the started percolating
- 25:13at the upper layer of uh neural networks
- 25:17and that's how the uh each of the uh
- 25:20informations now have a kind of global
- 25:22context which is available in the images
- 25:25that gets carried away and accordingly
- 25:27the improvement has been there in case
- 25:30of uh uh the other uh vision level
- 25:33information there has been a tremendous
- 25:34jump in uh uh basically the accuracy
- 25:37level which is expecting in after the
- 25:40development of alexnet and imageet
- 25:42cases. Then there was a kind of semantic
- 25:45uh unit. So this is this was the outcome
- 25:48of the there was few observations which
- 25:51was being made after the DNN based uh
- 25:54model started developing. So what has
- 25:57been observed that uh the earlier
- 25:59advantage of increasing the depth of CNN
- 26:03that started losing I mean maybe after
- 26:0520 or 30 layers the the the observation
- 26:08was there that uh as you go deeper and
- 26:11deeper then that there's no extra
- 26:12information gets extracted and the upper
- 26:15layer basically gives you the kind of
- 26:16noise level informations. So there is
- 26:19there has not been any improvement. So
- 26:22later on the unit architecture which
- 26:24basically uh has a kind of uh new
- 26:27systems where encoder and decoder
- 26:29architectures gots added with uh uh
- 26:32extra layer where the semantic
- 26:34informations from the lower level uh uh
- 26:38networks has been passed on to the upper
- 26:40layer and that's how the improvement has
- 26:42happened uh with the deep neural network
- 26:45cases. Then there was another
- 26:47development which has happened uh and
- 26:49there's a paper which was being
- 26:51published by Google's uh team that was
- 26:54basically attention all you need. It's a
- 26:56very small uh title but it has a very
- 26:59big impact in uh overall vision systems.
- 27:03So what they have uh trying to do or
- 27:06tried to do under the uh uh vision based
- 27:10systems. So they started uh a kind of
- 27:13model where they say that uh instead of
- 27:15predicting and characterizing and
- 27:17classification let's do a prediction of
- 27:20the neighborhoods or you can say that
- 27:23the next points which is being possible.
- 27:26So generally in the NLP it is being used
- 27:28like if you just put it a few text and
- 27:31the correspondingly what may be the next
- 27:33text that get predicted. So same kind of
- 27:36mechanism also got uh uh uh basically
- 27:40carried into the vision based systems as
- 27:42well at later stages. In between there
- 27:45has been a kind of physics based
- 27:47simulations and other model which has
- 27:49also been people started trying to
- 27:52integrate. So instead of completely
- 27:54relying only on what data says and the
- 27:56people have started realizing or
- 27:58creating certain physics based uh uh
- 28:01network a kind of uh model simulation
- 28:04environment where physics based
- 28:06conservation of certain laws maybe the
- 28:09conservation of energy conservation of
- 28:11uh momentum all those things has been
- 28:13embedded and that those were basically
- 28:15specifically useful in the cases of
- 28:18geospatial domain. For example, in the
- 28:20weather pattern, it is uh the cloud
- 28:23movement is not sufficient. Cloud also
- 28:25if you take a cert lots of images of the
- 28:28cloud movement that is not sufficient to
- 28:31give idea about the weather pattern. You
- 28:33also need to know the certain physics or
- 28:35atmospheric physics related to the
- 28:37neural uh basically the numerical
- 28:39weather prediction models that needs to
- 28:41be integrated to have a uh better uh
- 28:45forecasting models. So these were
- 28:47basically started coming in in 2019.
- 28:51Then there was been uh uh vision based
- 28:53transformation which was basically from
- 28:562017 transformer now got into started
- 28:59getting into vision based
- 29:00transformations and again uh the similar
- 29:03paper was there. So you just said I mean
- 29:05in in that paper basically it says that
- 29:08you need only 16x6 pixels uh as an input
- 29:12to get any kind of vision task. So
- 29:14that's was their claim and that has been
- 29:17developed in 2020. Then slowly there was
- 29:20uh also been observed that during these
- 29:24times as we had a lot of uh data or I
- 29:28should say the data generation systems
- 29:30has also increased you have a lot of
- 29:31data coming from different sources uh
- 29:35social media and other places. So slowly
- 29:37what has happened is the the input rate
- 29:41of the data has outpaced the development
- 29:43of the uh basically uh you you can say
- 29:47the supervised labeling approaches. So
- 29:50earlier these systems were being built
- 29:52on the basis of labeling of the
- 29:53different uh sections in the images and
- 29:56that was being used in the supervised
- 29:57cases but as the data generation things
- 30:00has outpaced. So the impact of that was
- 30:03that people were not able to use the uh
- 30:06the higher rate of the data sets and in
- 30:08that scenarios there was a kind of uh
- 30:11self supervised vision based system has
- 30:14started coming. So what uh model has
- 30:17basically started uh has been developed
- 30:20uh with a with a objective that instead
- 30:23of uh trying to label it let's create
- 30:26certain things in between. So in between
- 30:29means there will be a certain latent
- 30:30spaces which informations can be created
- 30:33and on the basis of that you can uh do
- 30:36uh uh a kind of reconstruction of the
- 30:38whole scene and images. So we'll see few
- 30:40of those algorithms uh or the steps
- 30:43which is in the coming slides. Then
- 30:46there has been uh further development uh
- 30:48slowly that self-supervised learning
- 30:50approach has uh been adapted uh by lot
- 30:54of geospatial community especially the
- 30:56people who have big pockets. So they
- 30:59could consume the 40 30 years 40 years
- 31:02of uh the whole LANCAT imagery uh into
- 31:06modeling and then they had a lot of GPU
- 31:09power. So those are actually GPU rich uh
- 31:11I should say uh companies they invested
- 31:14a lot of money and accordingly we had a
- 31:16lot of uh uh developed uh the foundation
- 31:19models started coming. So these basic
- 31:22objective of these foundation models are
- 31:24that whatever the general information
- 31:27which is present uh in in any kind of
- 31:30geospatial data. So that has to be
- 31:32mapped into a certain uh latent space or
- 31:35intermediate uh space then we can have a
- 31:39certain upstream task which can then be
- 31:41used for doing some kind of basic task
- 31:43like classifications segmentations and
- 31:46so on. Now there has been a uh again
- 31:49continuously the development has been
- 31:50there and there has been an integrations
- 31:52of multimodel foundation architectures.
- 31:55So what this does is basically instead
- 31:57of just relying and giving an image and
- 31:59then predicting it use uh it it has
- 32:02started coming in term of summary. So
- 32:04the the Google has started giving a lot
- 32:07of services where based on the
- 32:09prediction of the input data sets it
- 32:11gives you the weather conditions and
- 32:13based on the different weather
- 32:14parameters like precipitations uh cloud
- 32:17coverage which model gets generated then
- 32:20on the basis of that that gets uh
- 32:22converted into a certain smaller text
- 32:25which can be sent to the user that there
- 32:27is a chance of certain amount of rain or
- 32:29there will be heavy rain or
- 32:31thunderstorms. So these all those uh
- 32:33models uh which was basic project uh
- 32:36basic object the models basic objective
- 32:38was earlier to predict the weather
- 32:41patterns or the land use land maps uh
- 32:44patterns instead of that now they
- 32:46started giving a specific prescriptive
- 32:49or user level summarizations of what
- 32:52those uh model means. So that was a big
- 32:55leap and the development is still going
- 32:57on in these domains. So these on the
- 32:59bottom you can see there's a timeline of
- 33:01all these uh basic vision based models
- 33:04which we just uh discussed. So out of
- 33:07those I'll just focus only the uh basic
- 33:10structures or the basic things which has
- 33:13been there especially in in term of
- 33:14segmentations which was based on
- 33:17convolutional neural networks especially
- 33:18the image based uh networks and the
- 33:21unit. So uh that was the basic uh thing
- 33:25which which was there and then later on
- 33:27the NLP was being uh used uh and that uh
- 33:31and in in the case of NLP you have an
- 33:34attention based system. So what exactly
- 33:35is the attentionbased uh systems so in
- 33:39the attention based system what happens
- 33:40is that you will be given certain
- 33:42inputs. So the way in the in term of
- 33:45language context we can say that uh
- 33:47let's say we say that there is a man.
- 33:50[clears throat] So after that what is
- 33:52the next word it is going to come that
- 33:55we can do a predictions on the basis of
- 33:57the language context uh its uses the
- 34:00word after the man how many times uh the
- 34:03man comes after uh a word a or what what
- 34:07is the next word which is coming after
- 34:09the man and so on. So those kind of uh
- 34:12scenarios gets generated based on the
- 34:15historical or large data sets available
- 34:18in uh available from different corpus
- 34:21and so on. So accordingly we get a kind
- 34:24of prediction of the next level. So this
- 34:26is what the transformer's job is that
- 34:29taking all the input which is available
- 34:32uh maybe a sequence of word maybe
- 34:34sequence of image patches and so on and
- 34:36then on the basis of that it tries to
- 34:38predict what may be the next word or the
- 34:41next picture. So if you are creating a
- 34:44video so it will get it will give you it
- 34:46will take a set of sequence of the
- 34:48images and after some time it will also
- 34:50generate what may be the next sequence
- 34:53of the images. So that's the kind of
- 34:54attention mechanism is being created or
- 34:57in case of images there might be a cases
- 34:59for example you may be having a very big
- 35:01images out of that you may give randomly
- 35:04four or five different patches and then
- 35:06you say reconstruct the whole images. So
- 35:08those are actually the kind of cases uh
- 35:10which is being done with the help of
- 35:12attention based uh uh models and then
- 35:16we'll also see the foundation based uh
- 35:18model era where these vision based uh uh
- 35:23models how it got integrated to
- 35:25different geospatial kind of
- 35:27informations like you have a multiple
- 35:29spectral band so that multisspectral
- 35:32band data and the ge uh geographical
- 35:35locations those information then got
- 35:37integrated and accordingly the uh the
- 35:41basically there has been an improvement.
- 35:43So if you see that uh at the bottom
- 35:45there's an accuracy table. So now at
- 35:47least we are in a in a stage where we
- 35:50are having an accuracy of all these uh
- 35:53foundation uh level models which where
- 35:56accuracy of the uh data sets or the any
- 35:59kind of image in in term of
- 36:00classifications you see it is around 94
- 36:0395% which is at at par with the human
- 36:07level uh informations.
- 36:09So let's see that uh how those uh u
- 36:13method has evolved. So earlier we say
- 36:15supervised uh learning I think most of
- 36:18you may be aware of that. So what you do
- 36:20you give an input and you create a
- 36:22certain labels and you say this input
- 36:25basically means this. So there may be a
- 36:27kind of uh you may be having a satellite
- 36:29imagery which will be which may get
- 36:31collected uh from different band uh band
- 36:35sensor data sets or in in normal cases
- 36:37we say you have a just clicked
- 36:39photograph from a camera and then from
- 36:42that camera you can say okay this is the
- 36:44human being this is a dog this is a cat
- 36:46and so on. So that labeling you do and
- 36:48then basically the systems tries to
- 36:51learn that how to differentiate the
- 36:54input feature which is the pixel level
- 36:56concentrations or organizations of those
- 36:58pixels to understand okay in in case of
- 37:01dog what kind of organization of those
- 37:03pixel will be there or in case of human
- 37:05beings what kind of organization of
- 37:06those pixels will be there those are
- 37:08actually supervised learning approach
- 37:11under the unsupervised learning we don't
- 37:13uh know how many classes are there but
- 37:16instead of that we look for certain kind
- 37:18of similarity. So those similarity may
- 37:20be at the pixel level or maybe at a
- 37:23regional level or maybe uh at an image
- 37:26level itself and then uh the systems
- 37:29tries to group them in a different
- 37:31category. So let's say you'll be having
- 37:33a 10 photographs. So what system will
- 37:35say that okay these two photo first two
- 37:37photographs looks like uh place uh of
- 37:41one place another two or three
- 37:42photographs looks maybe having a certain
- 37:45other places and so on. So it will give
- 37:46you those kind of grouping but you uh it
- 37:49will not tell you what those grouping
- 37:50means. So as a human being you have to
- 37:53tell okay these groups basically means
- 37:55uh maybe a hilly region maybe a uh uh
- 37:58you can say the flate uh flood plane
- 38:02regions and so on. Those kind of things
- 38:04uh can be carried out and these are very
- 38:06standard process which is being followed
- 38:08in geospatial domain. In case of
- 38:11self-supervised learning what happens is
- 38:13so the users is not giving an uh input.
- 38:18So rather than what happens is the
- 38:20machine is going to take your input it
- 38:23itself will have a certain pipeline
- 38:25where it will try to distort it. So it
- 38:27will do uh certain error which may be
- 38:30introduced into the images and then
- 38:33finally uh uh it will try to see what is
- 38:36the impact of those changes or those uh
- 38:40alterations process which has it which
- 38:43has taken place and accordingly it will
- 38:45try to understand uh what are the basic
- 38:48underlying uh structure. For example,
- 38:51let's say if one image is being given.
- 38:53So if you rotate it that particular
- 38:55image or that image let's let's say of a
- 38:57dog. So it will remain as a dog. So you
- 39:00rotate it 90° 60° or 50° there's no
- 39:03impact of rotation. Finally at the end
- 39:05of the day you have to say that image
- 39:07belongs to a dog. So in that case the
- 39:09machines basically rotates it in a
- 39:11multiple uh sections and then it will
- 39:14give you a kind of output where it will
- 39:15say okay it has understood that uh even
- 39:19uh rotating it multiple cases whatever
- 39:21the feature differences were there that
- 39:24should be projected in such a way that
- 39:26it should uh tell that this is a
- 39:28particular object uh is related to one
- 39:30specific object as a whole. Then there
- 39:33are other uh approach which is the
- 39:34reinforcement learning approach. So
- 39:36which is again uh a kind of uh I
- 39:39couldn't I should say this without
- 39:41labeling approach. So there instead of
- 39:43telling each and every features what you
- 39:45tell the systems that you are right or
- 39:48wrong or there may be a certain degrees
- 39:50can also be provided. So on the basis of
- 39:53that uh machines takes the actions and
- 39:57then it learns on the basis of that. So
- 39:59these are the major four learning
- 40:01approaches. So generally the foundation
- 40:04models or geospaceial foundation model
- 40:06works on a principle of uh
- 40:08self-supervised uh learning method. So
- 40:11under the self-supervised learning
- 40:13method what you do is you take an
- 40:15images. This is just an example where
- 40:17you do uh a certain uh pretext task. So
- 40:21that is nothing but a kind of distorting
- 40:23the input images. So these distortions
- 40:25can be you can take an image break down
- 40:27into multiple component and then you
- 40:29delete certain portions and then you say
- 40:32please predict the deleted portions.
- 40:35Other cases you may take certain patches
- 40:37you may rotate it or make some uh noises
- 40:40in all all those patches. So finally
- 40:43system needs to learn and tell that okay
- 40:45what kind of uh distortion or noise
- 40:48might have been added which might have
- 40:50resulted in the cases uh in that cases.
- 40:54So in generally what happens is you take
- 40:56an input data then you do do a certain
- 40:58pre uh task basically rotating breaking
- 41:01masking those kind of uh uh task which
- 41:04you carry out and then you send those
- 41:07partial informations to the uh a set of
- 41:11uh neural networks uh uh large neural
- 41:13networks or DNN networks which basically
- 41:17takes those data sets and create a kind
- 41:19of intermediate multiple or I should say
- 41:22multi-dimensional data assets
- 41:24projections which finally can be used
- 41:27for different downstream task like you
- 41:29may be interested in doing some kind of
- 41:31classifications segmentations and
- 41:34control and so on. So for example let's
- 41:36say at the bottom you see a lot of uh
- 41:38pen chromatic colors uh or maybe you may
- 41:42be having a true color images or SAR
- 41:44images multisspectral images and so on.
- 41:46So these data sets may be given to the
- 41:49input. So at the input level it will
- 41:51just try to understand okay there is a
- 41:53huge intensity at certain places there
- 41:55is a lesser intensity at other places or
- 41:58your system may also try to get certain
- 42:01additional information like okay if
- 42:03there is a high intensity so it is
- 42:05surrounded by a lower intensity or what
- 42:08is the probability of having a higher
- 42:10intensities data sets to be having an
- 42:12higher uh intensity in surrounding
- 42:15areas. So for example in the SAR as you
- 42:17can see there is a very less number of
- 42:20uh uh high intensity color which is
- 42:23being surrounded by uh the uh the high
- 42:27another high intensity areas. But in
- 42:30case of pan chromatic you can see
- 42:31there's a white patches so which is
- 42:33quite bigger and so on. So this kind of
- 42:36distinction it will try to understand it
- 42:38will not know okay this bright color
- 42:40means what is it a pond or this is a is
- 42:43it a kind of river bed or it is a
- 42:46building those kind of information it
- 42:48doesn't have but it just knows that okay
- 42:50within the image there are certain
- 42:52places where uh you're having a high
- 42:55intensity and low intensities there's a
- 42:58texture variations there is an intensity
- 43:00variations and so on I mean those kind
- 43:02of basic uh uh uh the parameters which
- 43:07which basically gives just information
- 43:10in term of color intensity their
- 43:12relationship and so on. So once these
- 43:14information gets collected in a latent
- 43:16space then finally you can create a
- 43:18certain downstream task. So in that case
- 43:20the advantage is that we don't need to
- 43:22build separate separate model for one
- 43:25model for object detection another model
- 43:27for classifications third model is just
- 43:29doing a change detections. So that was
- 43:32the earlier uh approach. So in the CNN
- 43:34models you have to just take let's say
- 43:36for example using list three if you do a
- 43:38classifications if you try to use the
- 43:40same uh classification algorithms let's
- 43:43say on the landset probably you will not
- 43:45be able to do it. So so that kind of
- 43:48thing now has been avoided with the help
- 43:50of uh these uh uh self-supervised uh
- 43:54learning approach. So I'll just focus
- 43:56the basic two models which is being
- 43:59mostly used in uh remote sensing domain.
- 44:02So one is that you have called a CNN
- 44:05model. So CNN and DNN are all those
- 44:08families of of a models where what it
- 44:11does is you take a very large area then
- 44:14you create a certain window size of the
- 44:16smaller area. So those window size maybe
- 44:183x3 maybe 16x 16s and so on and then you
- 44:22can create a kind of multi-layer window.
- 44:25So for example as you can see in the uh
- 44:28left side of the CNN images. So all all
- 44:32the images has been divided into 16
- 44:34boxes. So each of these boxes will
- 44:36having their own informations. So like
- 44:39first four on the top or maybe you can
- 44:42say the topmost area which we see in the
- 44:44CNN models it's an urban areas and so
- 44:47on. But in between where you see a
- 44:49larger 4x4 boxes which has been merged
- 44:52together. So that looks at green uh
- 44:55areas. So CNN basic job is to understand
- 44:59the spatial context and on the basis of
- 45:02just to look into the textural
- 45:05informations. For example, it will just
- 45:07say that okay there is a kind of green
- 45:09area which is there inside a in in
- 45:13somewhere. Similarly there may be a
- 45:14certain areas which is having a high
- 45:16textured uh uh areas uh which may be of
- 45:20the urban buildups and so on. So what
- 45:23happens is it is not able to tell you
- 45:25that this green area basically means a
- 45:28forest or it is a golf course or maybe
- 45:32uh a kind of park and so on. That kind
- 45:34of context is uh the CNN is not able to
- 45:38give you why because uh the CNN is not
- 45:41having any kind of global context. So
- 45:44what what does it mean is let's say you
- 45:46take any boxes in the CNN one. So if you
- 45:50see this um one of those boxes so it
- 45:53sees okay on the left side there's
- 45:54another box and so on but it doesn't
- 45:57understand that this box actually
- 45:59belongs to a larger green patch area
- 46:02because it is not having understanding
- 46:04of the whole scene which is present in
- 46:06the image. So accordingly it loses that
- 46:09global context and it is not able to get
- 46:12you the informations that what exactly
- 46:14this green piece is there. So, so but in
- 46:19case of vision based transformer or the
- 46:21earlier which I was talking about the
- 46:23attention which we need or image of 16x6
- 46:26pixel those papers basically talks about
- 46:28that. So in those cases what happens is
- 46:31it the vision based global detention
- 46:34mechanism they gets the context of
- 46:37global context as well. So in that
- 46:40context it sees okay if there is a kind
- 46:42of green area and it is surrounded by a
- 46:45lot of uh uh uh non-green area which
- 46:48looks like a kind of uh builtup areas.
- 46:51So the green patch in between has to be
- 46:53an urban park. It cannot be a forest
- 46:56because within the city it is not
- 46:59expected that you're going to have a
- 47:00forest. these kind of uh global uh I
- 47:04should say the
- 47:06uh uh basically context which vision
- 47:09vision based transformer is able to get
- 47:12in and this is actually a quite long
- 47:14range. So CNN is basically having a very
- 47:16small field specific uh smaller range uh
- 47:20informations but vision based
- 47:22transformer are having a global context
- 47:24and accordingly it is going to give you
- 47:26a kind of uh uh larger context and and
- 47:30it is able to give you the uh
- 47:33classification or categorizations much
- 47:36better in those cases. So these are just
- 47:38an example like how does CNN basically
- 47:40relies mostly on convolutional and max
- 47:43pooling layers which is generally been
- 47:45used and where in case of vision based
- 47:48transformer basically relies on uh the
- 47:51attention based mechanisms. So what does
- 47:54attention based mechanism and how it
- 47:56works we'll just see in the coming uh
- 47:58slides as well. So let's see the other
- 48:01examples where the how does the CNN and
- 48:04vision based transformer models can be
- 48:07useful. So as at the bottom you can see
- 48:09that there is a kind of original images.
- 48:12In case of CNN if that same area you
- 48:15take the another image of another time
- 48:18and it is cloud cover or it is having
- 48:21some kind of uh let's say uh you're
- 48:24having some kind of fog or maybe defocus
- 48:27all those kind of problems are there. So
- 48:29even if you have a very well-trained CNN
- 48:32model but if the input data set gets
- 48:34distorted or uh uh in in in in those
- 48:38cases what happens is the CNN model will
- 48:41fail and they will not be able to give
- 48:43you the exact uh output uh and giving
- 48:46you the clear pictures what exactly it
- 48:48is there where in case of vision based
- 48:51transformer what happens is that any
- 48:53kind of augmentations or error or noises
- 48:55which gets uh added so what happens is
- 48:59even if there's an uh uh uh let's say
- 49:02the contrast changes or there may be a
- 49:04certain uh uh missing informations and
- 49:08so on still what it will able to do is
- 49:11it will take whatever the partial
- 49:13information is available and then it
- 49:15will try to fulfill the uh or I should
- 49:18say that it will try to fill the missing
- 49:21data sets. So once it is able to fill
- 49:24the missing data sets then it is able to
- 49:26uh reconstruct the whole images which is
- 49:30uh which which is which is actually uh
- 49:32you can say the filtering of the input
- 49:35uh noises has been removed from the
- 49:37input images and accordingly what
- 49:39happens is you get an output which gets
- 49:42generated and then that particular image
- 49:45can be used for classification like
- 49:47urban park and so on. So this is basic
- 49:49advantages where in case of uh vision
- 49:53based transformer model it not only
- 49:55takes the input but it corrects the
- 49:57input data sets and then that correction
- 50:00is being done on the basis of global
- 50:02context which it might have learned
- 50:04during the learning process and then
- 50:06finally it gives you the categorizations
- 50:08and or other downstream classes like
- 50:10classifications or the uh other uh uh
- 50:15basically differences or maybe the time
- 50:17based
- 50:18changes all those things it can it can
- 50:20do it on the basis of those
- 50:22reconstructed image. So this is the
- 50:24basic uh differences which the
- 50:27transformer has built uh has brought in
- 50:30into on the table table and accordingly
- 50:32what has happened is this has been
- 50:34exploited by a lot of uh different uh uh
- 50:38foundation models development. So let's
- 50:41take an example of what kind of uh
- 50:44whenever we say that uh I I had said
- 50:46earlier that in case of self-supervised
- 50:48learning what happens is your uh model
- 50:52takes the input it tries to distort it
- 50:55certain input. So for example at the uh
- 50:58there's basically a creating a fake
- 51:00images. So as you can see that uh uh
- 51:03this is basically a kind of one approach
- 51:05where you just provide one input and
- 51:08model uh there will be a certain uh
- 51:10channel or you can say uh uh one models
- 51:13will be there or one uh maybe simpler
- 51:16models may be there related to where
- 51:17just doing an pixelbased rotations or
- 51:21certain other augmentations distortions
- 51:23can be there that you may be having
- 51:24instead of just rotations you may be
- 51:26having a uh similarity measures uh CNN
- 51:30and DN models which will give you a
- 51:33multi uh images or different type of
- 51:35images which will be almost similar but
- 51:38in different cases for example as you
- 51:40can see in the deer pictures which is
- 51:42being taken as an input then what we'll
- 51:44do is it the in in the CNN based models
- 51:47uh which is basically an exampler CNN
- 51:50based models uh especially uh has come
- 51:53in 2015 so what it does it gives you the
- 51:56different perspective of the same deer
- 51:58in a different contrast colors
- 52:00and so on. So it get those it gets
- 52:03created. So all the other images which
- 52:04you see except from the first one is a
- 52:07fake one. So this is the first pre-text
- 52:10task which is being carried out by the
- 52:12any kind of uh I should say
- 52:14self-supervised learning approaches
- 52:16where it will generate a lot of such a
- 52:18fake data sets and then try to
- 52:22understand that even though those fake
- 52:24data sets may be generating from the uh
- 52:26from the same uh single images but all
- 52:29of those fake data sets actually gives
- 52:31you similar context like it is all these
- 52:33images which you see is basically a kind
- 52:35of image of a deer. there. Similarly,
- 52:38there are another augmentation or
- 52:39distortion method which is just trying
- 52:41to learn or create uh rotations. So, you
- 52:44take an images, you do a multiple
- 52:46multiple rotations. So, your objective
- 52:48is even if you give uh even the model is
- 52:52being given the first image or maybe any
- 52:54four distorted image which is rotated by
- 52:56a different angles, the model should be
- 52:59able to tell that all those four images
- 53:02are actually same as the first image. So
- 53:05that's the other kind of distortion
- 53:07which self-supervised uh systems tries
- 53:09to create and trying to understand what
- 53:12exactly is the impact of uh rotations.
- 53:16Then so those are basically of simpler
- 53:18cases. Then in case of positional
- 53:21augmentations what happens is you take a
- 53:23randomly a certain images and then you
- 53:26randomly create certain neighbor
- 53:28patches. So neighbor patches need not
- 53:30have to be connected which is generally
- 53:33we expect uh in other CNN models. So
- 53:36like uh for example in the image as you
- 53:38can see there are nine uh different
- 53:41boxes. Each of these boxes have a
- 53:43relative positioning of each other. So
- 53:46and accordingly we can say that patch
- 53:48number one is actually on the uh west or
- 53:52northwest to the uh blue pixel.
- 53:55Similarly patch number two is actually
- 53:57not to the blue or those kind of
- 54:00positioning is fixed. So in this case of
- 54:03augmentation images what happens is you
- 54:06just give at any two pair of uh these
- 54:09patches to the models and then uh the
- 54:12model should say and tell you that where
- 54:15this patch basically belongs to which
- 54:18particular positions. So advantage of
- 54:20this is as you can see now if you build
- 54:22such kind of uh relative positioning
- 54:24your system model is able to learn then
- 54:27what happens is instead of having a few
- 54:29receptive uh window specific uh CNN
- 54:33focused uh I should say informations now
- 54:36it gets extrapolated beyond that so it
- 54:39is able to see beyond that in a whole
- 54:42image context and the impact of that is
- 54:44that you'll be able to uh uh get a
- 54:48larger uh context. So as we can see if
- 54:51you're able to see that at the center
- 54:53you're having a blue boxes and there is
- 54:55another uh place and the third basically
- 54:58which is on the topmost uh I should say
- 55:01the right side topmost that is the
- 55:03position number three and if you're able
- 55:05to tell and do a prediction very well
- 55:08and as you can see that if you're able
- 55:10to understand that at the blue it's a
- 55:12basically a face of cat and the three
- 55:16basically is a kind of cases which is
- 55:19just on the top left and right then you
- 55:21can say okay three is actually maybe
- 55:23most probably a kind of ear of the cat
- 55:26it cannot be a kind of leg or other uh
- 55:30cases. So these kind of additional now
- 55:32you can create start creating a kind of
- 55:34contextual informations which might be
- 55:36quite useful. So in those cases
- 55:38generally uh uh is being done. So it's a
- 55:41basically relative positioning.
- 55:42Similarly the other approaches is
- 55:44basically jig jigsaw puzzling. So what
- 55:46you do you give all those uh uh you just
- 55:49take those nine points or the nine
- 55:52patches surrounding and you randomly
- 55:54arrange them. So model's job is just to
- 55:56give exact positionings and that's how
- 55:59these systems is going to learn. In some
- 56:02cases you can also feed it to some fake
- 56:04images which is not part of these images
- 56:07and systems should be able to tell okay
- 56:09this doesn't belongs to any of those
- 56:11eight places this is somewhere coming
- 56:13from outside. So these kind of
- 56:16augmentation patching is the uh certain
- 56:19learning approaches which gets added to
- 56:21the self-supervised uh learning. So just
- 56:25to give you an idea you have original
- 56:27image and then you do a certain uh
- 56:29augmentations. So certain augmentations
- 56:31may be quite easy for example just
- 56:33changing the color or making a red to
- 56:35blue or green or something like that. So
- 56:38those are quite augmentation uh those
- 56:40are called weak augmentation techniques.
- 56:42uh basically it means that it is just uh
- 56:45uh the overall the context of whole
- 56:48image is still present uh uh in term of
- 56:51objects uh which is being there. Then
- 56:54there are something called strong
- 56:55augmentations. So which is basically
- 56:57distorting in a much larger way. So
- 57:00which means not only decreasing
- 57:02increasing the brightness putting lot of
- 57:04additional images. For example, as you
- 57:06can see that if you improve, if you
- 57:08increase the saturations, the context of
- 57:11the dog itself is completely lost. Now
- 57:15the system has to kind of learn and find
- 57:17it out. So if if you put if your system
- 57:20is able to get hold of these stronger
- 57:23augmentations uh or the pre-text tasks
- 57:26which you give and it is able to still
- 57:28reconstruct the original images from
- 57:30these outputs then you can say your uh
- 57:33image is quite good uh your models is
- 57:36quite good and generally in the strong
- 57:38augmentation cases CNN models are bound
- 57:41to fail. So just to show you the example
- 57:44that the same images where you see a lot
- 57:46of haziness and other images where
- 57:49simply distorting or maybe rotating uh
- 57:52some portion of the cases images the uh
- 57:56in the second case if your model is able
- 57:58to get the uh context well like as a
- 58:02human being from the second image we
- 58:04will still be able to tell okay there's
- 58:06a park uh there are some uh in the
- 58:08surrounding area there might be a kind
- 58:10of few places where a built-up area
- 58:13might be there. But in the CNN based
- 58:15model such kind of things is not
- 58:17possible. It will fail. But VIT or
- 58:20vision transformation models basically
- 58:21will also will overcome all these strong
- 58:24augmentation based method uh uh strong
- 58:28augmentations in your data sets and
- 58:30accordingly it will able to give you a
- 58:32better output uh results. So let's focus
- 58:36uh some of the uh basic uh some of the
- 58:39basic uh models or most primitive models
- 58:42which we call it a mask image modeling
- 58:44cases. So as you uh as in this the
- 58:47process approach is very simple. So what
- 58:50you take you take an image randomly you
- 58:52create certain patches and remove the
- 58:55information which is present in those
- 58:56patches. So as you can see those blue
- 58:59boxes are the places where no
- 59:01information is there. So initially you
- 59:03have an image then what you do is you
- 59:06simply remove uh you create certain
- 59:08patches area and remove the information
- 59:10which is present in those cases and your
- 59:12model job is wherever these new the
- 59:16information is missing based on the
- 59:18remaining informations which is present
- 59:20in the image. So in this case at least
- 59:22we see around 95% of the pixel is still
- 59:26intact. So on the basis of that it
- 59:28should be able to get you the complete
- 59:31information of all those blue boxes as
- 59:34well. So width basically is going to uh
- 59:37provide you this uh uh uh this output
- 59:41through the mask imaging modeling. So
- 59:43the overall model how it works is that
- 59:45you just take certain input images
- 59:47randomly you create certain pi patches
- 59:50delete it send one pipeline where
- 59:52encoders will be put put and you just
- 59:54send the another pipeline where original
- 59:56image which you already had it and then
- 59:59you see that how much it is able to uh
- 1:00:02see that uh uh one whatever the partial
- 1:00:04information is there it is trying to
- 1:00:06reconstruct the whole images and then
- 1:00:08you see what was the original image and
- 1:00:10what was the reconstructed uh images. If
- 1:00:13there is a biasness, then you tell the
- 1:00:16reconstructed images uh uh to recreate
- 1:00:19the whole uh uh recreate and do some
- 1:00:22kind of correction in inside and
- 1:00:24accordingly the mask imaging modeling
- 1:00:26basically uh gets uh added uh basically
- 1:00:30gets improvement and accordingly it gets
- 1:00:32the global context and slowly it start
- 1:00:34learning and reconstructing the images.
- 1:00:37One uh basic uh understanding of all
- 1:00:40these mask image modeling approach you
- 1:00:42need to remember is that these all these
- 1:00:44requires a very huge data sets. So
- 1:00:47during the training phase you cannot
- 1:00:49just have a smaller image data sets and
- 1:00:51then you can create a uh global context
- 1:00:54because to understand the global context
- 1:00:56you need to have a variety of input
- 1:00:59images. So let's take a a run through of
- 1:01:02that image uh in the current examples.
- 1:01:05So as you can see one image has been
- 1:01:06broken down into uh 16 uh basically the
- 1:01:098x8 pixel of patches. Then from those
- 1:01:13patches each patches is being sent to a
- 1:01:16transformer. So basically each pixel
- 1:01:18let's say you are having a 256x 256
- 1:01:21pixels. So that will get divided into 8x
- 1:01:238. So each patches will be 16x 16.
- 1:01:27So you just create all those 16x6 image
- 1:01:29pixels. you just put them and flatten
- 1:01:31them into a single stream of the pixel
- 1:01:34side pixels and then you simply send it
- 1:01:37through some kind of uh uh basically
- 1:01:41standard CNN models like imageet reset
- 1:01:43and so on and accordingly you'll get a
- 1:01:46latent space which can be considered as
- 1:01:48a uh tokenizer. So it's basically
- 1:01:50creating a tokenizers. So and then
- 1:01:53finally you uh once that pixels and
- 1:01:55corresponding tokenizer is being sent
- 1:01:58then those tokenizers uh uh from those
- 1:02:00tokenizers tokens what you do you just
- 1:02:02remove certain uh certain informations
- 1:02:05and then you recreate the whole images.
- 1:02:08So as you can see from the mask input
- 1:02:11images once it gets an uh tokens. So
- 1:02:15these tokens are uh a kind of uh image
- 1:02:18based to uh token generator systems are
- 1:02:20there. So that what it will do is it
- 1:02:22will create a kind of uh a kind of
- 1:02:25generic information which may be present
- 1:02:27in in each of those patches. Then from
- 1:02:30those patches what is being done is
- 1:02:31simply you remove uh certain uh
- 1:02:34information. So as you can see there are
- 1:02:36two path data path gets created. One is
- 1:02:40basically taking the original image and
- 1:02:41second one is removing certain patches
- 1:02:45and then you realigning all those
- 1:02:46patches uh and send it to an encoder and
- 1:02:50then you say what are their positions in
- 1:02:52the original images. So once the
- 1:02:54corresponding position is being
- 1:02:57identified that position is being uh
- 1:03:00compared uh or is being put uh into the
- 1:03:04whole structures of the input data sets
- 1:03:07and finally it is being sent to a
- 1:03:09decoder to reconstruct the images. So
- 1:03:11you have an original uh uh mask uh image
- 1:03:15in uh basically tokenizers and second
- 1:03:18one is the reconstructed model which
- 1:03:19gives you the uh regenerated part. So
- 1:03:22there are certain parts where position
- 1:03:24has already been correctly predicted and
- 1:03:27then uh you reconstruct it. Then what
- 1:03:29you do you simply look in these uh token
- 1:03:32space or you can say the latent space
- 1:03:34where you try to find out what's the uh
- 1:03:37loss and accordingly you create a uh a
- 1:03:41kind of uh uh uh learning process as
- 1:03:44accordingly the whole weight gets
- 1:03:45adjusted to encoder or decoder uh
- 1:03:48systems. Finally as as there's no
- 1:03:51improvement is going to be there in
- 1:03:53reconstructed and original images your
- 1:03:56systems will you'll say that you'll stop
- 1:03:58the u uh basically the uh further
- 1:04:01learning process. So these are few
- 1:04:03examples of the same images and the
- 1:04:05corresponding algorithms has been given
- 1:04:07this it is available in mask image
- 1:04:10modeling paper which is available on
- 1:04:12archive. So you can just get it from
- 1:04:14there then.
- 1:04:17So once you have that uh uh
- 1:04:19reconstructions so that that was the
- 1:04:21first uh uh basically reconstruction
- 1:04:23approaches then similarly we can have a
- 1:04:25contrasting approaches. So what you do
- 1:04:27you take an image do some kind of uh uh
- 1:04:30augmentations for example in this case
- 1:04:32the image has been rotated and as well
- 1:04:34as the color has been changed you took
- 1:04:37original images you reconstructed the
- 1:04:39whole original images. So in the
- 1:04:42previous case in in case of
- 1:04:44reconstruction approach you do mapping
- 1:04:46in latent space or after the tokenizer
- 1:04:49or each patches is there. So you do the
- 1:04:51measurement in the latent space there.
- 1:04:54But in case of uh contrastive uh image
- 1:04:58modeling approach which is again a mask
- 1:04:59image modeling approach you take
- 1:05:02original image itself and the job of the
- 1:05:05encoder is to re to reconstruct the
- 1:05:09whole image back the way it was there uh
- 1:05:13in case of uh uh uh the original images.
- 1:05:17So whatever the partial information
- 1:05:19which is available so just the
- 1:05:21reconstruction and uh positioning is
- 1:05:23being carried out and accordingly you
- 1:05:25just see okay what are the information
- 1:05:27which is able to and how well it is able
- 1:05:28to map it. So accordingly the mask image
- 1:05:31modeling approach basically gets uh
- 1:05:32created. So this is just an algorithm
- 1:05:34for that. So the way we are having the
- 1:05:37in the contrastive uh mask imaging uh
- 1:05:40method we can also do lot of other
- 1:05:43augmentations. You can do some kind of
- 1:05:45rotations. You can do flipping. Color
- 1:05:47may be shifted or uh you can do some
- 1:05:50kind of glossial uh blurring. You can
- 1:05:52also do resizing. You can also do some
- 1:05:55kind of uh uh sun colorizations or the
- 1:05:59little bit of cloud simulations and so
- 1:06:01on. So all these different augmented
- 1:06:04input can go through this channel. So
- 1:06:07it's strong. You just create an uh
- 1:06:09augmentation. So any one of these
- 1:06:11augmentations can be applied and can be
- 1:06:14sent to an uh strong augmented output
- 1:06:17and then the whole chain can be
- 1:06:19recreated. So uh in case of mask image
- 1:06:22modeling you just try to remove certain
- 1:06:24sections uh using uh one of the
- 1:06:27approach. So you can use any other
- 1:06:29approach in the uh uh other uh remote
- 1:06:33sensing specific uh augmentations like
- 1:06:36solarizations is uh is a remote sensing
- 1:06:40specific algorithms where you can just
- 1:06:42say okay if data set is in one band how
- 1:06:44it will look into the second bands and
- 1:06:46so on. Similarly if there's a spectral
- 1:06:48shift so if there is a slight changes in
- 1:06:51uh let's say green band from 5.52
- 1:06:54nanometers to maybe 56.57 nanometers. So
- 1:06:57what can be the impact in overall
- 1:06:59images. So these kind of additional
- 1:07:01augmentations methods can be added to
- 1:07:04generate the uh uh augmented uh pipeline
- 1:07:08which will then you can use it the mask
- 1:07:10image modeling for the uh reconstruction
- 1:07:13of the images. So these are a few
- 1:07:15example and flow of these uh data
- 1:07:19augmentations and use of the vision
- 1:07:21transformation models especially in the
- 1:07:23remote sensing cases. So you take the
- 1:07:26strong augmentations you have any strong
- 1:07:28augmentations you can have any one of
- 1:07:30those which has been listed like uh
- 1:07:33strong jittering gshian spectral row and
- 1:07:36so on. So those kind of masking
- 1:07:39tokenization. So those kind of things
- 1:07:40you can add it and accordingly you can
- 1:07:43use transformer encoded and use the
- 1:07:46original image which is present so that
- 1:07:48the corresponding reconstructions can be
- 1:07:51used and this is the major uh I should
- 1:07:54say approach which is being used by a
- 1:07:57lot of uh larger uh I should say
- 1:08:00foundation models. So there has been a
- 1:08:03lot of I should say little uh a kind of
- 1:08:06uh uh sampling strategy uh or the
- 1:08:10augmentation strategies training
- 1:08:12strategies but fundamentally all of them
- 1:08:15actually follows the similar kind of uh
- 1:08:18pipeline.
- 1:08:20So in so basically uh so these these are
- 1:08:22basically the cases where three or four
- 1:08:25uh generally vision based transformer
- 1:08:27are basically on two image or a three
- 1:08:30band images but as we are aware that in
- 1:08:33case of remote sensing cases we are not
- 1:08:35going to have three bands but we'll be
- 1:08:37having a multiple band so more than 12
- 1:08:39bands is possible uh in some of the MSS
- 1:08:42data and in case of hyperspectral you
- 1:08:44may be having a bands in hundreds so
- 1:08:47that is also possible. So the uh these
- 1:08:51uh uh basically the uh all the vision
- 1:08:54based transformer uh based foundation
- 1:08:57model which is being used in the
- 1:08:58geospatial domains. So they have to
- 1:09:01adapt this multi- channelannel data sets
- 1:09:04of different uh uh uh I should say range
- 1:09:07of uh uh spectral band and uh it has to
- 1:09:12and it it also has a certain positional
- 1:09:15informations. So provisional information
- 1:09:17is quite important. For example, let's
- 1:09:19say if you go to Himalayas at glacial
- 1:09:24cases. So if you take multiple photos or
- 1:09:27multiple photograph of the same Himalaya
- 1:09:30at multiple times. So you are going to
- 1:09:33get or you will see mostly theis. So if
- 1:09:36I or the or I should say snow uh will be
- 1:09:40there or some kind of ice or snow may be
- 1:09:42observed. So in that case what will
- 1:09:45happen is always that area will be a
- 1:09:47kind of white colors. So if we if you
- 1:09:50know that these white colors are
- 1:09:52basically located at certain latitude
- 1:09:54and longitude. So we did not even have
- 1:09:57to look into the input images just
- 1:09:59knowing about the positions latitude and
- 1:10:02longitude and corresponding height. So
- 1:10:04we'll have a kind of pre uh training
- 1:10:07information of that areas. So these kind
- 1:10:10of additional fourdimensional or
- 1:10:12positional informations informations can
- 1:10:15be added or extracted along with the
- 1:10:18multi- uh I should say uh multiband
- 1:10:22spectral informations which becomes a
- 1:10:25kind of uh major hallmark of all the
- 1:10:28geospatial based uh foundation models.
- 1:10:31So just to give you an uh basic idea
- 1:10:34what exactly uh how it is being used. So
- 1:10:37one of the uh models which is being used
- 1:10:40is the clay models. Uh it is uh kind of
- 1:10:43land land cover map models. The basic uh
- 1:10:46foundation model uh in in this case what
- 1:10:48it does is it takes the input data of
- 1:10:51any band. So you can give this uh input
- 1:10:55to uh input to this model maybe of 250 m
- 1:10:58resolutions, 24 m resolution, lancet
- 1:11:01imagery or I mean you can give modest
- 1:11:04lancet or any kind of imagery and uh it
- 1:11:08is independent of the input bands. But
- 1:11:12once you give those inputs and
- 1:11:14corresponding uh band information is
- 1:11:16being fed then what happens is this will
- 1:11:19give you an correct classified image map
- 1:11:23uh which can be created from that image.
- 1:11:25So it means like it frees uh basically
- 1:11:27it lets you uh get uh I should say it is
- 1:11:32not dependent upon the input spectral
- 1:11:34band details. It is not dependent upon
- 1:11:37the uh the kind of input resolutions or
- 1:11:41the uh or the spatial resolutions
- 1:11:43spectral resolutions. It is independent
- 1:11:45of that. So how it is able to do this?
- 1:11:48This is basically there are two layer of
- 1:11:50uh
- 1:11:52basically adaptation which is being done
- 1:11:55apart from the part which we have
- 1:11:57discussed. So first one is that the
- 1:12:00reconstructions of the I should say the
- 1:12:03spectral informations. So let's say you
- 1:12:06are having a four band list three
- 1:12:08images. So what it will do is it will
- 1:12:10take a list three images but it will try
- 1:12:13to use those list three central bands
- 1:12:16and will try to construct a kind of
- 1:12:19intermediate continuous band data values
- 1:12:22to in a 128 channel. So it means like
- 1:12:25your uh you can think of that this is
- 1:12:27basically doing a uh uh linear
- 1:12:30transformations from four band into 128
- 1:12:34uh band informations or band level
- 1:12:36informations. So that weight it it
- 1:12:39doesn't gives you the exact uh I should
- 1:12:42say band information but the weight of
- 1:12:44each of those 128 channels which which
- 1:12:47is being fixed uh as an intermediate uh
- 1:12:50wave uh channel uh dimensions. So that
- 1:12:54basically gets uh uh stored and that
- 1:12:58weight is being used as an input to MEA
- 1:13:02reconstruction process. So as we have
- 1:13:03seen that MEA cases where you take an
- 1:13:06image break down into multiple segments
- 1:13:08and do some kind of random
- 1:13:10rearrangement. So your model's job is to
- 1:13:12reconstruct the whole uh uh reconstruct
- 1:13:16the whole images based on partial
- 1:13:18information. So at the middle of this
- 1:13:21whole layer what you see is basically
- 1:13:22the same B uh MEA reconstruction
- 1:13:25approach. Only thing is like instead of
- 1:13:27input images you also add the locations
- 1:13:31uh location embedding and as well as the
- 1:13:33latitude or longitude based time
- 1:13:36embedding informations and then you pass
- 1:13:38on to uh uh to the MEA reconstruction
- 1:13:42channels where you have encoder and
- 1:13:44decoder two channels. So this encoder
- 1:13:47and decoder channel is uh decoder
- 1:13:49channel is only active during the
- 1:13:50learning phases. In case of uh uh
- 1:13:54prediction phases this decoder is
- 1:13:56discarded. So you directly take the
- 1:13:58latent space which is present in the
- 1:14:00dimension D or the purple color which
- 1:14:02you see in solid D and these D uh latent
- 1:14:06space are then being used for different
- 1:14:10uh downstream classes uh uh basically
- 1:14:13cases like classifications and so on.
- 1:14:16So there has been a multiple development
- 1:14:18apart from MEA uh recently till recently
- 1:14:22there has been a kind of video mass
- 1:14:23encoder and so on. So these are well
- 1:14:25published uh papers which has been put
- 1:14:28into CBPR and other uh well uh
- 1:14:32conference uh organizations. So what
- 1:14:36happens is this same thing whatever we
- 1:14:38have done for the image classifications
- 1:14:40which is the clay is d is doing the same
- 1:14:43kind of task can also be done in case of
- 1:14:45numerical weather prediction model. So
- 1:14:47difference between the classifications
- 1:14:50and numerical weather prediction model
- 1:14:52is that it is time function. So in the
- 1:14:55classification cases that is time
- 1:14:57independent. So whatever image you give
- 1:14:59you simply get the corresponding output.
- 1:15:01But in case of uh weather predictions
- 1:15:04and so on. So you may be giving an
- 1:15:06instance maybe let's say in the morning
- 1:15:076:00 the different atmospheric condition
- 1:15:11uh conditions. So your model job is
- 1:15:13basically to give you the output
- 1:15:15for a future maybe 3 days 5 days 7 days
- 1:15:19predictions is there. So means it has to
- 1:15:21understand the underlying evolutions of
- 1:15:24uh different physical processes and then
- 1:15:28it should give you the output. So that
- 1:15:31is being done uh using the uh one of the
- 1:15:34models which is being developed by NASA
- 1:15:37and uh IBM that's called Priti. So Priti
- 1:15:41basically has two different uh P3 WXC uh
- 1:15:44that's the model exact name which is uh
- 1:15:47being provided by them. So what they
- 1:15:49have done is they have taken around uh
- 1:15:51uh 20 years of the whole numerical
- 1:15:54predicted data sets from uh WRF uh
- 1:15:58output data which is available from ECMF
- 1:16:01uh W systems and then uh uh they have
- 1:16:06they have used all those data sets to
- 1:16:08understand the different underlying
- 1:16:10weather pattern in multi-year levels and
- 1:16:13each day information is also been uh is
- 1:16:16generated uh and is that whole model
- 1:16:19also again works on the same MEA cases.
- 1:16:22Only thing is that there are two level
- 1:16:24of MEA. So instead of having a single
- 1:16:27initial patch and deleting, you may be
- 1:16:29having a whole group of patch which may
- 1:16:30be created and that's how the uh the
- 1:16:34models the foundation models which is
- 1:16:35being used for generating the uh weather
- 1:16:38predictions and so on. So these are the
- 1:16:40few two basic or I should say the time
- 1:16:43domain as well as uh the spatial domain
- 1:16:46weather prediction models uh which is
- 1:16:49currently available. Apart from that
- 1:16:51there are also uh some other models like
- 1:16:53SATM spectra ch uh GPT. So it is also a
- 1:16:58kind of temporal vision transform
- 1:16:59models. Then you have a scala and uh
- 1:17:02scale me. So these are basically small
- 1:17:04uh improvement taking into
- 1:17:06multi-reolution data set. For example,
- 1:17:08scale is able to take any data sets
- 1:17:10input which is from 30 cm to 100 m uh to
- 1:17:1410 m resolutions and so on. So these are
- 1:17:17the other few development which has
- 1:17:18happened in uh foundational models. Then
- 1:17:22there's a DOA model which is again is
- 1:17:24being used for generating or combining
- 1:17:27the data sets from the radar to optical.
- 1:17:29So radar it it the doa model has seen
- 1:17:32the sentinel one that is a radar data
- 1:17:34set 2 and so on and clay I think I have
- 1:17:38already uh discussed much in detail. So
- 1:17:41this is just a few summary of the models
- 1:17:44which is available in the geospatial
- 1:17:46domain and you can use uh any one of
- 1:17:48them and they are having their own uh
- 1:17:50advantages and uh and most of them can
- 1:17:53be used for their lot of downstream
- 1:17:55task. So once these model data sets are
- 1:17:57publicly available so as an end user you
- 1:18:00just need to take the input and create
- 1:18:01certain downstream tasks like
- 1:18:03classification change detections and so
- 1:18:05on that can be created. So with this uh
- 1:18:08we are just ending today's sessions and
- 1:18:10we'll be happy to take certain questions
- 1:18:12as a part of these sessions.
- 1:24:24Okay. So I think uh we can take certain
- 1:24:27questions. Uh so I think uh from the
- 1:24:31beginning there are few questions
- 1:24:36which is related to registrations and
- 1:24:37other quizzes. I think you might be
- 1:24:39knowing already that process. So I'm not
- 1:24:41going to talk all those. I'm just
- 1:24:43focusing on the current lectures. So
- 1:24:46first questions I'll take from Sans
- 1:24:48Sharma. So it's basically that how does
- 1:24:50the CL clay model basically integrate
- 1:24:53spatial spectral and temporal
- 1:24:55informations in its architecture to
- 1:24:58improve the remote sensing images. So
- 1:25:00first of all uh uh I should say that the
- 1:25:03clay model what it does is uh let's go
- 1:25:06back to uh maybe
- 1:25:11I think you might have seen uh
- 1:25:18the architecture which is
- 1:25:24model.
- 1:25:27Yeah.
- 1:25:30So if you see the uh if you observe the
- 1:25:33clay models, so one is called wave
- 1:25:36transformer. Okay. So that's the first
- 1:25:38stage and there are two stage wave
- 1:25:40transformer transformer which clay model
- 1:25:42uses. So in this case what happens is
- 1:25:45that uh the first wave transformer model
- 1:25:49which you are seeing that takes the
- 1:25:52center wavelength of the input data set.
- 1:25:54So let's say you are giving uh list four
- 1:25:56images. So list four images having a
- 1:25:58four bands. So red, green, uh and the
- 1:26:01corresponding swear and I bands. So what
- 1:26:04you have to give you have to just give
- 1:26:06the central wavelength as an input that
- 1:26:08has to be provided as an input. So the
- 1:26:11first stage uh uh transformer which you
- 1:26:13are seeing as a wave transformer. What
- 1:26:15it does is it takes those four band
- 1:26:18input images and considers them as a
- 1:26:23four distinct input points at those uh
- 1:26:26locations. So you may be having a 0455
- 1:26:3075 and let's say 1 uh uh uh 1.13 I mean
- 1:26:35that sorry 11.3 and so on. So these kind
- 1:26:38of different uh band uh sequencing is
- 1:26:41being put in. So using those four or
- 1:26:43five bands uh information or central
- 1:26:46wavelength what it does is it tries to
- 1:26:48create a 128 projected uh uh linear
- 1:26:52transformations of that data sets. Uh so
- 1:26:55what happens is that only using the four
- 1:26:57bands it knows out of 128 which it's
- 1:27:00expecting as an as a part of its own uh
- 1:27:03uh embedding spaces it will be able to
- 1:27:06see okay how much weightage has to be
- 1:27:08given to those 128 bands which
- 1:27:11internally it just maintains. So those
- 1:27:13are basically the uh internal embedding
- 1:27:15band which is being kept in. So
- 1:27:18accordingly what happens is that those
- 1:27:20weights are then being used uh different
- 1:27:22waiting system is being used. So that's
- 1:27:24how it is able to understand the
- 1:27:27different spectral information. So if
- 1:27:29some of the cases let's say and that
- 1:27:30that weight is getting again uh uh
- 1:27:33generated through 4year transformations
- 1:27:36uh model approach. So using the 4year
- 1:27:38transformation coefficients is being
- 1:27:40used to generate the weight. So if
- 1:27:42you're having a four bands so there will
- 1:27:44be certain weights which usually
- 1:27:46consider the four uh bands to give or
- 1:27:48arrive the weightage to all the 128 uh
- 1:27:52different uh embedding space. Similarly
- 1:27:54you may be having 12 bands. So
- 1:27:56accordingly those 12 bands will be used
- 1:27:58to generate the uh corresponding uh uh
- 1:28:02spectral embedding spaces that gets
- 1:28:03created. So that's how the spectral band
- 1:28:06is getting uh uh extracted. the temporal
- 1:28:09information it doesn't directly encodes
- 1:28:12it. So instead of that what it does is
- 1:28:14it simply does the normalizations. So
- 1:28:16during the modeling or the I should say
- 1:28:19the fine-tuning process what you need to
- 1:28:21do is you need to take the input and
- 1:28:23then you have to do a jet scaling of all
- 1:28:26the data sets. So uh the temporal
- 1:28:29information at such uh it it uh first or
- 1:28:32the image level simply does the kind of
- 1:28:34normalizations and the temp temporal
- 1:28:37information basically does the s cosine
- 1:28:39transformation. So that's that's how it
- 1:28:41is uh is being added to the embedding
- 1:28:44space in the next uh uh basically
- 1:28:48next uh uh me stage. So that's how it is
- 1:28:51going to give you the uh it it takes
- 1:28:54care of the overall multisspectral and
- 1:28:57as well as the uh temporal informations
- 1:29:00in in that sense. Then we say uh then
- 1:29:03Virra is asking how what does the govt
- 1:29:06learns and machine learning use in
- 1:29:08learning patterns to deep the image
- 1:29:10pixels. So basically I think probably uh
- 1:29:13the question is might be reframed that
- 1:29:16uh what does the goit learns and what
- 1:29:19may be the learning in other approach
- 1:29:21like CNN and so on. So CNN and other
- 1:29:24approaches may be looking for a kind of
- 1:29:27local context and pixel to pixel never
- 1:29:30ne never pixels to ne like some pixel
- 1:29:33are there so what is the relationship
- 1:29:34with the neighboring pixels. So
- 1:29:36accordingly what it will do it will get
- 1:29:37a kind of context. Okay, for example,
- 1:29:40let's say if you're having a single road
- 1:29:43pixels and it is surrounded by a tree,
- 1:29:46so you may be having a one pixel with
- 1:29:48the road and neighboring to that you may
- 1:29:50be having an agriculture pixels and so
- 1:29:52on. So you will be you'll be having that
- 1:29:53kind of informations. So when it
- 1:29:56understands this relationship okay so in
- 1:29:58most of the places let's say in the
- 1:29:59whole world around 80% of the areas it
- 1:30:02sees that at the just next to the road
- 1:30:05pixels of the black uh pixels there is a
- 1:30:08yellow color uh I should say the
- 1:30:10agriculture farmland is there. So
- 1:30:12accordingly what it will do is it will
- 1:30:14try to uh understand and tell that
- 1:30:18wherever black pixels are there. So that
- 1:30:20has to be a kind of road classes and it
- 1:30:23has to have a kind of continuity because
- 1:30:25in between suppose there is a mask of
- 1:30:27the trees but since it knows that this
- 1:30:30this pixel belongs to a uh road pixels.
- 1:30:34So in those cases even though there
- 1:30:35might be a overhead uh or let's say
- 1:30:38occlusion of the trees on the top of the
- 1:30:40road but still it will try to
- 1:30:42reconstructed. So it means that's the
- 1:30:44convolutional uh systems works but that
- 1:30:47works only in a localized uh field I
- 1:30:50should say or within the specific
- 1:30:52windows where in case of uh geoid it has
- 1:30:55a overall global context. So within the
- 1:30:57whole image this uh the systems learns
- 1:31:02what is the position of these pixels or
- 1:31:04what are the uh group of pixels which is
- 1:31:06a patch where where it is exactly
- 1:31:09located what kind of information is
- 1:31:12present in that patch and who all are
- 1:31:15the neighbor to this patch and so on. So
- 1:31:18correspondingly what happens is the
- 1:31:20patch is able to get local global and as
- 1:31:23well as internal context as a whole uh
- 1:31:27from the pixel information and
- 1:31:28accordingly it will be able to
- 1:31:30reconstruct the uh the whole uh uh I
- 1:31:33should say informations or the semantic
- 1:31:35information which is present. So this is
- 1:31:37how it takes uh that uh then there's a
- 1:31:42question of for disaster forecasting in
- 1:31:44NWB. Yes, definitely and that is the
- 1:31:46basic advantage which is in compared to
- 1:31:49uh weather forecast WRF models and so
- 1:31:52on. So in the forecast forecast models
- 1:31:54if you see that uh to do a predictions
- 1:31:57uh for let's say next 6 hours
- 1:32:01predictions even in the decent good uh
- 1:32:04machines you may be requiring hours of
- 1:32:07uh data processing even in this good
- 1:32:11number of HPC environment and in uh but
- 1:32:14in case of uh uh basically the
- 1:32:16foundation models like P3 WXCX you'll be
- 1:32:19able to get the output in seconds. So
- 1:32:21advantage of that is that you have a
- 1:32:23higher speed up. Uh then uh it also
- 1:32:27understands the overall natures or uh uh
- 1:32:30basically the prediction area and
- 1:32:32corresponding context which has happened
- 1:32:34in last 20 years uh 2020 years. So it
- 1:32:38means like it is having a it is going to
- 1:32:40give you a predictions not only from the
- 1:32:42current u I should say uh the
- 1:32:45atmospheric condition but it also have
- 1:32:47an understanding of what was the overall
- 1:32:50average weather patterns in each of the
- 1:32:53global uh glo uh as a whole uh within a
- 1:32:57globe. So bringing the combining those
- 1:33:00two informations and faster predictions
- 1:33:02it will be able to give you a better uh
- 1:33:05result. For example, if you just combine
- 1:33:06them with case of disaster areas. So
- 1:33:09what may happen is like if you if you
- 1:33:11just combine those uh disaster special
- 1:33:13embedding where uh large number of
- 1:33:15disaster happens for example landslides
- 1:33:18which is mostly found in the I should
- 1:33:21say hilly regions or maybe a cloud bus
- 1:33:24which is again in the hilly region. So
- 1:33:26if you just do a fine-tuning and attach
- 1:33:28additional uh elevation informations to
- 1:33:31these uh foundation models then it will
- 1:33:34not only give you the heavy predictions
- 1:33:36precipitations area but rather than it
- 1:33:39will also will be able to tell you that
- 1:33:41there is also likely to have some kind
- 1:33:43of landslide and so on. So it is
- 1:33:45definitely going to be much helpful in
- 1:33:48uh in in that sense and that's that's
- 1:33:50how it is going beyond the uh numerical
- 1:33:54prediction which just gives you the
- 1:33:55atmospheric condition uh atmospheric
- 1:33:57conditions from the different numerical
- 1:34:00weather prediction models but these uh
- 1:34:02AI models gives you beyond that okay
- 1:34:04let's say if there's a heavy rain what
- 1:34:06is going to the impact there will be
- 1:34:08flooding there will be no flooding what
- 1:34:10area will get inundated and so on so all
- 1:34:13these things can be possible possible
- 1:34:15through the numerical uh uh weather
- 1:34:17prediction uh foundation models that can
- 1:34:20be used. Okay. So quizzes I think it
- 1:34:23will be available uh on the portal. So
- 1:34:25you can look into that. That's the
- 1:34:26questions to Kala. Yeah. One animesh has
- 1:34:31also asked one questions related is that
- 1:34:34how does the clay dynamic embedding
- 1:34:36block process the variable number of
- 1:34:38spectral bands compared to the standard
- 1:34:41uh vision transformation. So wave uh
- 1:34:44models if you take it uh as I was
- 1:34:47telling so just do a kind of you can you
- 1:34:50can think of uh in very layman's
- 1:34:52language not exactly equivalent but just
- 1:34:55think of that you're having a
- 1:34:56multisspectral four band and you are
- 1:34:58trying to reconstruct the hyperspectral
- 1:35:01bands in 128 bands. So if you're having
- 1:35:04a hyperspectral 128 bands and uh you
- 1:35:08need to reconstruct it from the four
- 1:35:09bands then what you'll do you'll simply
- 1:35:11assign the certain weightage to those uh
- 1:35:15uh 128 bands in such a way that
- 1:35:18summation or or the uh I should say
- 1:35:20linear summation of those bands is
- 1:35:23equivalent to those four bands or five
- 1:35:24bands whatever it is available to you.
- 1:35:27So that's how it is able to uh create
- 1:35:30it. uh if you're interested in more math
- 1:35:32probably you can send me a mail I can
- 1:35:34just send you the uh refer you can refer
- 1:35:37the paper and there are also an publicly
- 1:35:40available uh I should say uh uh
- 1:35:42repository is also there where clay
- 1:35:44models uh source code is available so
- 1:35:47you can easily use and you can look more
- 1:35:50in detail uh from the github repository
- 1:35:53so this is also available on the hugging
- 1:35:55space there also uh you can do that uh
- 1:35:59how does the juda processing contributes
- 1:36:01to national development projects. So as
- 1:36:03any other technologies
- 1:36:05uh the uses applications are quite wide.
- 1:36:09It depends upon the uh the adaptation of
- 1:36:12these tools uh in an decision-m process.
- 1:36:15It will be helpful in those uh similar
- 1:36:17cases. Uh okay. So I think I have
- 1:36:21answered already that how does the NWP
- 1:36:24works for disaster forecasting. So I
- 1:36:26think with based and what their
- 1:36:28advantage so I think uh Axa Singsh I
- 1:36:31think I already have answered that and
- 1:36:34uh how does the model handle spatial and
- 1:36:36temporal irregularity for missing data
- 1:36:39across so that's what I was telling that
- 1:36:41as a part of MEA the model's objective
- 1:36:45is that whatever partial information it
- 1:36:47is having on the basis of that it has to
- 1:36:50reconstruct the whole condition. So
- 1:36:52there has been one paper and on priti
- 1:36:55and they are claiming that if even if
- 1:36:57you're having only five percent of the
- 1:37:00area where you may be having an
- 1:37:02atmospheric uh informations from that 5%
- 1:37:06of the informations the model is able to
- 1:37:08construct the remaining 95 and the
- 1:37:11accuracy of those reconstructed 95 area
- 1:37:14has 95 portion uh of the remaining area
- 1:37:18which was missing has been found to be
- 1:37:20around 85 to 90 uh% accuracy level. So
- 1:37:25this this is this this is a kind of uh
- 1:37:27regeneration uh data sets which is which
- 1:37:31is possible through uh such models. So I
- 1:37:34think uh you if you want if you're
- 1:37:36interested you can just have uh the
- 1:37:38better understanding by looking into
- 1:37:40codes and the other models uses uh
- 1:37:43especially the prit and the clay model
- 1:37:45all those are publicly available and
- 1:37:47available on the GitHub so you can use
- 1:37:49them and uh if you have any further
- 1:37:52questions just write me uh and I'll be
- 1:37:55able to answer all of them.
- 1:38:03Uh yeah there is one question from Niha
- 1:38:05how is SSL different from supervised
- 1:38:08learning. So as I have said that in case
- 1:38:11of supervised learning you have to
- 1:38:12recreate the labels. So as an uh uh
- 1:38:16person you have to label it create a
- 1:38:18mask and then you have to provide it to
- 1:38:20the learning algorithm where SSL uh what
- 1:38:24it does it instead of having a uniform
- 1:38:27labels it just randomly creates the uh
- 1:38:30certain augmentation and distortion in
- 1:38:32the images and try to reconstruct the
- 1:38:34whole images. So it's it's a kind of
- 1:38:36self-arning. So it's a kind of you can
- 1:38:38think of that one kid is sitting and
- 1:38:40he's trying to learn uh how to put one
- 1:38:43block on the top of the others. So you
- 1:38:45just pass on a certain blocks that
- 1:38:48person that uh kid may be putting it one
- 1:38:51block on the top maybe initially
- 1:38:53starting at the edges then slowly it
- 1:38:55will start putting at the center and
- 1:38:57then later on it learns that it has to
- 1:39:00put that another block exactly aligned
- 1:39:03with the central of uh gravity of each
- 1:39:06of those blocks. So that's how the uh
- 1:39:08models also learns. It just does certain
- 1:39:11distortion maybe rotation maybe deleting
- 1:39:13certain process and trying to recreate
- 1:39:15it again after rotation again
- 1:39:18reconstruct back the original
- 1:39:19orientation of the images and so on. So
- 1:39:21those kind of things is being done and
- 1:39:23that's how the uh available. So I think
- 1:39:26I have already answered about the priti
- 1:39:28wxc models. Yes, that is publicly
- 1:39:31available. There's another model from
- 1:39:33Google which is called graphcast.
- 1:39:35uh now cast. So those kind of models uh
- 1:39:38can also be used and those are also
- 1:39:40again uh AI based uh NWP models which
- 1:39:44can be used uh for uh different uh use
- 1:39:49cases.
- 1:39:50Uh I think rest uh are uh so I think
- 1:39:54probably more of these methods or
- 1:39:57questions which you have already asked.
- 1:39:59Yeah. 1 C1 is asking that what is the
- 1:40:01glossian blur and how can we understand
- 1:40:05what is the what the data is in the
- 1:40:07fourth dimension. So when we say four
- 1:40:09dimensions generally means having the
- 1:40:12additional spectral information spatial
- 1:40:14informations uh especially the latitude
- 1:40:17longitude and other band which is not
- 1:40:21being used in normal vision based
- 1:40:23systems. So all the other vision based
- 1:40:25system just uses three band images RGB
- 1:40:27images which generally get captured
- 1:40:29through normal camera but in case of
- 1:40:32geospatial domain we go beyond that. So
- 1:40:34you may be having a 12 band spectral
- 1:40:36band data sets. So all of the bands you
- 1:40:39cannot view at the same time. You can
- 1:40:42also include certain bands for example
- 1:40:44the terrain informations can be included
- 1:40:47during the training process. So the
- 1:40:49model also have an understanding and
- 1:40:51context of uh I should say terrain based
- 1:40:54information or positional informations.
- 1:40:56So these kind of additional uh
- 1:40:58informations you can put all all the all
- 1:41:00of them together in fourth uh uh
- 1:41:03dimensions.
- 1:41:06Yes. So Google Earth Engine is also
- 1:41:07having one uh foundation model
- 1:41:09especially the embedding space. I think
- 1:41:11it is having a 64 band uh currently
- 1:41:14currently I'm not able to recall its
- 1:41:16names but uh it is also having an
- 1:41:19foundation model outputs and the latent
- 1:41:21space uh data sets has been created and
- 1:41:24it is found to be quite uh uh I should
- 1:41:27say the accuracy is quite good which can
- 1:41:29be used for different classifications
- 1:41:31and segmentation task. So just have a
- 1:41:34look into Google Earth Engine and look
- 1:41:36for uh uh uh foundation models data
- 1:41:39sets. probably you'll get uh uh data uh
- 1:41:42I mean uh the data set which is
- 1:41:44available and create generated by Google
- 1:41:48okay how does we decide which weather
- 1:41:50model is best for particular reason so
- 1:41:53so the selection of any weather model is
- 1:41:57true or any kind of activity which we do
- 1:42:00is is always dependent upon the user's
- 1:42:03uh choices. So if you think certain
- 1:42:06models are good in your conditions you
- 1:42:08see there's a very good accuracy
- 1:42:09predictions or sometime you may be
- 1:42:12interested going beyond that maybe
- 1:42:14sensitivity analysis you can carry out
- 1:42:17so on the basis of that you take a
- 1:42:18decision which model is good so that
- 1:42:20decision has to be taken by the users uh
- 1:42:23and accordingly you can accept like in
- 1:42:25any other cases like WRF models or there
- 1:42:28are a lot of weather prediction models
- 1:42:29are there IMD is also providing the data
- 1:42:32sets but if you think that uh that model
- 1:42:34is good enough and it is giving you a
- 1:42:37very good understanding of what is going
- 1:42:38to happen. It is matches with the
- 1:42:40reality then you accept it otherwise you
- 1:42:42simply reject it. So same is also the
- 1:42:44case with any kind of other uh model uh
- 1:42:48as well. Uh yeah so then the one
- 1:42:50questions the use of latent layer in the
- 1:42:53clay model. So latent layer clay model
- 1:42:55as I was telling that when you when your
- 1:42:58model learns uh so during the uh
- 1:43:01learning process it understands the
- 1:43:03positioning or overall global context.
- 1:43:06So when I say global context what does
- 1:43:08it means? So it means that if given an
- 1:43:11input it is trying to understand okay
- 1:43:13what are the foreground what are the
- 1:43:14background what are the information
- 1:43:16which is present in very much in the
- 1:43:19near infrared uh regions what are the
- 1:43:22textural informations so these kind of
- 1:43:25uh the contextual is the texture is fine
- 1:43:28or granular. So all these information
- 1:43:31gets combined together in its latent
- 1:43:33space of uh 1024 bands. So you can think
- 1:43:37of that you are just providing a four
- 1:43:39band input images or maybe 12 band input
- 1:43:41images and finally you get 1024 uh bands
- 1:43:46output which has information not only in
- 1:43:49the spectral reason but it also in the
- 1:43:51spatial relationships and so on. So
- 1:43:54accordingly you get a very large latent
- 1:43:56space. So accordingly and that's what
- 1:43:58you can use it for different kind of uh
- 1:44:00classifications categorizations which
- 1:44:02can be used and uh that's how uh we can
- 1:44:06use it for different uh
- 1:44:09cases I think wit and other uh I think I
- 1:44:12have already talked about so I think I
- 1:44:14just took all the questions so probably
- 1:44:17almost I was able to answer all of those
- 1:44:20uh questions so uh hersel I think we
- 1:44:25today we didn't discuss anything related
- 1:44:27to vector raers. So probably you you
- 1:44:31need to brush up and have a look what
- 1:44:33exactly we discussed. Anyway, so thank
- 1:44:36you for uh uh joining these sessions.
- 1:44:39Have a nice day. Bye-bye.
About this transcript
This page contains the full transcript of Foundation Models for Geodata Processing by Dr. Ashutosh Kumar Jha by IIRS ISRO Digital Learning Programme, generated from the public captions YouTube serves with the video. The transcript has 12,906 words across 1,805 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.