YouTube2Text

Foundation Models for Geodata Processing by Dr. Ashutosh Kumar Jha — Transcript

by IIRS ISRO Digital Learning Programme · 12,906 words · 1,805 segments · language en · Watch on YouTube

Full transcript

  1. 20:34Okay, good evening. Uh,
  2. 20:37today we'll be having a discussion on
  3. 20:40trying to understand the foundation
  4. 20:41models for Joda processing. So these are
  5. 20:44basically
  6. 20:45uh certain recent development which has
  7. 20:48happened in uh geospatial domain. So
  8. 20:51we'll be talking about uh those
  9. 20:53development. So we will not be going
  10. 20:54much in details just trying to
  11. 20:56understand what is going on uh currently
  12. 20:59in uh especially in the foundation
  13. 21:02models which is related to uh G data
  14. 21:04processings.
  15. 21:06So basically uh we'll just browse
  16. 21:09through the different uh evolution of AI
  17. 21:11vision models that are being used
  18. 21:14currently and there are different stages
  19. 21:16which has led to uh certain development
  20. 21:19and that has uh resulted in a situations
  21. 21:22where now we are seeing that lot of
  22. 21:25vision AI vision based models are having
  23. 21:27an accuracies which is at par to the
  24. 21:30supervised level of uh accuracy or
  25. 21:33classifications techniques or may be
  26. 21:36related to the task which is a human can
  27. 21:38do that. So the current systems are at
  28. 21:40least uh approaching to that kind of uh
  29. 21:43uh performance. Then we'll see a certain
  30. 21:46uh basic understanding uh of uh which is
  31. 21:50needed to uh look into different kind of
  32. 21:52learning approaches which are being used
  33. 21:55by these uh different vision models and
  34. 21:59we'll also focus uh slightly more on uh
  35. 22:02CNN and as well as the vidbased models.
  36. 22:04So I will not be touching more detail in
  37. 22:06CNN but rather than on just uh giving
  38. 22:10details on visual uh vision uh level
  39. 22:12transformations and how it is uh
  40. 22:15different from the CNN and so on. Then
  41. 22:17I'll just extend that to a
  42. 22:18self-supervised learning approach which
  43. 22:20is being used and then we'll be having a
  44. 22:22few geospatial foundation models which
  45. 22:25will be look into that. So if you see uh
  46. 22:28in in in basically in the vision based
  47. 22:32systems which is mostly related to a
  48. 22:34neural based approach uh the development
  49. 22:36started quite early in 1958 and probably
  50. 22:40might be have seen the seifur data sets
  51. 22:43uh some of you might have seen that. So
  52. 22:45digital identifications were the first
  53. 22:47task which was being attempted and
  54. 22:49accordingly there was a lot of
  55. 22:50development has happened. So perceptron
  56. 22:52was basically a kind of computational
  57. 22:54models uh which is uh based on how a
  58. 22:59neuron u works. So it's basically
  59. 23:02mathematical mapping of that is being
  60. 23:04done in term of the weighted uh uh uh
  61. 23:08summations of the input and accordingly
  62. 23:10the the mapping has been done in term of
  63. 23:12nonlinear uh approach using the
  64. 23:14exponential function has been used. So
  65. 23:16that's the basic principle of
  66. 23:17perceptron. Then there has been a
  67. 23:20certain development uh which is related
  68. 23:22to CNN. So CNN basically doesn't works
  69. 23:25on a pixels or in in term of image we
  70. 23:28can say pixel level point data sets
  71. 23:31rather than it takes into the local
  72. 23:33context. So it basically takes a window
  73. 23:35around surrounding the central pixels
  74. 23:36and surrounding that there will be a few
  75. 23:38pixels and that will be applied to uh is
  76. 23:42being applied for CNN cases. So CNN has
  77. 23:44an advantage that it gets a kind of
  78. 23:47pixel informations along with the
  79. 23:49context and that's how the uh the
  80. 23:52classifications
  81. 23:54uh is being uh uh basically is being
  82. 23:57improved compared to pixel level
  83. 24:00perceptron based systems. Then we had uh
  84. 24:03recurrent neural networks or LSTM. So
  85. 24:05these are basically using the first
  86. 24:08level of output of uh one CNN or one
  87. 24:11neural network based systems giving the
  88. 24:13output and then feeding it to the next
  89. 24:16uh level. So generally it is being used
  90. 24:17in a temporal uh cases where in the CNN
  91. 24:21and uh perception related uh neural
  92. 24:24network you don't have any memory. So
  93. 24:26whatever has happened earlier in
  94. 24:28instances that is not being taken into
  95. 24:30consideration but LSS and RS RSTM
  96. 24:33basically takes into those uh into the
  97. 24:35considerations. Then there was a big
  98. 24:37leap in uh 2012 and the imageet
  99. 24:41classifications was been put by
  100. 24:45Google and that was a kind of remarkable
  101. 24:48achievement uh especially the using the
  102. 24:50deep neural network. So whatever
  103. 24:52convolutional neural network which was
  104. 24:54being used earlier that was restricted
  105. 24:56to maybe two or three layers but later
  106. 24:58on in the image net they have gone into
  107. 25:01quite deeper 18 20 layers of their depth
  108. 25:04has been there. So due to that what has
  109. 25:06happened is there has been a lot of
  110. 25:08contextual information also uh
  111. 25:11percolating in the started percolating
  112. 25:13at the upper layer of uh neural networks
  113. 25:17and that's how the uh each of the uh
  114. 25:20informations now have a kind of global
  115. 25:22context which is available in the images
  116. 25:25that gets carried away and accordingly
  117. 25:27the improvement has been there in case
  118. 25:30of uh uh the other uh vision level
  119. 25:33information there has been a tremendous
  120. 25:34jump in uh uh basically the accuracy
  121. 25:37level which is expecting in after the
  122. 25:40development of alexnet and imageet
  123. 25:42cases. Then there was a kind of semantic
  124. 25:45uh unit. So this is this was the outcome
  125. 25:48of the there was few observations which
  126. 25:51was being made after the DNN based uh
  127. 25:54model started developing. So what has
  128. 25:57been observed that uh the earlier
  129. 25:59advantage of increasing the depth of CNN
  130. 26:03that started losing I mean maybe after
  131. 26:0520 or 30 layers the the the observation
  132. 26:08was there that uh as you go deeper and
  133. 26:11deeper then that there's no extra
  134. 26:12information gets extracted and the upper
  135. 26:15layer basically gives you the kind of
  136. 26:16noise level informations. So there is
  137. 26:19there has not been any improvement. So
  138. 26:22later on the unit architecture which
  139. 26:24basically uh has a kind of uh new
  140. 26:27systems where encoder and decoder
  141. 26:29architectures gots added with uh uh
  142. 26:32extra layer where the semantic
  143. 26:34informations from the lower level uh uh
  144. 26:38networks has been passed on to the upper
  145. 26:40layer and that's how the improvement has
  146. 26:42happened uh with the deep neural network
  147. 26:45cases. Then there was another
  148. 26:47development which has happened uh and
  149. 26:49there's a paper which was being
  150. 26:51published by Google's uh team that was
  151. 26:54basically attention all you need. It's a
  152. 26:56very small uh title but it has a very
  153. 26:59big impact in uh overall vision systems.
  154. 27:03So what they have uh trying to do or
  155. 27:06tried to do under the uh uh vision based
  156. 27:10systems. So they started uh a kind of
  157. 27:13model where they say that uh instead of
  158. 27:15predicting and characterizing and
  159. 27:17classification let's do a prediction of
  160. 27:20the neighborhoods or you can say that
  161. 27:23the next points which is being possible.
  162. 27:26So generally in the NLP it is being used
  163. 27:28like if you just put it a few text and
  164. 27:31the correspondingly what may be the next
  165. 27:33text that get predicted. So same kind of
  166. 27:36mechanism also got uh uh uh basically
  167. 27:40carried into the vision based systems as
  168. 27:42well at later stages. In between there
  169. 27:45has been a kind of physics based
  170. 27:47simulations and other model which has
  171. 27:49also been people started trying to
  172. 27:52integrate. So instead of completely
  173. 27:54relying only on what data says and the
  174. 27:56people have started realizing or
  175. 27:58creating certain physics based uh uh
  176. 28:01network a kind of uh model simulation
  177. 28:04environment where physics based
  178. 28:06conservation of certain laws maybe the
  179. 28:09conservation of energy conservation of
  180. 28:11uh momentum all those things has been
  181. 28:13embedded and that those were basically
  182. 28:15specifically useful in the cases of
  183. 28:18geospatial domain. For example, in the
  184. 28:20weather pattern, it is uh the cloud
  185. 28:23movement is not sufficient. Cloud also
  186. 28:25if you take a cert lots of images of the
  187. 28:28cloud movement that is not sufficient to
  188. 28:31give idea about the weather pattern. You
  189. 28:33also need to know the certain physics or
  190. 28:35atmospheric physics related to the
  191. 28:37neural uh basically the numerical
  192. 28:39weather prediction models that needs to
  193. 28:41be integrated to have a uh better uh
  194. 28:45forecasting models. So these were
  195. 28:47basically started coming in in 2019.
  196. 28:51Then there was been uh uh vision based
  197. 28:53transformation which was basically from
  198. 28:562017 transformer now got into started
  199. 28:59getting into vision based
  200. 29:00transformations and again uh the similar
  201. 29:03paper was there. So you just said I mean
  202. 29:05in in that paper basically it says that
  203. 29:08you need only 16x6 pixels uh as an input
  204. 29:12to get any kind of vision task. So
  205. 29:14that's was their claim and that has been
  206. 29:17developed in 2020. Then slowly there was
  207. 29:20uh also been observed that during these
  208. 29:24times as we had a lot of uh data or I
  209. 29:28should say the data generation systems
  210. 29:30has also increased you have a lot of
  211. 29:31data coming from different sources uh
  212. 29:35social media and other places. So slowly
  213. 29:37what has happened is the the input rate
  214. 29:41of the data has outpaced the development
  215. 29:43of the uh basically uh you you can say
  216. 29:47the supervised labeling approaches. So
  217. 29:50earlier these systems were being built
  218. 29:52on the basis of labeling of the
  219. 29:53different uh sections in the images and
  220. 29:56that was being used in the supervised
  221. 29:57cases but as the data generation things
  222. 30:00has outpaced. So the impact of that was
  223. 30:03that people were not able to use the uh
  224. 30:06the higher rate of the data sets and in
  225. 30:08that scenarios there was a kind of uh
  226. 30:11self supervised vision based system has
  227. 30:14started coming. So what uh model has
  228. 30:17basically started uh has been developed
  229. 30:20uh with a with a objective that instead
  230. 30:23of uh trying to label it let's create
  231. 30:26certain things in between. So in between
  232. 30:29means there will be a certain latent
  233. 30:30spaces which informations can be created
  234. 30:33and on the basis of that you can uh do
  235. 30:36uh uh a kind of reconstruction of the
  236. 30:38whole scene and images. So we'll see few
  237. 30:40of those algorithms uh or the steps
  238. 30:43which is in the coming slides. Then
  239. 30:46there has been uh further development uh
  240. 30:48slowly that self-supervised learning
  241. 30:50approach has uh been adapted uh by lot
  242. 30:54of geospatial community especially the
  243. 30:56people who have big pockets. So they
  244. 30:59could consume the 40 30 years 40 years
  245. 31:02of uh the whole LANCAT imagery uh into
  246. 31:06modeling and then they had a lot of GPU
  247. 31:09power. So those are actually GPU rich uh
  248. 31:11I should say uh companies they invested
  249. 31:14a lot of money and accordingly we had a
  250. 31:16lot of uh uh developed uh the foundation
  251. 31:19models started coming. So these basic
  252. 31:22objective of these foundation models are
  253. 31:24that whatever the general information
  254. 31:27which is present uh in in any kind of
  255. 31:30geospatial data. So that has to be
  256. 31:32mapped into a certain uh latent space or
  257. 31:35intermediate uh space then we can have a
  258. 31:39certain upstream task which can then be
  259. 31:41used for doing some kind of basic task
  260. 31:43like classifications segmentations and
  261. 31:46so on. Now there has been a uh again
  262. 31:49continuously the development has been
  263. 31:50there and there has been an integrations
  264. 31:52of multimodel foundation architectures.
  265. 31:55So what this does is basically instead
  266. 31:57of just relying and giving an image and
  267. 31:59then predicting it use uh it it has
  268. 32:02started coming in term of summary. So
  269. 32:04the the Google has started giving a lot
  270. 32:07of services where based on the
  271. 32:09prediction of the input data sets it
  272. 32:11gives you the weather conditions and
  273. 32:13based on the different weather
  274. 32:14parameters like precipitations uh cloud
  275. 32:17coverage which model gets generated then
  276. 32:20on the basis of that that gets uh
  277. 32:22converted into a certain smaller text
  278. 32:25which can be sent to the user that there
  279. 32:27is a chance of certain amount of rain or
  280. 32:29there will be heavy rain or
  281. 32:31thunderstorms. So these all those uh
  282. 32:33models uh which was basic project uh
  283. 32:36basic object the models basic objective
  284. 32:38was earlier to predict the weather
  285. 32:41patterns or the land use land maps uh
  286. 32:44patterns instead of that now they
  287. 32:46started giving a specific prescriptive
  288. 32:49or user level summarizations of what
  289. 32:52those uh model means. So that was a big
  290. 32:55leap and the development is still going
  291. 32:57on in these domains. So these on the
  292. 32:59bottom you can see there's a timeline of
  293. 33:01all these uh basic vision based models
  294. 33:04which we just uh discussed. So out of
  295. 33:07those I'll just focus only the uh basic
  296. 33:10structures or the basic things which has
  297. 33:13been there especially in in term of
  298. 33:14segmentations which was based on
  299. 33:17convolutional neural networks especially
  300. 33:18the image based uh networks and the
  301. 33:21unit. So uh that was the basic uh thing
  302. 33:25which which was there and then later on
  303. 33:27the NLP was being uh used uh and that uh
  304. 33:31and in in the case of NLP you have an
  305. 33:34attention based system. So what exactly
  306. 33:35is the attentionbased uh systems so in
  307. 33:39the attention based system what happens
  308. 33:40is that you will be given certain
  309. 33:42inputs. So the way in the in term of
  310. 33:45language context we can say that uh
  311. 33:47let's say we say that there is a man.
  312. 33:50[clears throat] So after that what is
  313. 33:52the next word it is going to come that
  314. 33:55we can do a predictions on the basis of
  315. 33:57the language context uh its uses the
  316. 34:00word after the man how many times uh the
  317. 34:03man comes after uh a word a or what what
  318. 34:07is the next word which is coming after
  319. 34:09the man and so on. So those kind of uh
  320. 34:12scenarios gets generated based on the
  321. 34:15historical or large data sets available
  322. 34:18in uh available from different corpus
  323. 34:21and so on. So accordingly we get a kind
  324. 34:24of prediction of the next level. So this
  325. 34:26is what the transformer's job is that
  326. 34:29taking all the input which is available
  327. 34:32uh maybe a sequence of word maybe
  328. 34:34sequence of image patches and so on and
  329. 34:36then on the basis of that it tries to
  330. 34:38predict what may be the next word or the
  331. 34:41next picture. So if you are creating a
  332. 34:44video so it will get it will give you it
  333. 34:46will take a set of sequence of the
  334. 34:48images and after some time it will also
  335. 34:50generate what may be the next sequence
  336. 34:53of the images. So that's the kind of
  337. 34:54attention mechanism is being created or
  338. 34:57in case of images there might be a cases
  339. 34:59for example you may be having a very big
  340. 35:01images out of that you may give randomly
  341. 35:04four or five different patches and then
  342. 35:06you say reconstruct the whole images. So
  343. 35:08those are actually the kind of cases uh
  344. 35:10which is being done with the help of
  345. 35:12attention based uh uh models and then
  346. 35:16we'll also see the foundation based uh
  347. 35:18model era where these vision based uh uh
  348. 35:23models how it got integrated to
  349. 35:25different geospatial kind of
  350. 35:27informations like you have a multiple
  351. 35:29spectral band so that multisspectral
  352. 35:32band data and the ge uh geographical
  353. 35:35locations those information then got
  354. 35:37integrated and accordingly the uh the
  355. 35:41basically there has been an improvement.
  356. 35:43So if you see that uh at the bottom
  357. 35:45there's an accuracy table. So now at
  358. 35:47least we are in a in a stage where we
  359. 35:50are having an accuracy of all these uh
  360. 35:53foundation uh level models which where
  361. 35:56accuracy of the uh data sets or the any
  362. 35:59kind of image in in term of
  363. 36:00classifications you see it is around 94
  364. 36:0395% which is at at par with the human
  365. 36:07level uh informations.
  366. 36:09So let's see that uh how those uh u
  367. 36:13method has evolved. So earlier we say
  368. 36:15supervised uh learning I think most of
  369. 36:18you may be aware of that. So what you do
  370. 36:20you give an input and you create a
  371. 36:22certain labels and you say this input
  372. 36:25basically means this. So there may be a
  373. 36:27kind of uh you may be having a satellite
  374. 36:29imagery which will be which may get
  375. 36:31collected uh from different band uh band
  376. 36:35sensor data sets or in in normal cases
  377. 36:37we say you have a just clicked
  378. 36:39photograph from a camera and then from
  379. 36:42that camera you can say okay this is the
  380. 36:44human being this is a dog this is a cat
  381. 36:46and so on. So that labeling you do and
  382. 36:48then basically the systems tries to
  383. 36:51learn that how to differentiate the
  384. 36:54input feature which is the pixel level
  385. 36:56concentrations or organizations of those
  386. 36:58pixels to understand okay in in case of
  387. 37:01dog what kind of organization of those
  388. 37:03pixel will be there or in case of human
  389. 37:05beings what kind of organization of
  390. 37:06those pixels will be there those are
  391. 37:08actually supervised learning approach
  392. 37:11under the unsupervised learning we don't
  393. 37:13uh know how many classes are there but
  394. 37:16instead of that we look for certain kind
  395. 37:18of similarity. So those similarity may
  396. 37:20be at the pixel level or maybe at a
  397. 37:23regional level or maybe uh at an image
  398. 37:26level itself and then uh the systems
  399. 37:29tries to group them in a different
  400. 37:31category. So let's say you'll be having
  401. 37:33a 10 photographs. So what system will
  402. 37:35say that okay these two photo first two
  403. 37:37photographs looks like uh place uh of
  404. 37:41one place another two or three
  405. 37:42photographs looks maybe having a certain
  406. 37:45other places and so on. So it will give
  407. 37:46you those kind of grouping but you uh it
  408. 37:49will not tell you what those grouping
  409. 37:50means. So as a human being you have to
  410. 37:53tell okay these groups basically means
  411. 37:55uh maybe a hilly region maybe a uh uh
  412. 37:58you can say the flate uh flood plane
  413. 38:02regions and so on. Those kind of things
  414. 38:04uh can be carried out and these are very
  415. 38:06standard process which is being followed
  416. 38:08in geospatial domain. In case of
  417. 38:11self-supervised learning what happens is
  418. 38:13so the users is not giving an uh input.
  419. 38:18So rather than what happens is the
  420. 38:20machine is going to take your input it
  421. 38:23itself will have a certain pipeline
  422. 38:25where it will try to distort it. So it
  423. 38:27will do uh certain error which may be
  424. 38:30introduced into the images and then
  425. 38:33finally uh uh it will try to see what is
  426. 38:36the impact of those changes or those uh
  427. 38:40alterations process which has it which
  428. 38:43has taken place and accordingly it will
  429. 38:45try to understand uh what are the basic
  430. 38:48underlying uh structure. For example,
  431. 38:51let's say if one image is being given.
  432. 38:53So if you rotate it that particular
  433. 38:55image or that image let's let's say of a
  434. 38:57dog. So it will remain as a dog. So you
  435. 39:00rotate it 90° 60° or 50° there's no
  436. 39:03impact of rotation. Finally at the end
  437. 39:05of the day you have to say that image
  438. 39:07belongs to a dog. So in that case the
  439. 39:09machines basically rotates it in a
  440. 39:11multiple uh sections and then it will
  441. 39:14give you a kind of output where it will
  442. 39:15say okay it has understood that uh even
  443. 39:19uh rotating it multiple cases whatever
  444. 39:21the feature differences were there that
  445. 39:24should be projected in such a way that
  446. 39:26it should uh tell that this is a
  447. 39:28particular object uh is related to one
  448. 39:30specific object as a whole. Then there
  449. 39:33are other uh approach which is the
  450. 39:34reinforcement learning approach. So
  451. 39:36which is again uh a kind of uh I
  452. 39:39couldn't I should say this without
  453. 39:41labeling approach. So there instead of
  454. 39:43telling each and every features what you
  455. 39:45tell the systems that you are right or
  456. 39:48wrong or there may be a certain degrees
  457. 39:50can also be provided. So on the basis of
  458. 39:53that uh machines takes the actions and
  459. 39:57then it learns on the basis of that. So
  460. 39:59these are the major four learning
  461. 40:01approaches. So generally the foundation
  462. 40:04models or geospaceial foundation model
  463. 40:06works on a principle of uh
  464. 40:08self-supervised uh learning method. So
  465. 40:11under the self-supervised learning
  466. 40:13method what you do is you take an
  467. 40:15images. This is just an example where
  468. 40:17you do uh a certain uh pretext task. So
  469. 40:21that is nothing but a kind of distorting
  470. 40:23the input images. So these distortions
  471. 40:25can be you can take an image break down
  472. 40:27into multiple component and then you
  473. 40:29delete certain portions and then you say
  474. 40:32please predict the deleted portions.
  475. 40:35Other cases you may take certain patches
  476. 40:37you may rotate it or make some uh noises
  477. 40:40in all all those patches. So finally
  478. 40:43system needs to learn and tell that okay
  479. 40:45what kind of uh distortion or noise
  480. 40:48might have been added which might have
  481. 40:50resulted in the cases uh in that cases.
  482. 40:54So in generally what happens is you take
  483. 40:56an input data then you do do a certain
  484. 40:58pre uh task basically rotating breaking
  485. 41:01masking those kind of uh uh task which
  486. 41:04you carry out and then you send those
  487. 41:07partial informations to the uh a set of
  488. 41:11uh neural networks uh uh large neural
  489. 41:13networks or DNN networks which basically
  490. 41:17takes those data sets and create a kind
  491. 41:19of intermediate multiple or I should say
  492. 41:22multi-dimensional data assets
  493. 41:24projections which finally can be used
  494. 41:27for different downstream task like you
  495. 41:29may be interested in doing some kind of
  496. 41:31classifications segmentations and
  497. 41:34control and so on. So for example let's
  498. 41:36say at the bottom you see a lot of uh
  499. 41:38pen chromatic colors uh or maybe you may
  500. 41:42be having a true color images or SAR
  501. 41:44images multisspectral images and so on.
  502. 41:46So these data sets may be given to the
  503. 41:49input. So at the input level it will
  504. 41:51just try to understand okay there is a
  505. 41:53huge intensity at certain places there
  506. 41:55is a lesser intensity at other places or
  507. 41:58your system may also try to get certain
  508. 42:01additional information like okay if
  509. 42:03there is a high intensity so it is
  510. 42:05surrounded by a lower intensity or what
  511. 42:08is the probability of having a higher
  512. 42:10intensities data sets to be having an
  513. 42:12higher uh intensity in surrounding
  514. 42:15areas. So for example in the SAR as you
  515. 42:17can see there is a very less number of
  516. 42:20uh uh high intensity color which is
  517. 42:23being surrounded by uh the uh the high
  518. 42:27another high intensity areas. But in
  519. 42:30case of pan chromatic you can see
  520. 42:31there's a white patches so which is
  521. 42:33quite bigger and so on. So this kind of
  522. 42:36distinction it will try to understand it
  523. 42:38will not know okay this bright color
  524. 42:40means what is it a pond or this is a is
  525. 42:43it a kind of river bed or it is a
  526. 42:46building those kind of information it
  527. 42:48doesn't have but it just knows that okay
  528. 42:50within the image there are certain
  529. 42:52places where uh you're having a high
  530. 42:55intensity and low intensities there's a
  531. 42:58texture variations there is an intensity
  532. 43:00variations and so on I mean those kind
  533. 43:02of basic uh uh uh the parameters which
  534. 43:07which basically gives just information
  535. 43:10in term of color intensity their
  536. 43:12relationship and so on. So once these
  537. 43:14information gets collected in a latent
  538. 43:16space then finally you can create a
  539. 43:18certain downstream task. So in that case
  540. 43:20the advantage is that we don't need to
  541. 43:22build separate separate model for one
  542. 43:25model for object detection another model
  543. 43:27for classifications third model is just
  544. 43:29doing a change detections. So that was
  545. 43:32the earlier uh approach. So in the CNN
  546. 43:34models you have to just take let's say
  547. 43:36for example using list three if you do a
  548. 43:38classifications if you try to use the
  549. 43:40same uh classification algorithms let's
  550. 43:43say on the landset probably you will not
  551. 43:45be able to do it. So so that kind of
  552. 43:48thing now has been avoided with the help
  553. 43:50of uh these uh uh self-supervised uh
  554. 43:54learning approach. So I'll just focus
  555. 43:56the basic two models which is being
  556. 43:59mostly used in uh remote sensing domain.
  557. 44:02So one is that you have called a CNN
  558. 44:05model. So CNN and DNN are all those
  559. 44:08families of of a models where what it
  560. 44:11does is you take a very large area then
  561. 44:14you create a certain window size of the
  562. 44:16smaller area. So those window size maybe
  563. 44:183x3 maybe 16x 16s and so on and then you
  564. 44:22can create a kind of multi-layer window.
  565. 44:25So for example as you can see in the uh
  566. 44:28left side of the CNN images. So all all
  567. 44:32the images has been divided into 16
  568. 44:34boxes. So each of these boxes will
  569. 44:36having their own informations. So like
  570. 44:39first four on the top or maybe you can
  571. 44:42say the topmost area which we see in the
  572. 44:44CNN models it's an urban areas and so
  573. 44:47on. But in between where you see a
  574. 44:49larger 4x4 boxes which has been merged
  575. 44:52together. So that looks at green uh
  576. 44:55areas. So CNN basic job is to understand
  577. 44:59the spatial context and on the basis of
  578. 45:02just to look into the textural
  579. 45:05informations. For example, it will just
  580. 45:07say that okay there is a kind of green
  581. 45:09area which is there inside a in in
  582. 45:13somewhere. Similarly there may be a
  583. 45:14certain areas which is having a high
  584. 45:16textured uh uh areas uh which may be of
  585. 45:20the urban buildups and so on. So what
  586. 45:23happens is it is not able to tell you
  587. 45:25that this green area basically means a
  588. 45:28forest or it is a golf course or maybe
  589. 45:32uh a kind of park and so on. That kind
  590. 45:34of context is uh the CNN is not able to
  591. 45:38give you why because uh the CNN is not
  592. 45:41having any kind of global context. So
  593. 45:44what what does it mean is let's say you
  594. 45:46take any boxes in the CNN one. So if you
  595. 45:50see this um one of those boxes so it
  596. 45:53sees okay on the left side there's
  597. 45:54another box and so on but it doesn't
  598. 45:57understand that this box actually
  599. 45:59belongs to a larger green patch area
  600. 46:02because it is not having understanding
  601. 46:04of the whole scene which is present in
  602. 46:06the image. So accordingly it loses that
  603. 46:09global context and it is not able to get
  604. 46:12you the informations that what exactly
  605. 46:14this green piece is there. So, so but in
  606. 46:19case of vision based transformer or the
  607. 46:21earlier which I was talking about the
  608. 46:23attention which we need or image of 16x6
  609. 46:26pixel those papers basically talks about
  610. 46:28that. So in those cases what happens is
  611. 46:31it the vision based global detention
  612. 46:34mechanism they gets the context of
  613. 46:37global context as well. So in that
  614. 46:40context it sees okay if there is a kind
  615. 46:42of green area and it is surrounded by a
  616. 46:45lot of uh uh uh non-green area which
  617. 46:48looks like a kind of uh builtup areas.
  618. 46:51So the green patch in between has to be
  619. 46:53an urban park. It cannot be a forest
  620. 46:56because within the city it is not
  621. 46:59expected that you're going to have a
  622. 47:00forest. these kind of uh global uh I
  623. 47:04should say the
  624. 47:06uh uh basically context which vision
  625. 47:09vision based transformer is able to get
  626. 47:12in and this is actually a quite long
  627. 47:14range. So CNN is basically having a very
  628. 47:16small field specific uh smaller range uh
  629. 47:20informations but vision based
  630. 47:22transformer are having a global context
  631. 47:24and accordingly it is going to give you
  632. 47:26a kind of uh uh larger context and and
  633. 47:30it is able to give you the uh
  634. 47:33classification or categorizations much
  635. 47:36better in those cases. So these are just
  636. 47:38an example like how does CNN basically
  637. 47:40relies mostly on convolutional and max
  638. 47:43pooling layers which is generally been
  639. 47:45used and where in case of vision based
  640. 47:48transformer basically relies on uh the
  641. 47:51attention based mechanisms. So what does
  642. 47:54attention based mechanism and how it
  643. 47:56works we'll just see in the coming uh
  644. 47:58slides as well. So let's see the other
  645. 48:01examples where the how does the CNN and
  646. 48:04vision based transformer models can be
  647. 48:07useful. So as at the bottom you can see
  648. 48:09that there is a kind of original images.
  649. 48:12In case of CNN if that same area you
  650. 48:15take the another image of another time
  651. 48:18and it is cloud cover or it is having
  652. 48:21some kind of uh let's say uh you're
  653. 48:24having some kind of fog or maybe defocus
  654. 48:27all those kind of problems are there. So
  655. 48:29even if you have a very well-trained CNN
  656. 48:32model but if the input data set gets
  657. 48:34distorted or uh uh in in in in those
  658. 48:38cases what happens is the CNN model will
  659. 48:41fail and they will not be able to give
  660. 48:43you the exact uh output uh and giving
  661. 48:46you the clear pictures what exactly it
  662. 48:48is there where in case of vision based
  663. 48:51transformer what happens is that any
  664. 48:53kind of augmentations or error or noises
  665. 48:55which gets uh added so what happens is
  666. 48:59even if there's an uh uh uh let's say
  667. 49:02the contrast changes or there may be a
  668. 49:04certain uh uh missing informations and
  669. 49:08so on still what it will able to do is
  670. 49:11it will take whatever the partial
  671. 49:13information is available and then it
  672. 49:15will try to fulfill the uh or I should
  673. 49:18say that it will try to fill the missing
  674. 49:21data sets. So once it is able to fill
  675. 49:24the missing data sets then it is able to
  676. 49:26uh reconstruct the whole images which is
  677. 49:30uh which which is which is actually uh
  678. 49:32you can say the filtering of the input
  679. 49:35uh noises has been removed from the
  680. 49:37input images and accordingly what
  681. 49:39happens is you get an output which gets
  682. 49:42generated and then that particular image
  683. 49:45can be used for classification like
  684. 49:47urban park and so on. So this is basic
  685. 49:49advantages where in case of uh vision
  686. 49:53based transformer model it not only
  687. 49:55takes the input but it corrects the
  688. 49:57input data sets and then that correction
  689. 50:00is being done on the basis of global
  690. 50:02context which it might have learned
  691. 50:04during the learning process and then
  692. 50:06finally it gives you the categorizations
  693. 50:08and or other downstream classes like
  694. 50:10classifications or the uh other uh uh
  695. 50:15basically differences or maybe the time
  696. 50:17based
  697. 50:18changes all those things it can it can
  698. 50:20do it on the basis of those
  699. 50:22reconstructed image. So this is the
  700. 50:24basic uh differences which the
  701. 50:27transformer has built uh has brought in
  702. 50:30into on the table table and accordingly
  703. 50:32what has happened is this has been
  704. 50:34exploited by a lot of uh different uh uh
  705. 50:38foundation models development. So let's
  706. 50:41take an example of what kind of uh
  707. 50:44whenever we say that uh I I had said
  708. 50:46earlier that in case of self-supervised
  709. 50:48learning what happens is your uh model
  710. 50:52takes the input it tries to distort it
  711. 50:55certain input. So for example at the uh
  712. 50:58there's basically a creating a fake
  713. 51:00images. So as you can see that uh uh
  714. 51:03this is basically a kind of one approach
  715. 51:05where you just provide one input and
  716. 51:08model uh there will be a certain uh
  717. 51:10channel or you can say uh uh one models
  718. 51:13will be there or one uh maybe simpler
  719. 51:16models may be there related to where
  720. 51:17just doing an pixelbased rotations or
  721. 51:21certain other augmentations distortions
  722. 51:23can be there that you may be having
  723. 51:24instead of just rotations you may be
  724. 51:26having a uh similarity measures uh CNN
  725. 51:30and DN models which will give you a
  726. 51:33multi uh images or different type of
  727. 51:35images which will be almost similar but
  728. 51:38in different cases for example as you
  729. 51:40can see in the deer pictures which is
  730. 51:42being taken as an input then what we'll
  731. 51:44do is it the in in the CNN based models
  732. 51:47uh which is basically an exampler CNN
  733. 51:50based models uh especially uh has come
  734. 51:53in 2015 so what it does it gives you the
  735. 51:56different perspective of the same deer
  736. 51:58in a different contrast colors
  737. 52:00and so on. So it get those it gets
  738. 52:03created. So all the other images which
  739. 52:04you see except from the first one is a
  740. 52:07fake one. So this is the first pre-text
  741. 52:10task which is being carried out by the
  742. 52:12any kind of uh I should say
  743. 52:14self-supervised learning approaches
  744. 52:16where it will generate a lot of such a
  745. 52:18fake data sets and then try to
  746. 52:22understand that even though those fake
  747. 52:24data sets may be generating from the uh
  748. 52:26from the same uh single images but all
  749. 52:29of those fake data sets actually gives
  750. 52:31you similar context like it is all these
  751. 52:33images which you see is basically a kind
  752. 52:35of image of a deer. there. Similarly,
  753. 52:38there are another augmentation or
  754. 52:39distortion method which is just trying
  755. 52:41to learn or create uh rotations. So, you
  756. 52:44take an images, you do a multiple
  757. 52:46multiple rotations. So, your objective
  758. 52:48is even if you give uh even the model is
  759. 52:52being given the first image or maybe any
  760. 52:54four distorted image which is rotated by
  761. 52:56a different angles, the model should be
  762. 52:59able to tell that all those four images
  763. 53:02are actually same as the first image. So
  764. 53:05that's the other kind of distortion
  765. 53:07which self-supervised uh systems tries
  766. 53:09to create and trying to understand what
  767. 53:12exactly is the impact of uh rotations.
  768. 53:16Then so those are basically of simpler
  769. 53:18cases. Then in case of positional
  770. 53:21augmentations what happens is you take a
  771. 53:23randomly a certain images and then you
  772. 53:26randomly create certain neighbor
  773. 53:28patches. So neighbor patches need not
  774. 53:30have to be connected which is generally
  775. 53:33we expect uh in other CNN models. So
  776. 53:36like uh for example in the image as you
  777. 53:38can see there are nine uh different
  778. 53:41boxes. Each of these boxes have a
  779. 53:43relative positioning of each other. So
  780. 53:46and accordingly we can say that patch
  781. 53:48number one is actually on the uh west or
  782. 53:52northwest to the uh blue pixel.
  783. 53:55Similarly patch number two is actually
  784. 53:57not to the blue or those kind of
  785. 54:00positioning is fixed. So in this case of
  786. 54:03augmentation images what happens is you
  787. 54:06just give at any two pair of uh these
  788. 54:09patches to the models and then uh the
  789. 54:12model should say and tell you that where
  790. 54:15this patch basically belongs to which
  791. 54:18particular positions. So advantage of
  792. 54:20this is as you can see now if you build
  793. 54:22such kind of uh relative positioning
  794. 54:24your system model is able to learn then
  795. 54:27what happens is instead of having a few
  796. 54:29receptive uh window specific uh CNN
  797. 54:33focused uh I should say informations now
  798. 54:36it gets extrapolated beyond that so it
  799. 54:39is able to see beyond that in a whole
  800. 54:42image context and the impact of that is
  801. 54:44that you'll be able to uh uh get a
  802. 54:48larger uh context. So as we can see if
  803. 54:51you're able to see that at the center
  804. 54:53you're having a blue boxes and there is
  805. 54:55another uh place and the third basically
  806. 54:58which is on the topmost uh I should say
  807. 55:01the right side topmost that is the
  808. 55:03position number three and if you're able
  809. 55:05to tell and do a prediction very well
  810. 55:08and as you can see that if you're able
  811. 55:10to understand that at the blue it's a
  812. 55:12basically a face of cat and the three
  813. 55:16basically is a kind of cases which is
  814. 55:19just on the top left and right then you
  815. 55:21can say okay three is actually maybe
  816. 55:23most probably a kind of ear of the cat
  817. 55:26it cannot be a kind of leg or other uh
  818. 55:30cases. So these kind of additional now
  819. 55:32you can create start creating a kind of
  820. 55:34contextual informations which might be
  821. 55:36quite useful. So in those cases
  822. 55:38generally uh uh is being done. So it's a
  823. 55:41basically relative positioning.
  824. 55:42Similarly the other approaches is
  825. 55:44basically jig jigsaw puzzling. So what
  826. 55:46you do you give all those uh uh you just
  827. 55:49take those nine points or the nine
  828. 55:52patches surrounding and you randomly
  829. 55:54arrange them. So model's job is just to
  830. 55:56give exact positionings and that's how
  831. 55:59these systems is going to learn. In some
  832. 56:02cases you can also feed it to some fake
  833. 56:04images which is not part of these images
  834. 56:07and systems should be able to tell okay
  835. 56:09this doesn't belongs to any of those
  836. 56:11eight places this is somewhere coming
  837. 56:13from outside. So these kind of
  838. 56:16augmentation patching is the uh certain
  839. 56:19learning approaches which gets added to
  840. 56:21the self-supervised uh learning. So just
  841. 56:25to give you an idea you have original
  842. 56:27image and then you do a certain uh
  843. 56:29augmentations. So certain augmentations
  844. 56:31may be quite easy for example just
  845. 56:33changing the color or making a red to
  846. 56:35blue or green or something like that. So
  847. 56:38those are quite augmentation uh those
  848. 56:40are called weak augmentation techniques.
  849. 56:42uh basically it means that it is just uh
  850. 56:45uh the overall the context of whole
  851. 56:48image is still present uh uh in term of
  852. 56:51objects uh which is being there. Then
  853. 56:54there are something called strong
  854. 56:55augmentations. So which is basically
  855. 56:57distorting in a much larger way. So
  856. 57:00which means not only decreasing
  857. 57:02increasing the brightness putting lot of
  858. 57:04additional images. For example, as you
  859. 57:06can see that if you improve, if you
  860. 57:08increase the saturations, the context of
  861. 57:11the dog itself is completely lost. Now
  862. 57:15the system has to kind of learn and find
  863. 57:17it out. So if if you put if your system
  864. 57:20is able to get hold of these stronger
  865. 57:23augmentations uh or the pre-text tasks
  866. 57:26which you give and it is able to still
  867. 57:28reconstruct the original images from
  868. 57:30these outputs then you can say your uh
  869. 57:33image is quite good uh your models is
  870. 57:36quite good and generally in the strong
  871. 57:38augmentation cases CNN models are bound
  872. 57:41to fail. So just to show you the example
  873. 57:44that the same images where you see a lot
  874. 57:46of haziness and other images where
  875. 57:49simply distorting or maybe rotating uh
  876. 57:52some portion of the cases images the uh
  877. 57:56in the second case if your model is able
  878. 57:58to get the uh context well like as a
  879. 58:02human being from the second image we
  880. 58:04will still be able to tell okay there's
  881. 58:06a park uh there are some uh in the
  882. 58:08surrounding area there might be a kind
  883. 58:10of few places where a built-up area
  884. 58:13might be there. But in the CNN based
  885. 58:15model such kind of things is not
  886. 58:17possible. It will fail. But VIT or
  887. 58:20vision transformation models basically
  888. 58:21will also will overcome all these strong
  889. 58:24augmentation based method uh uh strong
  890. 58:28augmentations in your data sets and
  891. 58:30accordingly it will able to give you a
  892. 58:32better output uh results. So let's focus
  893. 58:36uh some of the uh basic uh some of the
  894. 58:39basic uh models or most primitive models
  895. 58:42which we call it a mask image modeling
  896. 58:44cases. So as you uh as in this the
  897. 58:47process approach is very simple. So what
  898. 58:50you take you take an image randomly you
  899. 58:52create certain patches and remove the
  900. 58:55information which is present in those
  901. 58:56patches. So as you can see those blue
  902. 58:59boxes are the places where no
  903. 59:01information is there. So initially you
  904. 59:03have an image then what you do is you
  905. 59:06simply remove uh you create certain
  906. 59:08patches area and remove the information
  907. 59:10which is present in those cases and your
  908. 59:12model job is wherever these new the
  909. 59:16information is missing based on the
  910. 59:18remaining informations which is present
  911. 59:20in the image. So in this case at least
  912. 59:22we see around 95% of the pixel is still
  913. 59:26intact. So on the basis of that it
  914. 59:28should be able to get you the complete
  915. 59:31information of all those blue boxes as
  916. 59:34well. So width basically is going to uh
  917. 59:37provide you this uh uh uh this output
  918. 59:41through the mask imaging modeling. So
  919. 59:43the overall model how it works is that
  920. 59:45you just take certain input images
  921. 59:47randomly you create certain pi patches
  922. 59:50delete it send one pipeline where
  923. 59:52encoders will be put put and you just
  924. 59:54send the another pipeline where original
  925. 59:56image which you already had it and then
  926. 59:59you see that how much it is able to uh
  927. 1:00:02see that uh uh one whatever the partial
  928. 1:00:04information is there it is trying to
  929. 1:00:06reconstruct the whole images and then
  930. 1:00:08you see what was the original image and
  931. 1:00:10what was the reconstructed uh images. If
  932. 1:00:13there is a biasness, then you tell the
  933. 1:00:16reconstructed images uh uh to recreate
  934. 1:00:19the whole uh uh recreate and do some
  935. 1:00:22kind of correction in inside and
  936. 1:00:24accordingly the mask imaging modeling
  937. 1:00:26basically uh gets uh added uh basically
  938. 1:00:30gets improvement and accordingly it gets
  939. 1:00:32the global context and slowly it start
  940. 1:00:34learning and reconstructing the images.
  941. 1:00:37One uh basic uh understanding of all
  942. 1:00:40these mask image modeling approach you
  943. 1:00:42need to remember is that these all these
  944. 1:00:44requires a very huge data sets. So
  945. 1:00:47during the training phase you cannot
  946. 1:00:49just have a smaller image data sets and
  947. 1:00:51then you can create a uh global context
  948. 1:00:54because to understand the global context
  949. 1:00:56you need to have a variety of input
  950. 1:00:59images. So let's take a a run through of
  951. 1:01:02that image uh in the current examples.
  952. 1:01:05So as you can see one image has been
  953. 1:01:06broken down into uh 16 uh basically the
  954. 1:01:098x8 pixel of patches. Then from those
  955. 1:01:13patches each patches is being sent to a
  956. 1:01:16transformer. So basically each pixel
  957. 1:01:18let's say you are having a 256x 256
  958. 1:01:21pixels. So that will get divided into 8x
  959. 1:01:238. So each patches will be 16x 16.
  960. 1:01:27So you just create all those 16x6 image
  961. 1:01:29pixels. you just put them and flatten
  962. 1:01:31them into a single stream of the pixel
  963. 1:01:34side pixels and then you simply send it
  964. 1:01:37through some kind of uh uh basically
  965. 1:01:41standard CNN models like imageet reset
  966. 1:01:43and so on and accordingly you'll get a
  967. 1:01:46latent space which can be considered as
  968. 1:01:48a uh tokenizer. So it's basically
  969. 1:01:50creating a tokenizers. So and then
  970. 1:01:53finally you uh once that pixels and
  971. 1:01:55corresponding tokenizer is being sent
  972. 1:01:58then those tokenizers uh uh from those
  973. 1:02:00tokenizers tokens what you do you just
  974. 1:02:02remove certain uh certain informations
  975. 1:02:05and then you recreate the whole images.
  976. 1:02:08So as you can see from the mask input
  977. 1:02:11images once it gets an uh tokens. So
  978. 1:02:15these tokens are uh a kind of uh image
  979. 1:02:18based to uh token generator systems are
  980. 1:02:20there. So that what it will do is it
  981. 1:02:22will create a kind of uh a kind of
  982. 1:02:25generic information which may be present
  983. 1:02:27in in each of those patches. Then from
  984. 1:02:30those patches what is being done is
  985. 1:02:31simply you remove uh certain uh
  986. 1:02:34information. So as you can see there are
  987. 1:02:36two path data path gets created. One is
  988. 1:02:40basically taking the original image and
  989. 1:02:41second one is removing certain patches
  990. 1:02:45and then you realigning all those
  991. 1:02:46patches uh and send it to an encoder and
  992. 1:02:50then you say what are their positions in
  993. 1:02:52the original images. So once the
  994. 1:02:54corresponding position is being
  995. 1:02:57identified that position is being uh
  996. 1:03:00compared uh or is being put uh into the
  997. 1:03:04whole structures of the input data sets
  998. 1:03:07and finally it is being sent to a
  999. 1:03:09decoder to reconstruct the images. So
  1000. 1:03:11you have an original uh uh mask uh image
  1001. 1:03:15in uh basically tokenizers and second
  1002. 1:03:18one is the reconstructed model which
  1003. 1:03:19gives you the uh regenerated part. So
  1004. 1:03:22there are certain parts where position
  1005. 1:03:24has already been correctly predicted and
  1006. 1:03:27then uh you reconstruct it. Then what
  1007. 1:03:29you do you simply look in these uh token
  1008. 1:03:32space or you can say the latent space
  1009. 1:03:34where you try to find out what's the uh
  1010. 1:03:37loss and accordingly you create a uh a
  1011. 1:03:41kind of uh uh uh learning process as
  1012. 1:03:44accordingly the whole weight gets
  1013. 1:03:45adjusted to encoder or decoder uh
  1014. 1:03:48systems. Finally as as there's no
  1015. 1:03:51improvement is going to be there in
  1016. 1:03:53reconstructed and original images your
  1017. 1:03:56systems will you'll say that you'll stop
  1018. 1:03:58the u uh basically the uh further
  1019. 1:04:01learning process. So these are few
  1020. 1:04:03examples of the same images and the
  1021. 1:04:05corresponding algorithms has been given
  1022. 1:04:07this it is available in mask image
  1023. 1:04:10modeling paper which is available on
  1024. 1:04:12archive. So you can just get it from
  1025. 1:04:14there then.
  1026. 1:04:17So once you have that uh uh
  1027. 1:04:19reconstructions so that that was the
  1028. 1:04:21first uh uh basically reconstruction
  1029. 1:04:23approaches then similarly we can have a
  1030. 1:04:25contrasting approaches. So what you do
  1031. 1:04:27you take an image do some kind of uh uh
  1032. 1:04:30augmentations for example in this case
  1033. 1:04:32the image has been rotated and as well
  1034. 1:04:34as the color has been changed you took
  1035. 1:04:37original images you reconstructed the
  1036. 1:04:39whole original images. So in the
  1037. 1:04:42previous case in in case of
  1038. 1:04:44reconstruction approach you do mapping
  1039. 1:04:46in latent space or after the tokenizer
  1040. 1:04:49or each patches is there. So you do the
  1041. 1:04:51measurement in the latent space there.
  1042. 1:04:54But in case of uh contrastive uh image
  1043. 1:04:58modeling approach which is again a mask
  1044. 1:04:59image modeling approach you take
  1045. 1:05:02original image itself and the job of the
  1046. 1:05:05encoder is to re to reconstruct the
  1047. 1:05:09whole image back the way it was there uh
  1048. 1:05:13in case of uh uh uh the original images.
  1049. 1:05:17So whatever the partial information
  1050. 1:05:19which is available so just the
  1051. 1:05:21reconstruction and uh positioning is
  1052. 1:05:23being carried out and accordingly you
  1053. 1:05:25just see okay what are the information
  1054. 1:05:27which is able to and how well it is able
  1055. 1:05:28to map it. So accordingly the mask image
  1056. 1:05:31modeling approach basically gets uh
  1057. 1:05:32created. So this is just an algorithm
  1058. 1:05:34for that. So the way we are having the
  1059. 1:05:37in the contrastive uh mask imaging uh
  1060. 1:05:40method we can also do lot of other
  1061. 1:05:43augmentations. You can do some kind of
  1062. 1:05:45rotations. You can do flipping. Color
  1063. 1:05:47may be shifted or uh you can do some
  1064. 1:05:50kind of glossial uh blurring. You can
  1065. 1:05:52also do resizing. You can also do some
  1066. 1:05:55kind of uh uh sun colorizations or the
  1067. 1:05:59little bit of cloud simulations and so
  1068. 1:06:01on. So all these different augmented
  1069. 1:06:04input can go through this channel. So
  1070. 1:06:07it's strong. You just create an uh
  1071. 1:06:09augmentation. So any one of these
  1072. 1:06:11augmentations can be applied and can be
  1073. 1:06:14sent to an uh strong augmented output
  1074. 1:06:17and then the whole chain can be
  1075. 1:06:19recreated. So uh in case of mask image
  1076. 1:06:22modeling you just try to remove certain
  1077. 1:06:24sections uh using uh one of the
  1078. 1:06:27approach. So you can use any other
  1079. 1:06:29approach in the uh uh other uh remote
  1080. 1:06:33sensing specific uh augmentations like
  1081. 1:06:36solarizations is uh is a remote sensing
  1082. 1:06:40specific algorithms where you can just
  1083. 1:06:42say okay if data set is in one band how
  1084. 1:06:44it will look into the second bands and
  1085. 1:06:46so on. Similarly if there's a spectral
  1086. 1:06:48shift so if there is a slight changes in
  1087. 1:06:51uh let's say green band from 5.52
  1088. 1:06:54nanometers to maybe 56.57 nanometers. So
  1089. 1:06:57what can be the impact in overall
  1090. 1:06:59images. So these kind of additional
  1091. 1:07:01augmentations methods can be added to
  1092. 1:07:04generate the uh uh augmented uh pipeline
  1093. 1:07:08which will then you can use it the mask
  1094. 1:07:10image modeling for the uh reconstruction
  1095. 1:07:13of the images. So these are a few
  1096. 1:07:15example and flow of these uh data
  1097. 1:07:19augmentations and use of the vision
  1098. 1:07:21transformation models especially in the
  1099. 1:07:23remote sensing cases. So you take the
  1100. 1:07:26strong augmentations you have any strong
  1101. 1:07:28augmentations you can have any one of
  1102. 1:07:30those which has been listed like uh
  1103. 1:07:33strong jittering gshian spectral row and
  1104. 1:07:36so on. So those kind of masking
  1105. 1:07:39tokenization. So those kind of things
  1106. 1:07:40you can add it and accordingly you can
  1107. 1:07:43use transformer encoded and use the
  1108. 1:07:46original image which is present so that
  1109. 1:07:48the corresponding reconstructions can be
  1110. 1:07:51used and this is the major uh I should
  1111. 1:07:54say approach which is being used by a
  1112. 1:07:57lot of uh larger uh I should say
  1113. 1:08:00foundation models. So there has been a
  1114. 1:08:03lot of I should say little uh a kind of
  1115. 1:08:06uh uh sampling strategy uh or the
  1116. 1:08:10augmentation strategies training
  1117. 1:08:12strategies but fundamentally all of them
  1118. 1:08:15actually follows the similar kind of uh
  1119. 1:08:18pipeline.
  1120. 1:08:20So in so basically uh so these these are
  1121. 1:08:22basically the cases where three or four
  1122. 1:08:25uh generally vision based transformer
  1123. 1:08:27are basically on two image or a three
  1124. 1:08:30band images but as we are aware that in
  1125. 1:08:33case of remote sensing cases we are not
  1126. 1:08:35going to have three bands but we'll be
  1127. 1:08:37having a multiple band so more than 12
  1128. 1:08:39bands is possible uh in some of the MSS
  1129. 1:08:42data and in case of hyperspectral you
  1130. 1:08:44may be having a bands in hundreds so
  1131. 1:08:47that is also possible. So the uh these
  1132. 1:08:51uh uh basically the uh all the vision
  1133. 1:08:54based transformer uh based foundation
  1134. 1:08:57model which is being used in the
  1135. 1:08:58geospatial domains. So they have to
  1136. 1:09:01adapt this multi- channelannel data sets
  1137. 1:09:04of different uh uh uh I should say range
  1138. 1:09:07of uh uh spectral band and uh it has to
  1139. 1:09:12and it it also has a certain positional
  1140. 1:09:15informations. So provisional information
  1141. 1:09:17is quite important. For example, let's
  1142. 1:09:19say if you go to Himalayas at glacial
  1143. 1:09:24cases. So if you take multiple photos or
  1144. 1:09:27multiple photograph of the same Himalaya
  1145. 1:09:30at multiple times. So you are going to
  1146. 1:09:33get or you will see mostly theis. So if
  1147. 1:09:36I or the or I should say snow uh will be
  1148. 1:09:40there or some kind of ice or snow may be
  1149. 1:09:42observed. So in that case what will
  1150. 1:09:45happen is always that area will be a
  1151. 1:09:47kind of white colors. So if we if you
  1152. 1:09:50know that these white colors are
  1153. 1:09:52basically located at certain latitude
  1154. 1:09:54and longitude. So we did not even have
  1155. 1:09:57to look into the input images just
  1156. 1:09:59knowing about the positions latitude and
  1157. 1:10:02longitude and corresponding height. So
  1158. 1:10:04we'll have a kind of pre uh training
  1159. 1:10:07information of that areas. So these kind
  1160. 1:10:10of additional fourdimensional or
  1161. 1:10:12positional informations informations can
  1162. 1:10:15be added or extracted along with the
  1163. 1:10:18multi- uh I should say uh multiband
  1164. 1:10:22spectral informations which becomes a
  1165. 1:10:25kind of uh major hallmark of all the
  1166. 1:10:28geospatial based uh foundation models.
  1167. 1:10:31So just to give you an uh basic idea
  1168. 1:10:34what exactly uh how it is being used. So
  1169. 1:10:37one of the uh models which is being used
  1170. 1:10:40is the clay models. Uh it is uh kind of
  1171. 1:10:43land land cover map models. The basic uh
  1172. 1:10:46foundation model uh in in this case what
  1173. 1:10:48it does is it takes the input data of
  1174. 1:10:51any band. So you can give this uh input
  1175. 1:10:55to uh input to this model maybe of 250 m
  1176. 1:10:58resolutions, 24 m resolution, lancet
  1177. 1:11:01imagery or I mean you can give modest
  1178. 1:11:04lancet or any kind of imagery and uh it
  1179. 1:11:08is independent of the input bands. But
  1180. 1:11:12once you give those inputs and
  1181. 1:11:14corresponding uh band information is
  1182. 1:11:16being fed then what happens is this will
  1183. 1:11:19give you an correct classified image map
  1184. 1:11:23uh which can be created from that image.
  1185. 1:11:25So it means like it frees uh basically
  1186. 1:11:27it lets you uh get uh I should say it is
  1187. 1:11:32not dependent upon the input spectral
  1188. 1:11:34band details. It is not dependent upon
  1189. 1:11:37the uh the kind of input resolutions or
  1190. 1:11:41the uh or the spatial resolutions
  1191. 1:11:43spectral resolutions. It is independent
  1192. 1:11:45of that. So how it is able to do this?
  1193. 1:11:48This is basically there are two layer of
  1194. 1:11:50uh
  1195. 1:11:52basically adaptation which is being done
  1196. 1:11:55apart from the part which we have
  1197. 1:11:57discussed. So first one is that the
  1198. 1:12:00reconstructions of the I should say the
  1199. 1:12:03spectral informations. So let's say you
  1200. 1:12:06are having a four band list three
  1201. 1:12:08images. So what it will do is it will
  1202. 1:12:10take a list three images but it will try
  1203. 1:12:13to use those list three central bands
  1204. 1:12:16and will try to construct a kind of
  1205. 1:12:19intermediate continuous band data values
  1206. 1:12:22to in a 128 channel. So it means like
  1207. 1:12:25your uh you can think of that this is
  1208. 1:12:27basically doing a uh uh linear
  1209. 1:12:30transformations from four band into 128
  1210. 1:12:34uh band informations or band level
  1211. 1:12:36informations. So that weight it it
  1212. 1:12:39doesn't gives you the exact uh I should
  1213. 1:12:42say band information but the weight of
  1214. 1:12:44each of those 128 channels which which
  1215. 1:12:47is being fixed uh as an intermediate uh
  1216. 1:12:50wave uh channel uh dimensions. So that
  1217. 1:12:54basically gets uh uh stored and that
  1218. 1:12:58weight is being used as an input to MEA
  1219. 1:13:02reconstruction process. So as we have
  1220. 1:13:03seen that MEA cases where you take an
  1221. 1:13:06image break down into multiple segments
  1222. 1:13:08and do some kind of random
  1223. 1:13:10rearrangement. So your model's job is to
  1224. 1:13:12reconstruct the whole uh uh reconstruct
  1225. 1:13:16the whole images based on partial
  1226. 1:13:18information. So at the middle of this
  1227. 1:13:21whole layer what you see is basically
  1228. 1:13:22the same B uh MEA reconstruction
  1229. 1:13:25approach. Only thing is like instead of
  1230. 1:13:27input images you also add the locations
  1231. 1:13:31uh location embedding and as well as the
  1232. 1:13:33latitude or longitude based time
  1233. 1:13:36embedding informations and then you pass
  1234. 1:13:38on to uh uh to the MEA reconstruction
  1235. 1:13:42channels where you have encoder and
  1236. 1:13:44decoder two channels. So this encoder
  1237. 1:13:47and decoder channel is uh decoder
  1238. 1:13:49channel is only active during the
  1239. 1:13:50learning phases. In case of uh uh
  1240. 1:13:54prediction phases this decoder is
  1241. 1:13:56discarded. So you directly take the
  1242. 1:13:58latent space which is present in the
  1243. 1:14:00dimension D or the purple color which
  1244. 1:14:02you see in solid D and these D uh latent
  1245. 1:14:06space are then being used for different
  1246. 1:14:10uh downstream classes uh uh basically
  1247. 1:14:13cases like classifications and so on.
  1248. 1:14:16So there has been a multiple development
  1249. 1:14:18apart from MEA uh recently till recently
  1250. 1:14:22there has been a kind of video mass
  1251. 1:14:23encoder and so on. So these are well
  1252. 1:14:25published uh papers which has been put
  1253. 1:14:28into CBPR and other uh well uh
  1254. 1:14:32conference uh organizations. So what
  1255. 1:14:36happens is this same thing whatever we
  1256. 1:14:38have done for the image classifications
  1257. 1:14:40which is the clay is d is doing the same
  1258. 1:14:43kind of task can also be done in case of
  1259. 1:14:45numerical weather prediction model. So
  1260. 1:14:47difference between the classifications
  1261. 1:14:50and numerical weather prediction model
  1262. 1:14:52is that it is time function. So in the
  1263. 1:14:55classification cases that is time
  1264. 1:14:57independent. So whatever image you give
  1265. 1:14:59you simply get the corresponding output.
  1266. 1:15:01But in case of uh weather predictions
  1267. 1:15:04and so on. So you may be giving an
  1268. 1:15:06instance maybe let's say in the morning
  1269. 1:15:076:00 the different atmospheric condition
  1270. 1:15:11uh conditions. So your model job is
  1271. 1:15:13basically to give you the output
  1272. 1:15:15for a future maybe 3 days 5 days 7 days
  1273. 1:15:19predictions is there. So means it has to
  1274. 1:15:21understand the underlying evolutions of
  1275. 1:15:24uh different physical processes and then
  1276. 1:15:28it should give you the output. So that
  1277. 1:15:31is being done uh using the uh one of the
  1278. 1:15:34models which is being developed by NASA
  1279. 1:15:37and uh IBM that's called Priti. So Priti
  1280. 1:15:41basically has two different uh P3 WXC uh
  1281. 1:15:44that's the model exact name which is uh
  1282. 1:15:47being provided by them. So what they
  1283. 1:15:49have done is they have taken around uh
  1284. 1:15:51uh 20 years of the whole numerical
  1285. 1:15:54predicted data sets from uh WRF uh
  1286. 1:15:58output data which is available from ECMF
  1287. 1:16:01uh W systems and then uh uh they have
  1288. 1:16:06they have used all those data sets to
  1289. 1:16:08understand the different underlying
  1290. 1:16:10weather pattern in multi-year levels and
  1291. 1:16:13each day information is also been uh is
  1292. 1:16:16generated uh and is that whole model
  1293. 1:16:19also again works on the same MEA cases.
  1294. 1:16:22Only thing is that there are two level
  1295. 1:16:24of MEA. So instead of having a single
  1296. 1:16:27initial patch and deleting, you may be
  1297. 1:16:29having a whole group of patch which may
  1298. 1:16:30be created and that's how the uh the
  1299. 1:16:34models the foundation models which is
  1300. 1:16:35being used for generating the uh weather
  1301. 1:16:38predictions and so on. So these are the
  1302. 1:16:40few two basic or I should say the time
  1303. 1:16:43domain as well as uh the spatial domain
  1304. 1:16:46weather prediction models uh which is
  1305. 1:16:49currently available. Apart from that
  1306. 1:16:51there are also uh some other models like
  1307. 1:16:53SATM spectra ch uh GPT. So it is also a
  1308. 1:16:58kind of temporal vision transform
  1309. 1:16:59models. Then you have a scala and uh
  1310. 1:17:02scale me. So these are basically small
  1311. 1:17:04uh improvement taking into
  1312. 1:17:06multi-reolution data set. For example,
  1313. 1:17:08scale is able to take any data sets
  1314. 1:17:10input which is from 30 cm to 100 m uh to
  1315. 1:17:1410 m resolutions and so on. So these are
  1316. 1:17:17the other few development which has
  1317. 1:17:18happened in uh foundational models. Then
  1318. 1:17:22there's a DOA model which is again is
  1319. 1:17:24being used for generating or combining
  1320. 1:17:27the data sets from the radar to optical.
  1321. 1:17:29So radar it it the doa model has seen
  1322. 1:17:32the sentinel one that is a radar data
  1323. 1:17:34set 2 and so on and clay I think I have
  1324. 1:17:38already uh discussed much in detail. So
  1325. 1:17:41this is just a few summary of the models
  1326. 1:17:44which is available in the geospatial
  1327. 1:17:46domain and you can use uh any one of
  1328. 1:17:48them and they are having their own uh
  1329. 1:17:50advantages and uh and most of them can
  1330. 1:17:53be used for their lot of downstream
  1331. 1:17:55task. So once these model data sets are
  1332. 1:17:57publicly available so as an end user you
  1333. 1:18:00just need to take the input and create
  1334. 1:18:01certain downstream tasks like
  1335. 1:18:03classification change detections and so
  1336. 1:18:05on that can be created. So with this uh
  1337. 1:18:08we are just ending today's sessions and
  1338. 1:18:10we'll be happy to take certain questions
  1339. 1:18:12as a part of these sessions.
  1340. 1:24:24Okay. So I think uh we can take certain
  1341. 1:24:27questions. Uh so I think uh from the
  1342. 1:24:31beginning there are few questions
  1343. 1:24:36which is related to registrations and
  1344. 1:24:37other quizzes. I think you might be
  1345. 1:24:39knowing already that process. So I'm not
  1346. 1:24:41going to talk all those. I'm just
  1347. 1:24:43focusing on the current lectures. So
  1348. 1:24:46first questions I'll take from Sans
  1349. 1:24:48Sharma. So it's basically that how does
  1350. 1:24:50the CL clay model basically integrate
  1351. 1:24:53spatial spectral and temporal
  1352. 1:24:55informations in its architecture to
  1353. 1:24:58improve the remote sensing images. So
  1354. 1:25:00first of all uh uh I should say that the
  1355. 1:25:03clay model what it does is uh let's go
  1356. 1:25:06back to uh maybe
  1357. 1:25:11I think you might have seen uh
  1358. 1:25:18the architecture which is
  1359. 1:25:24model.
  1360. 1:25:27Yeah.
  1361. 1:25:30So if you see the uh if you observe the
  1362. 1:25:33clay models, so one is called wave
  1363. 1:25:36transformer. Okay. So that's the first
  1364. 1:25:38stage and there are two stage wave
  1365. 1:25:40transformer transformer which clay model
  1366. 1:25:42uses. So in this case what happens is
  1367. 1:25:45that uh the first wave transformer model
  1368. 1:25:49which you are seeing that takes the
  1369. 1:25:52center wavelength of the input data set.
  1370. 1:25:54So let's say you are giving uh list four
  1371. 1:25:56images. So list four images having a
  1372. 1:25:58four bands. So red, green, uh and the
  1373. 1:26:01corresponding swear and I bands. So what
  1374. 1:26:04you have to give you have to just give
  1375. 1:26:06the central wavelength as an input that
  1376. 1:26:08has to be provided as an input. So the
  1377. 1:26:11first stage uh uh transformer which you
  1378. 1:26:13are seeing as a wave transformer. What
  1379. 1:26:15it does is it takes those four band
  1380. 1:26:18input images and considers them as a
  1381. 1:26:23four distinct input points at those uh
  1382. 1:26:26locations. So you may be having a 0455
  1383. 1:26:3075 and let's say 1 uh uh uh 1.13 I mean
  1384. 1:26:35that sorry 11.3 and so on. So these kind
  1385. 1:26:38of different uh band uh sequencing is
  1386. 1:26:41being put in. So using those four or
  1387. 1:26:43five bands uh information or central
  1388. 1:26:46wavelength what it does is it tries to
  1389. 1:26:48create a 128 projected uh uh linear
  1390. 1:26:52transformations of that data sets. Uh so
  1391. 1:26:55what happens is that only using the four
  1392. 1:26:57bands it knows out of 128 which it's
  1393. 1:27:00expecting as an as a part of its own uh
  1394. 1:27:03uh embedding spaces it will be able to
  1395. 1:27:06see okay how much weightage has to be
  1396. 1:27:08given to those 128 bands which
  1397. 1:27:11internally it just maintains. So those
  1398. 1:27:13are basically the uh internal embedding
  1399. 1:27:15band which is being kept in. So
  1400. 1:27:18accordingly what happens is that those
  1401. 1:27:20weights are then being used uh different
  1402. 1:27:22waiting system is being used. So that's
  1403. 1:27:24how it is able to understand the
  1404. 1:27:27different spectral information. So if
  1405. 1:27:29some of the cases let's say and that
  1406. 1:27:30that weight is getting again uh uh
  1407. 1:27:33generated through 4year transformations
  1408. 1:27:36uh model approach. So using the 4year
  1409. 1:27:38transformation coefficients is being
  1410. 1:27:40used to generate the weight. So if
  1411. 1:27:42you're having a four bands so there will
  1412. 1:27:44be certain weights which usually
  1413. 1:27:46consider the four uh bands to give or
  1414. 1:27:48arrive the weightage to all the 128 uh
  1415. 1:27:52different uh embedding space. Similarly
  1416. 1:27:54you may be having 12 bands. So
  1417. 1:27:56accordingly those 12 bands will be used
  1418. 1:27:58to generate the uh corresponding uh uh
  1419. 1:28:02spectral embedding spaces that gets
  1420. 1:28:03created. So that's how the spectral band
  1421. 1:28:06is getting uh uh extracted. the temporal
  1422. 1:28:09information it doesn't directly encodes
  1423. 1:28:12it. So instead of that what it does is
  1424. 1:28:14it simply does the normalizations. So
  1425. 1:28:16during the modeling or the I should say
  1426. 1:28:19the fine-tuning process what you need to
  1427. 1:28:21do is you need to take the input and
  1428. 1:28:23then you have to do a jet scaling of all
  1429. 1:28:26the data sets. So uh the temporal
  1430. 1:28:29information at such uh it it uh first or
  1431. 1:28:32the image level simply does the kind of
  1432. 1:28:34normalizations and the temp temporal
  1433. 1:28:37information basically does the s cosine
  1434. 1:28:39transformation. So that's that's how it
  1435. 1:28:41is uh is being added to the embedding
  1436. 1:28:44space in the next uh uh basically
  1437. 1:28:48next uh uh me stage. So that's how it is
  1438. 1:28:51going to give you the uh it it takes
  1439. 1:28:54care of the overall multisspectral and
  1440. 1:28:57as well as the uh temporal informations
  1441. 1:29:00in in that sense. Then we say uh then
  1442. 1:29:03Virra is asking how what does the govt
  1443. 1:29:06learns and machine learning use in
  1444. 1:29:08learning patterns to deep the image
  1445. 1:29:10pixels. So basically I think probably uh
  1446. 1:29:13the question is might be reframed that
  1447. 1:29:16uh what does the goit learns and what
  1448. 1:29:19may be the learning in other approach
  1449. 1:29:21like CNN and so on. So CNN and other
  1450. 1:29:24approaches may be looking for a kind of
  1451. 1:29:27local context and pixel to pixel never
  1452. 1:29:30ne never pixels to ne like some pixel
  1453. 1:29:33are there so what is the relationship
  1454. 1:29:34with the neighboring pixels. So
  1455. 1:29:36accordingly what it will do it will get
  1456. 1:29:37a kind of context. Okay, for example,
  1457. 1:29:40let's say if you're having a single road
  1458. 1:29:43pixels and it is surrounded by a tree,
  1459. 1:29:46so you may be having a one pixel with
  1460. 1:29:48the road and neighboring to that you may
  1461. 1:29:50be having an agriculture pixels and so
  1462. 1:29:52on. So you will be you'll be having that
  1463. 1:29:53kind of informations. So when it
  1464. 1:29:56understands this relationship okay so in
  1465. 1:29:58most of the places let's say in the
  1466. 1:29:59whole world around 80% of the areas it
  1467. 1:30:02sees that at the just next to the road
  1468. 1:30:05pixels of the black uh pixels there is a
  1469. 1:30:08yellow color uh I should say the
  1470. 1:30:10agriculture farmland is there. So
  1471. 1:30:12accordingly what it will do is it will
  1472. 1:30:14try to uh understand and tell that
  1473. 1:30:18wherever black pixels are there. So that
  1474. 1:30:20has to be a kind of road classes and it
  1475. 1:30:23has to have a kind of continuity because
  1476. 1:30:25in between suppose there is a mask of
  1477. 1:30:27the trees but since it knows that this
  1478. 1:30:30this pixel belongs to a uh road pixels.
  1479. 1:30:34So in those cases even though there
  1480. 1:30:35might be a overhead uh or let's say
  1481. 1:30:38occlusion of the trees on the top of the
  1482. 1:30:40road but still it will try to
  1483. 1:30:42reconstructed. So it means that's the
  1484. 1:30:44convolutional uh systems works but that
  1485. 1:30:47works only in a localized uh field I
  1486. 1:30:50should say or within the specific
  1487. 1:30:52windows where in case of uh geoid it has
  1488. 1:30:55a overall global context. So within the
  1489. 1:30:57whole image this uh the systems learns
  1490. 1:31:02what is the position of these pixels or
  1491. 1:31:04what are the uh group of pixels which is
  1492. 1:31:06a patch where where it is exactly
  1493. 1:31:09located what kind of information is
  1494. 1:31:12present in that patch and who all are
  1495. 1:31:15the neighbor to this patch and so on. So
  1496. 1:31:18correspondingly what happens is the
  1497. 1:31:20patch is able to get local global and as
  1498. 1:31:23well as internal context as a whole uh
  1499. 1:31:27from the pixel information and
  1500. 1:31:28accordingly it will be able to
  1501. 1:31:30reconstruct the uh the whole uh uh I
  1502. 1:31:33should say informations or the semantic
  1503. 1:31:35information which is present. So this is
  1504. 1:31:37how it takes uh that uh then there's a
  1505. 1:31:42question of for disaster forecasting in
  1506. 1:31:44NWB. Yes, definitely and that is the
  1507. 1:31:46basic advantage which is in compared to
  1508. 1:31:49uh weather forecast WRF models and so
  1509. 1:31:52on. So in the forecast forecast models
  1510. 1:31:54if you see that uh to do a predictions
  1511. 1:31:57uh for let's say next 6 hours
  1512. 1:32:01predictions even in the decent good uh
  1513. 1:32:04machines you may be requiring hours of
  1514. 1:32:07uh data processing even in this good
  1515. 1:32:11number of HPC environment and in uh but
  1516. 1:32:14in case of uh uh basically the
  1517. 1:32:16foundation models like P3 WXCX you'll be
  1518. 1:32:19able to get the output in seconds. So
  1519. 1:32:21advantage of that is that you have a
  1520. 1:32:23higher speed up. Uh then uh it also
  1521. 1:32:27understands the overall natures or uh uh
  1522. 1:32:30basically the prediction area and
  1523. 1:32:32corresponding context which has happened
  1524. 1:32:34in last 20 years uh 2020 years. So it
  1525. 1:32:38means like it is having a it is going to
  1526. 1:32:40give you a predictions not only from the
  1527. 1:32:42current u I should say uh the
  1528. 1:32:45atmospheric condition but it also have
  1529. 1:32:47an understanding of what was the overall
  1530. 1:32:50average weather patterns in each of the
  1531. 1:32:53global uh glo uh as a whole uh within a
  1532. 1:32:57globe. So bringing the combining those
  1533. 1:33:00two informations and faster predictions
  1534. 1:33:02it will be able to give you a better uh
  1535. 1:33:05result. For example, if you just combine
  1536. 1:33:06them with case of disaster areas. So
  1537. 1:33:09what may happen is like if you if you
  1538. 1:33:11just combine those uh disaster special
  1539. 1:33:13embedding where uh large number of
  1540. 1:33:15disaster happens for example landslides
  1541. 1:33:18which is mostly found in the I should
  1542. 1:33:21say hilly regions or maybe a cloud bus
  1543. 1:33:24which is again in the hilly region. So
  1544. 1:33:26if you just do a fine-tuning and attach
  1545. 1:33:28additional uh elevation informations to
  1546. 1:33:31these uh foundation models then it will
  1547. 1:33:34not only give you the heavy predictions
  1548. 1:33:36precipitations area but rather than it
  1549. 1:33:39will also will be able to tell you that
  1550. 1:33:41there is also likely to have some kind
  1551. 1:33:43of landslide and so on. So it is
  1552. 1:33:45definitely going to be much helpful in
  1553. 1:33:48uh in in that sense and that's that's
  1554. 1:33:50how it is going beyond the uh numerical
  1555. 1:33:54prediction which just gives you the
  1556. 1:33:55atmospheric condition uh atmospheric
  1557. 1:33:57conditions from the different numerical
  1558. 1:34:00weather prediction models but these uh
  1559. 1:34:02AI models gives you beyond that okay
  1560. 1:34:04let's say if there's a heavy rain what
  1561. 1:34:06is going to the impact there will be
  1562. 1:34:08flooding there will be no flooding what
  1563. 1:34:10area will get inundated and so on so all
  1564. 1:34:13these things can be possible possible
  1565. 1:34:15through the numerical uh uh weather
  1566. 1:34:17prediction uh foundation models that can
  1567. 1:34:20be used. Okay. So quizzes I think it
  1568. 1:34:23will be available uh on the portal. So
  1569. 1:34:25you can look into that. That's the
  1570. 1:34:26questions to Kala. Yeah. One animesh has
  1571. 1:34:31also asked one questions related is that
  1572. 1:34:34how does the clay dynamic embedding
  1573. 1:34:36block process the variable number of
  1574. 1:34:38spectral bands compared to the standard
  1575. 1:34:41uh vision transformation. So wave uh
  1576. 1:34:44models if you take it uh as I was
  1577. 1:34:47telling so just do a kind of you can you
  1578. 1:34:50can think of uh in very layman's
  1579. 1:34:52language not exactly equivalent but just
  1580. 1:34:55think of that you're having a
  1581. 1:34:56multisspectral four band and you are
  1582. 1:34:58trying to reconstruct the hyperspectral
  1583. 1:35:01bands in 128 bands. So if you're having
  1584. 1:35:04a hyperspectral 128 bands and uh you
  1585. 1:35:08need to reconstruct it from the four
  1586. 1:35:09bands then what you'll do you'll simply
  1587. 1:35:11assign the certain weightage to those uh
  1588. 1:35:15uh 128 bands in such a way that
  1589. 1:35:18summation or or the uh I should say
  1590. 1:35:20linear summation of those bands is
  1591. 1:35:23equivalent to those four bands or five
  1592. 1:35:24bands whatever it is available to you.
  1593. 1:35:27So that's how it is able to uh create
  1594. 1:35:30it. uh if you're interested in more math
  1595. 1:35:32probably you can send me a mail I can
  1596. 1:35:34just send you the uh refer you can refer
  1597. 1:35:37the paper and there are also an publicly
  1598. 1:35:40available uh I should say uh uh
  1599. 1:35:42repository is also there where clay
  1600. 1:35:44models uh source code is available so
  1601. 1:35:47you can easily use and you can look more
  1602. 1:35:50in detail uh from the github repository
  1603. 1:35:53so this is also available on the hugging
  1604. 1:35:55space there also uh you can do that uh
  1605. 1:35:59how does the juda processing contributes
  1606. 1:36:01to national development projects. So as
  1607. 1:36:03any other technologies
  1608. 1:36:05uh the uses applications are quite wide.
  1609. 1:36:09It depends upon the uh the adaptation of
  1610. 1:36:12these tools uh in an decision-m process.
  1611. 1:36:15It will be helpful in those uh similar
  1612. 1:36:17cases. Uh okay. So I think I have
  1613. 1:36:21answered already that how does the NWP
  1614. 1:36:24works for disaster forecasting. So I
  1615. 1:36:26think with based and what their
  1616. 1:36:28advantage so I think uh Axa Singsh I
  1617. 1:36:31think I already have answered that and
  1618. 1:36:34uh how does the model handle spatial and
  1619. 1:36:36temporal irregularity for missing data
  1620. 1:36:39across so that's what I was telling that
  1621. 1:36:41as a part of MEA the model's objective
  1622. 1:36:45is that whatever partial information it
  1623. 1:36:47is having on the basis of that it has to
  1624. 1:36:50reconstruct the whole condition. So
  1625. 1:36:52there has been one paper and on priti
  1626. 1:36:55and they are claiming that if even if
  1627. 1:36:57you're having only five percent of the
  1628. 1:37:00area where you may be having an
  1629. 1:37:02atmospheric uh informations from that 5%
  1630. 1:37:06of the informations the model is able to
  1631. 1:37:08construct the remaining 95 and the
  1632. 1:37:11accuracy of those reconstructed 95 area
  1633. 1:37:14has 95 portion uh of the remaining area
  1634. 1:37:18which was missing has been found to be
  1635. 1:37:20around 85 to 90 uh% accuracy level. So
  1636. 1:37:25this this is this this is a kind of uh
  1637. 1:37:27regeneration uh data sets which is which
  1638. 1:37:31is possible through uh such models. So I
  1639. 1:37:34think uh you if you want if you're
  1640. 1:37:36interested you can just have uh the
  1641. 1:37:38better understanding by looking into
  1642. 1:37:40codes and the other models uses uh
  1643. 1:37:43especially the prit and the clay model
  1644. 1:37:45all those are publicly available and
  1645. 1:37:47available on the GitHub so you can use
  1646. 1:37:49them and uh if you have any further
  1647. 1:37:52questions just write me uh and I'll be
  1648. 1:37:55able to answer all of them.
  1649. 1:38:03Uh yeah there is one question from Niha
  1650. 1:38:05how is SSL different from supervised
  1651. 1:38:08learning. So as I have said that in case
  1652. 1:38:11of supervised learning you have to
  1653. 1:38:12recreate the labels. So as an uh uh
  1654. 1:38:16person you have to label it create a
  1655. 1:38:18mask and then you have to provide it to
  1656. 1:38:20the learning algorithm where SSL uh what
  1657. 1:38:24it does it instead of having a uniform
  1658. 1:38:27labels it just randomly creates the uh
  1659. 1:38:30certain augmentation and distortion in
  1660. 1:38:32the images and try to reconstruct the
  1661. 1:38:34whole images. So it's it's a kind of
  1662. 1:38:36self-arning. So it's a kind of you can
  1663. 1:38:38think of that one kid is sitting and
  1664. 1:38:40he's trying to learn uh how to put one
  1665. 1:38:43block on the top of the others. So you
  1666. 1:38:45just pass on a certain blocks that
  1667. 1:38:48person that uh kid may be putting it one
  1668. 1:38:51block on the top maybe initially
  1669. 1:38:53starting at the edges then slowly it
  1670. 1:38:55will start putting at the center and
  1671. 1:38:57then later on it learns that it has to
  1672. 1:39:00put that another block exactly aligned
  1673. 1:39:03with the central of uh gravity of each
  1674. 1:39:06of those blocks. So that's how the uh
  1675. 1:39:08models also learns. It just does certain
  1676. 1:39:11distortion maybe rotation maybe deleting
  1677. 1:39:13certain process and trying to recreate
  1678. 1:39:15it again after rotation again
  1679. 1:39:18reconstruct back the original
  1680. 1:39:19orientation of the images and so on. So
  1681. 1:39:21those kind of things is being done and
  1682. 1:39:23that's how the uh available. So I think
  1683. 1:39:26I have already answered about the priti
  1684. 1:39:28wxc models. Yes, that is publicly
  1685. 1:39:31available. There's another model from
  1686. 1:39:33Google which is called graphcast.
  1687. 1:39:35uh now cast. So those kind of models uh
  1688. 1:39:38can also be used and those are also
  1689. 1:39:40again uh AI based uh NWP models which
  1690. 1:39:44can be used uh for uh different uh use
  1691. 1:39:49cases.
  1692. 1:39:50Uh I think rest uh are uh so I think
  1693. 1:39:54probably more of these methods or
  1694. 1:39:57questions which you have already asked.
  1695. 1:39:59Yeah. 1 C1 is asking that what is the
  1696. 1:40:01glossian blur and how can we understand
  1697. 1:40:05what is the what the data is in the
  1698. 1:40:07fourth dimension. So when we say four
  1699. 1:40:09dimensions generally means having the
  1700. 1:40:12additional spectral information spatial
  1701. 1:40:14informations uh especially the latitude
  1702. 1:40:17longitude and other band which is not
  1703. 1:40:21being used in normal vision based
  1704. 1:40:23systems. So all the other vision based
  1705. 1:40:25system just uses three band images RGB
  1706. 1:40:27images which generally get captured
  1707. 1:40:29through normal camera but in case of
  1708. 1:40:32geospatial domain we go beyond that. So
  1709. 1:40:34you may be having a 12 band spectral
  1710. 1:40:36band data sets. So all of the bands you
  1711. 1:40:39cannot view at the same time. You can
  1712. 1:40:42also include certain bands for example
  1713. 1:40:44the terrain informations can be included
  1714. 1:40:47during the training process. So the
  1715. 1:40:49model also have an understanding and
  1716. 1:40:51context of uh I should say terrain based
  1717. 1:40:54information or positional informations.
  1718. 1:40:56So these kind of additional uh
  1719. 1:40:58informations you can put all all the all
  1720. 1:41:00of them together in fourth uh uh
  1721. 1:41:03dimensions.
  1722. 1:41:06Yes. So Google Earth Engine is also
  1723. 1:41:07having one uh foundation model
  1724. 1:41:09especially the embedding space. I think
  1725. 1:41:11it is having a 64 band uh currently
  1726. 1:41:14currently I'm not able to recall its
  1727. 1:41:16names but uh it is also having an
  1728. 1:41:19foundation model outputs and the latent
  1729. 1:41:21space uh data sets has been created and
  1730. 1:41:24it is found to be quite uh uh I should
  1731. 1:41:27say the accuracy is quite good which can
  1732. 1:41:29be used for different classifications
  1733. 1:41:31and segmentation task. So just have a
  1734. 1:41:34look into Google Earth Engine and look
  1735. 1:41:36for uh uh uh foundation models data
  1736. 1:41:39sets. probably you'll get uh uh data uh
  1737. 1:41:42I mean uh the data set which is
  1738. 1:41:44available and create generated by Google
  1739. 1:41:48okay how does we decide which weather
  1740. 1:41:50model is best for particular reason so
  1741. 1:41:53so the selection of any weather model is
  1742. 1:41:57true or any kind of activity which we do
  1743. 1:42:00is is always dependent upon the user's
  1744. 1:42:03uh choices. So if you think certain
  1745. 1:42:06models are good in your conditions you
  1746. 1:42:08see there's a very good accuracy
  1747. 1:42:09predictions or sometime you may be
  1748. 1:42:12interested going beyond that maybe
  1749. 1:42:14sensitivity analysis you can carry out
  1750. 1:42:17so on the basis of that you take a
  1751. 1:42:18decision which model is good so that
  1752. 1:42:20decision has to be taken by the users uh
  1753. 1:42:23and accordingly you can accept like in
  1754. 1:42:25any other cases like WRF models or there
  1755. 1:42:28are a lot of weather prediction models
  1756. 1:42:29are there IMD is also providing the data
  1757. 1:42:32sets but if you think that uh that model
  1758. 1:42:34is good enough and it is giving you a
  1759. 1:42:37very good understanding of what is going
  1760. 1:42:38to happen. It is matches with the
  1761. 1:42:40reality then you accept it otherwise you
  1762. 1:42:42simply reject it. So same is also the
  1763. 1:42:44case with any kind of other uh model uh
  1764. 1:42:48as well. Uh yeah so then the one
  1765. 1:42:50questions the use of latent layer in the
  1766. 1:42:53clay model. So latent layer clay model
  1767. 1:42:55as I was telling that when you when your
  1768. 1:42:58model learns uh so during the uh
  1769. 1:43:01learning process it understands the
  1770. 1:43:03positioning or overall global context.
  1771. 1:43:06So when I say global context what does
  1772. 1:43:08it means? So it means that if given an
  1773. 1:43:11input it is trying to understand okay
  1774. 1:43:13what are the foreground what are the
  1775. 1:43:14background what are the information
  1776. 1:43:16which is present in very much in the
  1777. 1:43:19near infrared uh regions what are the
  1778. 1:43:22textural informations so these kind of
  1779. 1:43:25uh the contextual is the texture is fine
  1780. 1:43:28or granular. So all these information
  1781. 1:43:31gets combined together in its latent
  1782. 1:43:33space of uh 1024 bands. So you can think
  1783. 1:43:37of that you are just providing a four
  1784. 1:43:39band input images or maybe 12 band input
  1785. 1:43:41images and finally you get 1024 uh bands
  1786. 1:43:46output which has information not only in
  1787. 1:43:49the spectral reason but it also in the
  1788. 1:43:51spatial relationships and so on. So
  1789. 1:43:54accordingly you get a very large latent
  1790. 1:43:56space. So accordingly and that's what
  1791. 1:43:58you can use it for different kind of uh
  1792. 1:44:00classifications categorizations which
  1793. 1:44:02can be used and uh that's how uh we can
  1794. 1:44:06use it for different uh
  1795. 1:44:09cases I think wit and other uh I think I
  1796. 1:44:12have already talked about so I think I
  1797. 1:44:14just took all the questions so probably
  1798. 1:44:17almost I was able to answer all of those
  1799. 1:44:20uh questions so uh hersel I think we
  1800. 1:44:25today we didn't discuss anything related
  1801. 1:44:27to vector raers. So probably you you
  1802. 1:44:31need to brush up and have a look what
  1803. 1:44:33exactly we discussed. Anyway, so thank
  1804. 1:44:36you for uh uh joining these sessions.
  1805. 1:44:39Have a nice day. Bye-bye.

About this transcript

This page contains the full transcript of Foundation Models for Geodata Processing by Dr. Ashutosh Kumar Jha by IIRS ISRO Digital Learning Programme, generated from the public captions YouTube serves with the video. The transcript has 12,906 words across 1,805 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.