YouTube2Text

Sensing Color | Image Sensing — Transcript

by First Principles of Computer Vision · 2,810 words · 406 segments · language en · Watch on YouTube

Full transcript

  1. 0:02now let's take a look at what it means
  2. 0:04to measure color using an image sensor
  3. 0:08we know that color has to do with the
  4. 0:10wavelength of light so let's assume that
  5. 0:13the incoming flux from a particular
  6. 0:15point in the scene is p of Lambda which
  7. 0:17is a photon flux as a function of
  8. 0:20wavelength that Photon flux arrives at
  9. 0:23your silicon pixel and the pixel
  10. 0:25converts it into electron flux I and the
  11. 0:30quantum efficiency of the material that
  12. 0:33you use to convert photons to electrons
  13. 0:37is defined as the ratio of the electron
  14. 0:41flux generated by the material to the
  15. 0:44photon flux that is incident on the
  16. 0:46material for a given wavelength of light
  17. 0:50so the quantum efficiency is a function
  18. 0:52of wavelength it's called Q of Lambda so
  19. 0:55let's take a look at what the quantum
  20. 0:56efficiency is for silicon so in the
  21. 0:59scale of silicon when you go to High
  22. 1:01wavelengths remember that visible light
  23. 1:03is between 400 nanom and 700 nanom now
  24. 1:07when you go to somewhere around 1,000
  25. 1:09nanometers it turns out that the Silicon
  26. 1:12silicon's Quantum efficiency is one it's
  27. 1:16almost perfect meaning to say that every
  28. 1:18Photon that it receives of that
  29. 1:20wavelength gets converted to an
  30. 1:23electron but this Quantum efficiency
  31. 1:25begins to fall as you go down in
  32. 1:27wavelength and somewhere around 400
  33. 1:30nanom it drops to almost zero so one can
  34. 1:34say that silicon is virtually
  35. 1:36transparent for wavelengths above 1,000
  36. 1:39or so nanometers and it gets almost
  37. 1:42opaque for wavelengths below 400
  38. 1:45[Music]
  39. 1:46nanom so now let's take a look at what
  40. 1:49happens when you have a single
  41. 1:50wavelength of light that's Lambda is
  42. 1:52equal to Lambda I this is called
  43. 1:54monochromatic light and let's say that
  44. 1:57the flux is p Lambda I
  45. 2:00so we know from the definition of
  46. 2:02quantum efficiency that the electron
  47. 2:04flux I is equal to Q Lambda i p Lambda I
  48. 2:09but now we know that for any point in
  49. 2:10the scene you have a range of
  50. 2:12wavelengths coming from it that it
  51. 2:14reflects light with Photon flux which
  52. 2:17varies as a function of Lambda this is
  53. 2:20called a spectral distribution P
  54. 2:23Lambda so in this case what is the
  55. 2:27electron flux going to be well let's
  56. 2:30take a look at a narrow band of
  57. 2:33wavelengths which goes from Lambda to
  58. 2:37Lambda + D Lambda so in this case the
  59. 2:41flux that's arriving at the pixel is p
  60. 2:45Lambda D Lambda and if you want to find
  61. 2:48the electron flux generated by the pixel
  62. 2:51for the entire spectral distribution
  63. 2:54it's nothing but Q Lambda P Lambda D
  64. 2:57Lambda integrated from 0 to Infinity
  65. 3:00that's the total number of electrons
  66. 3:03that you're going to get for the
  67. 3:04incoming
  68. 3:06light so now the question is if I give
  69. 3:09you I which is the electron flux you of
  70. 3:12course know the quantum efficiency Q of
  71. 3:14Lambda can you tell me what P of Lambda
  72. 3:17is this function P of Lambda the
  73. 3:18spectral
  74. 3:20distribution well the answer is no
  75. 3:23because there are many P lambdas that I
  76. 3:25can multiply with Q Lambda to get the
  77. 3:27same I
  78. 3:30so how do we measure P of Lambda well we
  79. 3:32use filters so in front of your pixel
  80. 3:35you're going to place a filter fi of
  81. 3:37Lambda and then the question is what
  82. 3:38filter should you use to get all of P
  83. 3:42Lambda well what we can do is use a
  84. 3:46filter which is narrow band that is make
  85. 3:48sure that f i Lambda is a Delta function
  86. 3:53positioned as Lambda I what is a Delta
  87. 3:56function it's infinitesimally thin and
  88. 3:59infinitely allall but it has an area
  89. 4:02equal to 1 that's a Delta function for
  90. 4:04you right here Delta Lambda minus Lambda
  91. 4:07I is a Delta function centered at Lambda
  92. 4:11I so now what do you get you get I equal
  93. 4:14to Q Lambda P Lambda fi Lambda D Lambda
  94. 4:20integral and we know that fi Lambda is
  95. 4:23Delta Lambda minus Lambda I it's a Delta
  96. 4:27function and what this gives you is the
  97. 4:30value of Q Lambda P Lambda at Lambda I
  98. 4:34this is called the sifting property of a
  99. 4:36Delta function it basically picks out
  100. 4:39that value wherever it's centered of the
  101. 4:41rest of the function so you get Q Lambda
  102. 4:43i p Lambda I isal to I so you're able to
  103. 4:47read out the value of P Lambda at Lambda
  104. 4:50I and now the question is how many
  105. 4:53filters do you need to recover all of P
  106. 4:56Lambda
  107. 4:58well first glance you would say an
  108. 5:01infinite number of uh filters but it
  109. 5:05turns out that you don't need an
  110. 5:06infinite number of filters if P Lambda
  111. 5:09is a smooth function we will see later
  112. 5:11in our image processing lectures that
  113. 5:13you can use a smaller number of
  114. 5:15measurements to recover P Lambda without
  115. 5:18losing any
  116. 5:22information so now that brings us to the
  117. 5:25topic of color what is color it turns
  118. 5:28out that color is not a physical
  119. 5:30quantity that you measure it is human
  120. 5:34response to different wavelengths of
  121. 5:36light so as I mentioned earlier the
  122. 5:40visible light spectrum the light that is
  123. 5:42visible to us lies between 400 nanom
  124. 5:45which is bluish velsh to 700 nanom which
  125. 5:50is red anything that is lower than 400
  126. 5:53Nom or right next to 400 nanom is called
  127. 5:56ultraviolet light we cannot see it and
  128. 5:59any anything that's beyond red 700 nanom
  129. 6:02is infrared light again we can't see it
  130. 6:05now this actually brings up an
  131. 6:07interesting point which is we can design
  132. 6:09cameras to measure information in the
  133. 6:12ultra violet Spectrum or infrared
  134. 6:14Spectrum as well in other words we can
  135. 6:17design computer vision systems that go
  136. 6:20well beyond the visible light spectrum
  137. 6:23and therefore are able to perceive
  138. 6:25things that you and I simply cannot
  139. 6:26perceive that's one of the advantages of
  140. 6:29using
  141. 6:30uh computer vision it allows us to walk
  142. 6:33into visual worlds right next to our
  143. 6:36visual world but but cannot be perceived
  144. 6:39by us so that's the visible light
  145. 6:42spectrum and so now the question is do
  146. 6:44we humans actually recover the spectral
  147. 6:47distribution P of Lambda within the spe
  148. 6:50the visible light spectrum for any given
  149. 6:52point well it turns out the human eye
  150. 6:56the sensor of the human eye which is the
  151. 6:58retina
  152. 7:00has two types of pixels rods and
  153. 7:04cones and the rods are not really
  154. 7:07sensitive to the color of light they're
  155. 7:09just looking for the brightness of the
  156. 7:11incoming light the cones are sensitive
  157. 7:14to color and they're essentially three
  158. 7:17types of cones these are neurochemical
  159. 7:20sensors that respond to three types of
  160. 7:23of color so to speak so let's take a
  161. 7:26look at these rods and cones so here you
  162. 7:28have the human ey we've seen this before
  163. 7:32you have your Cora you have your lens
  164. 7:35and the lens forms an image on the
  165. 7:37retina which is sitting back
  166. 7:39here and the retina like I said earlier
  167. 7:43also does a little bit of early
  168. 7:45processing but this is where all your
  169. 7:47pixels are sitting so let's take a look
  170. 7:49at the architecture if you will of the
  171. 7:52retina you see that you have your rods
  172. 7:54and cones or your pixels right here and
  173. 7:57then this information goes into
  174. 8:00bipolar cells which are sitting right
  175. 8:02here and that goes into gangon cells
  176. 8:05right here so there's a little bit of
  177. 8:07visual processing early visual
  178. 8:09processing happening and then that image
  179. 8:12that semi-processed image is passed
  180. 8:14through the optic nerve to the visual
  181. 8:17cortex in the brain so question for you
  182. 8:21where is light coming from which
  183. 8:22direction is light coming from in this
  184. 8:25diagram well you would assume that if
  185. 8:27the pixels are sitting here light really
  186. 8:29should be coming from the bottom it
  187. 8:31turns out that's not the case light is
  188. 8:34actually coming from the top and passing
  189. 8:37through the ganglion and and bipolar
  190. 8:39cells to get to your rods and cones
  191. 8:43interesting fact not exactly show why
  192. 8:46nature decided to to design the retina
  193. 8:49that way now let's come back to the rods
  194. 8:52and cones so this is a scanning electron
  195. 8:54microscope image of rods and cones you
  196. 8:57can see that the rods actually look like
  197. 8:59rods they're cylindrical more or less
  198. 9:02and the cones look like cones and that's
  199. 9:04why they're called
  200. 9:05so right here and they essentially do
  201. 9:10the same thing they take light in and
  202. 9:12they're able to generate a act
  203. 9:16activation an Impulse that goes off to
  204. 9:19the brain however they respond to
  205. 9:23different types of light so to speak so
  206. 9:26they have different proteins rods have r
  207. 9:29opsin as a protein and cones have
  208. 9:33photopsin as the
  209. 9:35protein and so rods are able to measure
  210. 9:39what I would call Black and whitish
  211. 9:41images they come into effect
  212. 9:43particularly when you're walking out in
  213. 9:45darkness when there's very little light
  214. 9:47say Moonlight say for instance you may
  215. 9:50have noticed that you're not able to
  216. 9:51actually discern the colors of things in
  217. 9:53the scene you just know whether there's
  218. 9:56light or no light so you get a very dim
  219. 9:59IM image which is
  220. 10:01colorless the cones when there's plenty
  221. 10:03of light are able to discern the colors
  222. 10:06of various things in the scene and so
  223. 10:10Vision using rods is called scotopic
  224. 10:13vision and vision using cones is called
  225. 10:16photopic Vision these are two different
  226. 10:18modes that the eye operates in and they
  227. 10:22PR once they measure something in terms
  228. 10:24of photon flux they send this as nerve
  229. 10:28impulses down around the optic nerve for
  230. 10:31further
  231. 10:33processing so now let's take a look at
  232. 10:36the distribution of cones particularly
  233. 10:39let's set aside rods for a minute you
  234. 10:41have essentially three types of cones
  235. 10:44red cones green cones and blue cones we
  236. 10:47call them these because they respond to
  237. 10:50reddish light greenish light and bluish
  238. 10:52light and you see their spatial
  239. 10:55Distribution on the retina and it turns
  240. 10:58out that they are maximally most dense
  241. 11:01they have maximum resolution in the
  242. 11:04phobia which is the place where you have
  243. 11:07maximum accurity when you look at
  244. 11:09something in the scene that something is
  245. 11:11falling on the fobia and everything else
  246. 11:13is peripheral vision so that's where you
  247. 11:16have maximum number of rods and cones in
  248. 11:19terms of relative numbers of rods and
  249. 11:21cones you have roughly 120 million rods
  250. 11:26versus only about 7 million cones on the
  251. 11:30retina and most of your cones like I
  252. 11:32said are in the fobal region right here
  253. 11:36very high density and then falling off
  254. 11:38very fast and then when it comes to the
  255. 11:41rods you have almost no rods in the
  256. 11:43center of the fobia and then you have
  257. 11:45lots of rods very high density and then
  258. 11:48the resolution begins to fall off as you
  259. 11:50go to the periphery of the field of
  260. 11:54view and right here you see that you
  261. 11:57don't have any rods of
  262. 12:00cones any idea what that is well that's
  263. 12:02the blind spot that's where the image
  264. 12:06goes through the optic nerve back to the
  265. 12:08brain and that's a place where you don't
  266. 12:11have any rods and cones in the eye now
  267. 12:13you must be wondering why is it that
  268. 12:15you're not seeing your blind spot all
  269. 12:17the time it turns out that your brain is
  270. 12:19filling in visual information in that
  271. 12:21region and so we believe we have an
  272. 12:24absolutely continuous image but it
  273. 12:27actually has a hole in it as we'll see
  274. 12:30shortly so now let's talk about the
  275. 12:33spectral response just like we talked
  276. 12:36about the quantum efficiency of silicon
  277. 12:38we can talk about the spectral responses
  278. 12:40of the cones the three types of cones so
  279. 12:43the spectral responses are called the
  280. 12:46tri stimulus curves so think about these
  281. 12:48as exactly like the quantum
  282. 12:51efficiencies efficiency of silicon but
  283. 12:53these are the quantum efficiencies of
  284. 12:56the red green and blue cones
  285. 13:00and so you have HB Lambda HG Lambda and
  286. 13:04HR
  287. 13:06Lambda and so now what do these cones
  288. 13:09measure well very much as in the case of
  289. 13:12our silicon pixel you have HR Lambda P
  290. 13:16Lambda P Lambda again is a spectral
  291. 13:18distribution of the incoming light the
  292. 13:21integral of that that gives you one
  293. 13:23number R similarly you get G and you get
  294. 13:27B so you get these three numbers
  295. 13:29corresponding to any incoming spectral
  296. 13:31distribution you don't get the complete
  297. 13:33spectral
  298. 13:34distribution so so we see here that the
  299. 13:38I does not measure the complete spectral
  300. 13:42distribution P of Lambda it basically
  301. 13:44gives you three numbers corresponding to
  302. 13:46any P of Lambda and so that brings us
  303. 13:49back to the old problem which is that
  304. 13:52there are multiple P lambdas there's an
  305. 13:55entire Continuum of P lambdas that
  306. 13:57actually generate the same RGB values
  307. 14:00and these are called metamers so here
  308. 14:03you're seeing different P
  309. 14:06lambdas significantly different P
  310. 14:08lambdas which actually produce the same
  311. 14:11RGB values in the ey and in fact the
  312. 14:14values at De they produce are roughly
  313. 14:17115 60 and 108 and the the the color
  314. 14:21that we perceive is this color here
  315. 14:23which is purplish purple or violet or
  316. 14:26whatever that color is magenta perhaps
  317. 14:29that's the color we perceive for all of
  318. 14:32these spectral distributions so the
  319. 14:34existence of these metamers is telling
  320. 14:36you that there's a lot of spectral
  321. 14:38distributions that exist in the world
  322. 14:41which are very different from one
  323. 14:42another which we humans perceived to be
  324. 14:45the same
  325. 14:47color so that brings us to Young's
  326. 14:50experiment on mixing colors so the what
  327. 14:54Young found out is that you can take
  328. 14:57three wavelengths of light just three
  329. 14:59wavelengths specific wavelengths of
  330. 15:01light and you can mix them in this case
  331. 15:05he's using a different projector for
  332. 15:08each wavelength of light and you can see
  333. 15:10that wherever the circles the projected
  334. 15:13circles overlap you're getting these
  335. 15:15different colors and it turns out that
  336. 15:17just with these three wavelengths you
  337. 15:19can actually reproduce the sensation of
  338. 15:23virtually all the colors that you and I
  339. 15:25are able to see and so these three uh
  340. 15:29colors are wavelengths are
  341. 15:32650 530 and 410 nanom now of course you
  342. 15:36can change these wavelengths and you
  343. 15:38would get very similar effects but three
  344. 15:40wavelengths of light are enough to
  345. 15:42produce virtually all the colors that
  346. 15:45you and I are able to perceive and this
  347. 15:47is why cameras and displays have red
  348. 15:52green and blue filters on
  349. 15:54them so let's take a look at how we can
  350. 15:57capture a color image so one way to do
  351. 15:59this is by using What's called the
  352. 16:01dicro prism this prism is a fairly
  353. 16:06sophisticated uh piece of
  354. 16:08Optics and when you show it an image
  355. 16:12that image gets split into three
  356. 16:15components a reddish image a bluish
  357. 16:18image and a greenish
  358. 16:19image and so now what you can do is you
  359. 16:22can place a sensor on these three faces
  360. 16:25that you're seeing right here of this
  361. 16:27dichroic prism and this sensor would
  362. 16:30capture a red image uh a green image and
  363. 16:34this one would capture a blue image
  364. 16:36these are not the images are not red
  365. 16:38green and blue they're capturing those
  366. 16:40wavelengths of light and they're
  367. 16:42perfectly aligned so you can now stack
  368. 16:44these three images and you essentially
  369. 16:45have red green and blue values at each
  370. 16:49pixel so that's one way to do it albe it
  371. 16:52a little bit uh bulky and a little bit
  372. 16:56um advanced in terms of requiring
  373. 16:59alignment and precision and so on it
  374. 17:02would be nice to be able to measure
  375. 17:03color using a single chip Just One
  376. 17:06sensor and that's what's usually done at
  377. 17:08least in consumer cameras and that's
  378. 17:11done using what's called a color Mosaic
  379. 17:14so in this case again you see pixels but
  380. 17:17you have different color uh filters in
  381. 17:20front of each pixel these are the red
  382. 17:23pixels and these are the green pixels
  383. 17:25and the blue pixels so any given pixel
  384. 17:29only measures one particular color and
  385. 17:33so at any location in the image you're
  386. 17:35only measuring one color but you know
  387. 17:39that you have neighbors that measure the
  388. 17:41other colors and so the basic idea is
  389. 17:43you take an image which has red green
  390. 17:47and blue
  391. 17:48values as shown here so each pixel is
  392. 17:51either red green and blue and then you
  393. 17:54use interpolation which is a technique
  394. 17:56we'll talk about later to Recon
  395. 17:59construct the full red green blue image
  396. 18:02the color image so in other words if
  397. 18:05you've measured blue at this point you
  398. 18:08then say given my neighbors my red and
  399. 18:12green neighbors what red and green
  400. 18:15values would I have measured at this
  401. 18:17point if I had a green pixel there and a
  402. 18:20red pixel there so you end up with your
  403. 18:22red your color image this process of
  404. 18:26interpolating this uh Mosaic image to
  405. 18:30get a full color image is called
  406. 18:32demosaicing

About this transcript

This page contains the full transcript of Sensing Color | Image Sensing by First Principles of Computer Vision, generated from the public captions YouTube serves with the video. The transcript has 2,810 words across 406 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.