Sensing Color | Image Sensing — Transcript
Full transcript
- 0:02now let's take a look at what it means
- 0:04to measure color using an image sensor
- 0:08we know that color has to do with the
- 0:10wavelength of light so let's assume that
- 0:13the incoming flux from a particular
- 0:15point in the scene is p of Lambda which
- 0:17is a photon flux as a function of
- 0:20wavelength that Photon flux arrives at
- 0:23your silicon pixel and the pixel
- 0:25converts it into electron flux I and the
- 0:30quantum efficiency of the material that
- 0:33you use to convert photons to electrons
- 0:37is defined as the ratio of the electron
- 0:41flux generated by the material to the
- 0:44photon flux that is incident on the
- 0:46material for a given wavelength of light
- 0:50so the quantum efficiency is a function
- 0:52of wavelength it's called Q of Lambda so
- 0:55let's take a look at what the quantum
- 0:56efficiency is for silicon so in the
- 0:59scale of silicon when you go to High
- 1:01wavelengths remember that visible light
- 1:03is between 400 nanom and 700 nanom now
- 1:07when you go to somewhere around 1,000
- 1:09nanometers it turns out that the Silicon
- 1:12silicon's Quantum efficiency is one it's
- 1:16almost perfect meaning to say that every
- 1:18Photon that it receives of that
- 1:20wavelength gets converted to an
- 1:23electron but this Quantum efficiency
- 1:25begins to fall as you go down in
- 1:27wavelength and somewhere around 400
- 1:30nanom it drops to almost zero so one can
- 1:34say that silicon is virtually
- 1:36transparent for wavelengths above 1,000
- 1:39or so nanometers and it gets almost
- 1:42opaque for wavelengths below 400
- 1:45[Music]
- 1:46nanom so now let's take a look at what
- 1:49happens when you have a single
- 1:50wavelength of light that's Lambda is
- 1:52equal to Lambda I this is called
- 1:54monochromatic light and let's say that
- 1:57the flux is p Lambda I
- 2:00so we know from the definition of
- 2:02quantum efficiency that the electron
- 2:04flux I is equal to Q Lambda i p Lambda I
- 2:09but now we know that for any point in
- 2:10the scene you have a range of
- 2:12wavelengths coming from it that it
- 2:14reflects light with Photon flux which
- 2:17varies as a function of Lambda this is
- 2:20called a spectral distribution P
- 2:23Lambda so in this case what is the
- 2:27electron flux going to be well let's
- 2:30take a look at a narrow band of
- 2:33wavelengths which goes from Lambda to
- 2:37Lambda + D Lambda so in this case the
- 2:41flux that's arriving at the pixel is p
- 2:45Lambda D Lambda and if you want to find
- 2:48the electron flux generated by the pixel
- 2:51for the entire spectral distribution
- 2:54it's nothing but Q Lambda P Lambda D
- 2:57Lambda integrated from 0 to Infinity
- 3:00that's the total number of electrons
- 3:03that you're going to get for the
- 3:04incoming
- 3:06light so now the question is if I give
- 3:09you I which is the electron flux you of
- 3:12course know the quantum efficiency Q of
- 3:14Lambda can you tell me what P of Lambda
- 3:17is this function P of Lambda the
- 3:18spectral
- 3:20distribution well the answer is no
- 3:23because there are many P lambdas that I
- 3:25can multiply with Q Lambda to get the
- 3:27same I
- 3:30so how do we measure P of Lambda well we
- 3:32use filters so in front of your pixel
- 3:35you're going to place a filter fi of
- 3:37Lambda and then the question is what
- 3:38filter should you use to get all of P
- 3:42Lambda well what we can do is use a
- 3:46filter which is narrow band that is make
- 3:48sure that f i Lambda is a Delta function
- 3:53positioned as Lambda I what is a Delta
- 3:56function it's infinitesimally thin and
- 3:59infinitely allall but it has an area
- 4:02equal to 1 that's a Delta function for
- 4:04you right here Delta Lambda minus Lambda
- 4:07I is a Delta function centered at Lambda
- 4:11I so now what do you get you get I equal
- 4:14to Q Lambda P Lambda fi Lambda D Lambda
- 4:20integral and we know that fi Lambda is
- 4:23Delta Lambda minus Lambda I it's a Delta
- 4:27function and what this gives you is the
- 4:30value of Q Lambda P Lambda at Lambda I
- 4:34this is called the sifting property of a
- 4:36Delta function it basically picks out
- 4:39that value wherever it's centered of the
- 4:41rest of the function so you get Q Lambda
- 4:43i p Lambda I isal to I so you're able to
- 4:47read out the value of P Lambda at Lambda
- 4:50I and now the question is how many
- 4:53filters do you need to recover all of P
- 4:56Lambda
- 4:58well first glance you would say an
- 5:01infinite number of uh filters but it
- 5:05turns out that you don't need an
- 5:06infinite number of filters if P Lambda
- 5:09is a smooth function we will see later
- 5:11in our image processing lectures that
- 5:13you can use a smaller number of
- 5:15measurements to recover P Lambda without
- 5:18losing any
- 5:22information so now that brings us to the
- 5:25topic of color what is color it turns
- 5:28out that color is not a physical
- 5:30quantity that you measure it is human
- 5:34response to different wavelengths of
- 5:36light so as I mentioned earlier the
- 5:40visible light spectrum the light that is
- 5:42visible to us lies between 400 nanom
- 5:45which is bluish velsh to 700 nanom which
- 5:50is red anything that is lower than 400
- 5:53Nom or right next to 400 nanom is called
- 5:56ultraviolet light we cannot see it and
- 5:59any anything that's beyond red 700 nanom
- 6:02is infrared light again we can't see it
- 6:05now this actually brings up an
- 6:07interesting point which is we can design
- 6:09cameras to measure information in the
- 6:12ultra violet Spectrum or infrared
- 6:14Spectrum as well in other words we can
- 6:17design computer vision systems that go
- 6:20well beyond the visible light spectrum
- 6:23and therefore are able to perceive
- 6:25things that you and I simply cannot
- 6:26perceive that's one of the advantages of
- 6:29using
- 6:30uh computer vision it allows us to walk
- 6:33into visual worlds right next to our
- 6:36visual world but but cannot be perceived
- 6:39by us so that's the visible light
- 6:42spectrum and so now the question is do
- 6:44we humans actually recover the spectral
- 6:47distribution P of Lambda within the spe
- 6:50the visible light spectrum for any given
- 6:52point well it turns out the human eye
- 6:56the sensor of the human eye which is the
- 6:58retina
- 7:00has two types of pixels rods and
- 7:04cones and the rods are not really
- 7:07sensitive to the color of light they're
- 7:09just looking for the brightness of the
- 7:11incoming light the cones are sensitive
- 7:14to color and they're essentially three
- 7:17types of cones these are neurochemical
- 7:20sensors that respond to three types of
- 7:23of color so to speak so let's take a
- 7:26look at these rods and cones so here you
- 7:28have the human ey we've seen this before
- 7:32you have your Cora you have your lens
- 7:35and the lens forms an image on the
- 7:37retina which is sitting back
- 7:39here and the retina like I said earlier
- 7:43also does a little bit of early
- 7:45processing but this is where all your
- 7:47pixels are sitting so let's take a look
- 7:49at the architecture if you will of the
- 7:52retina you see that you have your rods
- 7:54and cones or your pixels right here and
- 7:57then this information goes into
- 8:00bipolar cells which are sitting right
- 8:02here and that goes into gangon cells
- 8:05right here so there's a little bit of
- 8:07visual processing early visual
- 8:09processing happening and then that image
- 8:12that semi-processed image is passed
- 8:14through the optic nerve to the visual
- 8:17cortex in the brain so question for you
- 8:21where is light coming from which
- 8:22direction is light coming from in this
- 8:25diagram well you would assume that if
- 8:27the pixels are sitting here light really
- 8:29should be coming from the bottom it
- 8:31turns out that's not the case light is
- 8:34actually coming from the top and passing
- 8:37through the ganglion and and bipolar
- 8:39cells to get to your rods and cones
- 8:43interesting fact not exactly show why
- 8:46nature decided to to design the retina
- 8:49that way now let's come back to the rods
- 8:52and cones so this is a scanning electron
- 8:54microscope image of rods and cones you
- 8:57can see that the rods actually look like
- 8:59rods they're cylindrical more or less
- 9:02and the cones look like cones and that's
- 9:04why they're called
- 9:05so right here and they essentially do
- 9:10the same thing they take light in and
- 9:12they're able to generate a act
- 9:16activation an Impulse that goes off to
- 9:19the brain however they respond to
- 9:23different types of light so to speak so
- 9:26they have different proteins rods have r
- 9:29opsin as a protein and cones have
- 9:33photopsin as the
- 9:35protein and so rods are able to measure
- 9:39what I would call Black and whitish
- 9:41images they come into effect
- 9:43particularly when you're walking out in
- 9:45darkness when there's very little light
- 9:47say Moonlight say for instance you may
- 9:50have noticed that you're not able to
- 9:51actually discern the colors of things in
- 9:53the scene you just know whether there's
- 9:56light or no light so you get a very dim
- 9:59IM image which is
- 10:01colorless the cones when there's plenty
- 10:03of light are able to discern the colors
- 10:06of various things in the scene and so
- 10:10Vision using rods is called scotopic
- 10:13vision and vision using cones is called
- 10:16photopic Vision these are two different
- 10:18modes that the eye operates in and they
- 10:22PR once they measure something in terms
- 10:24of photon flux they send this as nerve
- 10:28impulses down around the optic nerve for
- 10:31further
- 10:33processing so now let's take a look at
- 10:36the distribution of cones particularly
- 10:39let's set aside rods for a minute you
- 10:41have essentially three types of cones
- 10:44red cones green cones and blue cones we
- 10:47call them these because they respond to
- 10:50reddish light greenish light and bluish
- 10:52light and you see their spatial
- 10:55Distribution on the retina and it turns
- 10:58out that they are maximally most dense
- 11:01they have maximum resolution in the
- 11:04phobia which is the place where you have
- 11:07maximum accurity when you look at
- 11:09something in the scene that something is
- 11:11falling on the fobia and everything else
- 11:13is peripheral vision so that's where you
- 11:16have maximum number of rods and cones in
- 11:19terms of relative numbers of rods and
- 11:21cones you have roughly 120 million rods
- 11:26versus only about 7 million cones on the
- 11:30retina and most of your cones like I
- 11:32said are in the fobal region right here
- 11:36very high density and then falling off
- 11:38very fast and then when it comes to the
- 11:41rods you have almost no rods in the
- 11:43center of the fobia and then you have
- 11:45lots of rods very high density and then
- 11:48the resolution begins to fall off as you
- 11:50go to the periphery of the field of
- 11:54view and right here you see that you
- 11:57don't have any rods of
- 12:00cones any idea what that is well that's
- 12:02the blind spot that's where the image
- 12:06goes through the optic nerve back to the
- 12:08brain and that's a place where you don't
- 12:11have any rods and cones in the eye now
- 12:13you must be wondering why is it that
- 12:15you're not seeing your blind spot all
- 12:17the time it turns out that your brain is
- 12:19filling in visual information in that
- 12:21region and so we believe we have an
- 12:24absolutely continuous image but it
- 12:27actually has a hole in it as we'll see
- 12:30shortly so now let's talk about the
- 12:33spectral response just like we talked
- 12:36about the quantum efficiency of silicon
- 12:38we can talk about the spectral responses
- 12:40of the cones the three types of cones so
- 12:43the spectral responses are called the
- 12:46tri stimulus curves so think about these
- 12:48as exactly like the quantum
- 12:51efficiencies efficiency of silicon but
- 12:53these are the quantum efficiencies of
- 12:56the red green and blue cones
- 13:00and so you have HB Lambda HG Lambda and
- 13:04HR
- 13:06Lambda and so now what do these cones
- 13:09measure well very much as in the case of
- 13:12our silicon pixel you have HR Lambda P
- 13:16Lambda P Lambda again is a spectral
- 13:18distribution of the incoming light the
- 13:21integral of that that gives you one
- 13:23number R similarly you get G and you get
- 13:27B so you get these three numbers
- 13:29corresponding to any incoming spectral
- 13:31distribution you don't get the complete
- 13:33spectral
- 13:34distribution so so we see here that the
- 13:38I does not measure the complete spectral
- 13:42distribution P of Lambda it basically
- 13:44gives you three numbers corresponding to
- 13:46any P of Lambda and so that brings us
- 13:49back to the old problem which is that
- 13:52there are multiple P lambdas there's an
- 13:55entire Continuum of P lambdas that
- 13:57actually generate the same RGB values
- 14:00and these are called metamers so here
- 14:03you're seeing different P
- 14:06lambdas significantly different P
- 14:08lambdas which actually produce the same
- 14:11RGB values in the ey and in fact the
- 14:14values at De they produce are roughly
- 14:17115 60 and 108 and the the the color
- 14:21that we perceive is this color here
- 14:23which is purplish purple or violet or
- 14:26whatever that color is magenta perhaps
- 14:29that's the color we perceive for all of
- 14:32these spectral distributions so the
- 14:34existence of these metamers is telling
- 14:36you that there's a lot of spectral
- 14:38distributions that exist in the world
- 14:41which are very different from one
- 14:42another which we humans perceived to be
- 14:45the same
- 14:47color so that brings us to Young's
- 14:50experiment on mixing colors so the what
- 14:54Young found out is that you can take
- 14:57three wavelengths of light just three
- 14:59wavelengths specific wavelengths of
- 15:01light and you can mix them in this case
- 15:05he's using a different projector for
- 15:08each wavelength of light and you can see
- 15:10that wherever the circles the projected
- 15:13circles overlap you're getting these
- 15:15different colors and it turns out that
- 15:17just with these three wavelengths you
- 15:19can actually reproduce the sensation of
- 15:23virtually all the colors that you and I
- 15:25are able to see and so these three uh
- 15:29colors are wavelengths are
- 15:32650 530 and 410 nanom now of course you
- 15:36can change these wavelengths and you
- 15:38would get very similar effects but three
- 15:40wavelengths of light are enough to
- 15:42produce virtually all the colors that
- 15:45you and I are able to perceive and this
- 15:47is why cameras and displays have red
- 15:52green and blue filters on
- 15:54them so let's take a look at how we can
- 15:57capture a color image so one way to do
- 15:59this is by using What's called the
- 16:01dicro prism this prism is a fairly
- 16:06sophisticated uh piece of
- 16:08Optics and when you show it an image
- 16:12that image gets split into three
- 16:15components a reddish image a bluish
- 16:18image and a greenish
- 16:19image and so now what you can do is you
- 16:22can place a sensor on these three faces
- 16:25that you're seeing right here of this
- 16:27dichroic prism and this sensor would
- 16:30capture a red image uh a green image and
- 16:34this one would capture a blue image
- 16:36these are not the images are not red
- 16:38green and blue they're capturing those
- 16:40wavelengths of light and they're
- 16:42perfectly aligned so you can now stack
- 16:44these three images and you essentially
- 16:45have red green and blue values at each
- 16:49pixel so that's one way to do it albe it
- 16:52a little bit uh bulky and a little bit
- 16:56um advanced in terms of requiring
- 16:59alignment and precision and so on it
- 17:02would be nice to be able to measure
- 17:03color using a single chip Just One
- 17:06sensor and that's what's usually done at
- 17:08least in consumer cameras and that's
- 17:11done using what's called a color Mosaic
- 17:14so in this case again you see pixels but
- 17:17you have different color uh filters in
- 17:20front of each pixel these are the red
- 17:23pixels and these are the green pixels
- 17:25and the blue pixels so any given pixel
- 17:29only measures one particular color and
- 17:33so at any location in the image you're
- 17:35only measuring one color but you know
- 17:39that you have neighbors that measure the
- 17:41other colors and so the basic idea is
- 17:43you take an image which has red green
- 17:47and blue
- 17:48values as shown here so each pixel is
- 17:51either red green and blue and then you
- 17:54use interpolation which is a technique
- 17:56we'll talk about later to Recon
- 17:59construct the full red green blue image
- 18:02the color image so in other words if
- 18:05you've measured blue at this point you
- 18:08then say given my neighbors my red and
- 18:12green neighbors what red and green
- 18:15values would I have measured at this
- 18:17point if I had a green pixel there and a
- 18:20red pixel there so you end up with your
- 18:22red your color image this process of
- 18:26interpolating this uh Mosaic image to
- 18:30get a full color image is called
- 18:32demosaicing
About this transcript
This page contains the full transcript of Sensing Color | Image Sensing by First Principles of Computer Vision, generated from the public captions YouTube serves with the video. The transcript has 2,810 words across 406 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.