GCI World 2026 September Session2 During Lecture — Transcript
Full transcript
- 0:00We will begin lecture two on a
- 0:25manage your data in Python. Then we will
- 0:28look at onedimensional.
- 0:31>> Just realized
- 0:33that you might
- 0:35not hearing the audio.
- 0:39Hold on a second. Sorry.
- 0:49How about this now?
- 0:57I think you can.
- 0:58>> We will begin lecture two on efficient
- 1:00data manipulation using non-fi.
- 1:03Throughout the next three lectures, we
- 1:06will look at different types of Python
- 1:08libraries that are often used for data
- 1:10science. This lecture will specifically
- 1:13focus on the NumPy library.
- 1:16First, I will introduce how to use NumPy
- 1:18to create ND array, a convenient and
- 1:21useful way to store and manage your data
- 1:24in Python. Then we will look at
- 1:26one-dimensional arrays to learn how
- 1:28numpy behaves such as universal
- 1:31functions that enable elementwise
- 1:33calculations without loops, broadcasting
- 1:36that allows operations on arrays of
- 1:38different shapes, aggregation functions
- 1:41to compute core statistics, and indexing
- 1:44that helps extracting individual values
- 1:46or entire subarrays for quick data
- 1:49reorganization or analysis. We will then
- 1:52look at the case of 2D arrays. When we
- 1:55extend to 2D, we will see how concepts
- 1:58such as axis and index specification
- 2:01become important.
- 2:03This class will focus on working with 1D
- 2:06and 2D arrays. We will not explore 3D or
- 2:10higher dimensional arrays this time and
- 2:12we'll skip more advanced topics like
- 2:15defining machine learning models or
- 2:17diving deep into linear algebra
- 2:19concepts.
- 2:21Before we get into Numpy, let's review
- 2:24some of the basics of Python grammar.
- 2:27First, variables were names that act as
- 2:30references to objects or values stored
- 2:32in memory. Whatever on the right side of
- 2:35the equals sign was stored into whatever
- 2:37was written on the left side of the
- 2:39sign, which can then be evaluated or
- 2:41referenced.
- 2:43Operators are symbols to indicate some
- 2:45kind of calculation between two
- 2:47variables. There are several types of
- 2:50operators such as arithmetic operators
- 2:52and comparison operators.
- 2:55Collections are a type of data structure
- 2:57that stores multiple values in some
- 3:00organized structure. Depending on the
- 3:02structure, it can be categorized into
- 3:05ordered management where order of the
- 3:07values matter such as lists and tpples
- 3:10and keyword management where order does
- 3:13not matter.
- 3:15In particular, a list is a type of data
- 3:18structure that stores multiple values in
- 3:21sequential order. Each value is assigned
- 3:24an index which indicates where the value
- 3:26is stored in that list.
- 3:28Two-dimensional lists can also be
- 3:30created as lists of lists. Similarly to
- 3:34lists, we can access specific values or
- 3:37elements by specifying the location
- 3:39using a process called indexing or
- 3:42slicing.
- 3:44Finally, for statements are a type of
- 3:46operation used to repeat the same
- 3:48process over multiple cases. What we
- 3:51loop over can be arbitrary. It can be
- 3:54some range, some list-like structure,
- 3:56and so on.
- 3:58Now, let's start by looking at what
- 4:00numpy is for and how to get started with
- 4:03it. Data mining is often described as a
- 4:06process to discover and extract valuable
- 4:09insights from large data sets. Christm
- 4:13is a framework which is used for smooth
- 4:15transition from understanding the
- 4:17business problem and the data itself
- 4:20through data preparation and modeling
- 4:23all the way to evaluation and expansion
- 4:26of our findings.
- 4:28The main goal here is to learn basic
- 4:30data manipulation using numpy so raw
- 4:33data can be transformed into a form that
- 4:36is ready for more advanced analyses
- 4:38later on.
- 4:40NumPy is a Python library suited for
- 4:43complex scientific computations. In
- 4:46Python, there are many kinds of
- 4:48libraries or collections of tools and
- 4:51functions that people can use to perform
- 4:53common tasks without having to write the
- 4:55code from scratch themselves. You can
- 4:58choose to use whichever library you want
- 5:00to use that best suits the task. Numpy
- 5:03is a library for handling large
- 5:05multi-dimensional arrays using a data
- 5:08structure named numpy.nd array using
- 5:11universal functions. It can handle
- 5:14intricate calculations without loops.
- 5:16Numpy is primarily written in C which
- 5:19makes its operations very fast.
- 5:22Additionally, NumPy also offers a broad
- 5:25range of functions for scientific
- 5:27computing. So plenty of tools are at our
- 5:30disposal for exploring and analyzing
- 5:32data. For example, we can simply write
- 5:35some equals a + b to add two arrays,
- 5:39which is much quicker than creating a
- 5:41new list, looping over every index, and
- 5:44appending the results. This simplicity
- 5:47means there will be less time to spend
- 5:49on repetitive tasks, and more time
- 5:51exploring the data.
- 5:54Now that we understand what numpy is,
- 5:56let's begin using it. As a convention,
- 5:59numpy is abbreviated as np, which makes
- 6:02the code shorter and easier to read. To
- 6:06import numpy into Python script, write
- 6:08import numpy as np. Using this
- 6:11abbreviation not only saves time when
- 6:14typing, but also helps keep the code
- 6:16clean and consistent.
- 6:19We can use NPI's functions by typing np
- 6:22followed by a dot and the function name
- 6:25then passing the object we want to
- 6:27process inside the parenthesis. For
- 6:29example, if we want to create an array,
- 6:32simply write array equals np.ray of 568.
- 6:38Here array is the function being used
- 6:41and 568 is the list converted into a
- 6:44numpy array. This straightforward syntax
- 6:47allows us to perform a wide range of
- 6:49operations on the data easily while
- 6:52making the code cleaner and more
- 6:54efficient.
- 6:56It's important to remember that
- 6:57functions are not something you need to
- 6:59memorize, but rather something you
- 7:01should look up when needed. Any library
- 7:04offers a vast array of functions to
- 7:06handle various tasks, and trying to
- 7:09remember all of them can be
- 7:10overwhelming. Instead, focus on
- 7:13understanding the core concepts and know
- 7:16how to efficiently find the functions
- 7:18you require. If there's a particular
- 7:20operation you'd like to try, don't
- 7:22hesitate to search for it in the
- 7:24official documentation or other reliable
- 7:27resources.
- 7:29Now, let's get into using numpy for 1D
- 7:31array.
- 7:37I'll just stop
- 7:39>> here moment and I think
- 7:44lot of you are facing that you cannot
- 7:48open this n book right I think yeah I
- 7:52wasn't able to open this book too so
- 7:56we're just figuring out um this issue at
- 8:01the moment so um
- 8:04out.
- 8:06Yeah. So, I'll just go back to the
- 8:07video, but um once that issue
- 8:12solves resolved, um I'll move on to the
- 8:16hands-on exercise.
- 8:20Before we dive into not by arrays, let's
- 8:23review Python's built-in data type
- 8:25called a list. A list is simply a
- 8:28collection of values enclosed in square
- 8:30brackets and separated by commas. For
- 8:33example, you can create a list like
- 8:36this. List A equals 0 1 2 3 4.
- 8:41While lists are versatile and easy to
- 8:43use, they aren't always the most
- 8:45efficient choice for numerical
- 8:47computations, especially when dealing
- 8:49with large data sets. This is where
- 8:52numpy arrays come into play. In numpy,
- 8:56arrays are handled using a data
- 8:58structure called numpy. Nd array. This
- 9:01ND array is similar to a list but is
- 9:04optimized for numerical computations and
- 9:07offers a lot more functions. To create
- 9:10an ND array, use the np.array function
- 9:13and pass in a list. For example, we can
- 9:17pass the list of from the previous slide
- 9:19to define it as an ND array. With this
- 9:22ND array, various nonby functions can be
- 9:25applied to perform efficient
- 9:27computations on the data. Let's take a
- 9:30closer look at how lists compare to
- 9:32nonpi.nd array. While a list in Python
- 9:35is versatile and can store different
- 9:37types of values, this flexibility can
- 9:40cause a drawback when performing
- 9:42numerical computations as operations
- 9:45between lists are often cumbersome and
- 9:48inefficient. In contrast, a numpy.nd ND
- 9:52array is designed to store elements of
- 9:54the same type which not only assists
- 9:57smooth arithmetic operations but also
- 10:00significantly enhance performance.
- 10:02Moreover, handling multi-dimensional
- 10:05data with lists requires nested
- 10:07structures that can quickly become
- 10:09complex and hard to manage. Numpy
- 10:13simplifies this process by efficiently
- 10:15managing multi-dimensional arrays
- 10:17through the concept of axes, allowing us
- 10:20to perform complex data manipulations
- 10:23with ease.
- 10:27Let's look at an example. Please open
- 10:30the notebook.
- 10:33Um, so I think
- 10:36that's not the best way, but I think I
- 10:39found a way to open the notebook. Um,
- 10:44which is if you um right click this
- 10:48notebook and open um click click open
- 10:52with and open in the new tab. I think
- 10:55you're able to
- 10:59open the notebook.
- 11:02Yeah. Yeah. I think that works. I'm not
- 11:05sure why you cannot
- 11:08um that ways. But um well now please um
- 11:13use that way to um open the notic.
- 11:17So, um I'll just
- 11:21um do a little bit of hands on for now.
- 11:27And
- 11:32so, um at the start, um you need to
- 11:36import the library we're using which is
- 11:39numpy. So um please click this um you
- 11:44need to run the cell
- 11:47do numpy
- 11:53and I think it should be working.
- 12:01Yeah.
- 12:04If you see this um green check mark,
- 12:07that means
- 12:10you're all good to go. And
- 12:16yeah, I
- 12:18And
- 12:19did we go?
- 12:26We are running a little bit out of time.
- 12:29So I'll
- 12:31go back to the lecture video for now.
- 12:34I'll coming back and
- 12:40like
- 12:41few minutes.
- 12:46First to handle numpy we need to import
- 12:50the library.
- 12:52There are two methods to import a
- 12:54library.
- 12:56You can import the entire library with
- 12:58an alias such as import library name as
- 13:01alias or you can import specific
- 13:04features from the library with from
- 13:06library name import feature name.
- 13:10Starting with the first method, it is
- 13:13common to import numpy with the alias.
- 13:16So we will import it this way.
- 13:21You can also import specific modules or
- 13:23functions from a certain library. To do
- 13:26so you can write from library name
- 13:29import module name.
- 13:34Here we will import random module and
- 13:36linel module. Here linel stands for
- 13:40linear algebra.
- 13:45Now let's take a look at the code.
- 13:50First let's look at how to create a
- 13:52numpy and d array. Any list-like
- 13:55structure such as Python lists and
- 13:57tpples can be converted into nonpendd
- 14:00array using ep.array function. Take a
- 14:03look at the example. We first define a
- 14:06Python list and then convert it to array
- 14:09using np.ray.
- 14:14To convert the nd array back to list, we
- 14:17can use the to list method. If a is a nd
- 14:20array, we write a do.to to list to
- 14:23convert it back to Python list.
- 14:28Before we dive further, let me explain
- 14:30some key aspects of numpy.
- 14:36Why is numpy so good at dealing with
- 14:37arrays? This is because numpai's
- 14:40functions are universal.
- 14:43Universal functions or efk in numpy are
- 14:47essential for performing elementwise
- 14:49operations on arrays. A EUK operates on
- 14:52each element of a ND array, allowing you
- 14:56to execute calculations efficiently
- 14:58without writing explicit loops. For
- 15:01example, if you have two arrays A and B,
- 15:05adding them together using A + B will
- 15:08produce a new array where each element
- 15:10is the sum of the corresponding elements
- 15:12in A and B. So if a is 24 and b is 21,
- 15:19the result will be 45.
- 15:22This operation is not only concise but
- 15:26also leverages namp's optimized
- 15:28performance because the plus operator
- 15:30automatically invokes the euunk and
- 15:33p.ab.
- 15:34The same principle applies to other
- 15:36arithmetic operations such as
- 15:39subtraction and multiplication.
- 15:42For other arithmetic operations, it
- 15:45would look something like this where
- 15:47different eupk are called for each
- 15:49operation.
- 15:51Implementing the code would look like
- 15:52this.
- 15:54Numpai offers a wide range of functions
- 15:56from exponential functions to
- 15:58trigonometric functions and much more.
- 16:01By leveraging these functions, we can
- 16:04handle a wide range of mathematical
- 16:06operations seamlessly, making the data
- 16:09processing tasks more streamlined and
- 16:12effective.
- 16:13Compare this with elementwise addition
- 16:16using Python lists. When you use the
- 16:19plus operator between two lists, it
- 16:21doesn't add the corresponding elements
- 16:23together. Instead, it concatenates the
- 16:26lists, combining them into a single
- 16:29list. That means that for addition you
- 16:32need to use a for loop instead making
- 16:34the code cumbersome. Moreover, trying to
- 16:37use other arithmetic operators like
- 16:39minus times or divide with lists will
- 16:43lead to errors. In contrast, with numpy
- 16:46arrays, these operations are
- 16:49straightforward and efficient thanks to
- 16:51universal functions.
- 16:56Let's go back to the notebook.
- 17:01To get ourselves more familiarized with
- 17:04NumPy, we will actually look at a
- 17:06publicly available data set and do
- 17:09simple analysis using NumPy.
- 17:13We will use a sample data set from Noah
- 17:15which contains data of daily temperature
- 17:18and precipitation recorded at one of the
- 17:21weather observing stations. The data is
- 17:24structured in 2D table format. Each row
- 17:27represents one entry of data and each
- 17:30column represents an attribute of that
- 17:32data such as station name, elevation,
- 17:35latitude and longitude, recorded date
- 17:39and recorded data. In this data set,
- 17:42three data are recorded. Maximum
- 17:45temperature, minimum temperature, and
- 17:47precipitation per day.
- 17:51We will load the data set using a nonby
- 17:53function enthy.load load txt function.
- 17:56Running the cell will download data and
- 17:59store it as variable raw.
- 18:04In the first half of this lecture, we
- 18:06will only use the data from the first
- 18:08week. Run the cell to select only the
- 18:11data from the first week. We will store
- 18:14the data as week 1 t-max, week 1 t-min
- 18:17and week 1 PRCP.
- 18:23Like Python lists, the length of an
- 18:25onedimensional array can be obtained
- 18:27using len function. A more common way in
- 18:30numpy is to access the shape attribute
- 18:34which returns the length per dimension
- 18:36of the array in array format.
- 18:42Now let's look at some numpy functions
- 18:44using the data set. For example, the
- 18:47temperature and precipitation are
- 18:49measured in t of degrees and t of
- 18:53millime respectively. However, it is
- 18:56more common to convert them to celsius
- 18:59and millime respectively. We can do so
- 19:02by dividing the array by 10. For
- 19:04example, running week 1 tmax /10 will
- 19:08divide each value in the array. Run the
- 19:11code and you will see how the operation
- 19:13is universal.
- 19:18We can then calculate the temperature
- 19:19range. Week one t-max minus week 1 t
- 19:23min.
- 19:27Other functions such as max, sum, mean,
- 19:30and standard deviation are also
- 19:33implemented in numpy.
- 19:38One point to note is that numpy treats
- 19:40division by zero slightly differently.
- 19:43In standard Python, trying to divide a
- 19:46number by zero led to zero division
- 19:48error.
- 19:52However, in numpy, this will result with
- 19:55an nd array of insf or epi. INF to be
- 20:00precise represents infinity. This means
- 20:03that although the code will not
- 20:05terminate, you may experience issues
- 20:07when trying to do further analysis using
- 20:09the returned array.
- 20:14Other math functions are also computed
- 20:17universally. For example, one can
- 20:20convert skewed data using a log
- 20:22function. Converting back is also
- 20:24universal.
- 20:29Now let's look at how NPI handles
- 20:31conditional operations with NDR. This is
- 20:39>> um so I want to look at
- 20:44the practice question in the notebook
- 20:49is
- 20:51practice question 2.1. So please open
- 20:56this section.
- 20:58Um in this question oh um what it's
- 21:01asking me is to create two
- 21:05one-dimensional np dot in the array
- 21:08arrays with a length of three with one
- 21:14with one with even numbers as elements
- 21:16and one with odd numbers elements then
- 21:20use the type function to confirm that
- 21:23they are type of mp in the array. So um
- 21:28here I don't know why
- 21:31done but um yeah so we want two arrays
- 21:37with even numbers and odd numbers so
- 21:41we're just going to call it a even
- 21:47this suggestion generative AI is um she
- 21:51sent me as but um yeah Um
- 21:56going to make it red.
- 22:00Yeah. So we're just going to call it
- 22:02even and odd and two four and six.
- 22:08One, three, and five.
- 22:11Then we can just try print it
- 22:17even and print
- 22:21print
- 22:22odd.
- 22:25So there it um it works fine and what
- 22:32it's also asking me is use the type
- 22:35function to confirm that they are type
- 22:37of um npd array. So want to print
- 22:45type
- 22:47even. So basically yeah
- 22:51let's see how it goes.
- 22:54Yeah, it's
- 22:56a numpy in the array type. So yeah, it
- 23:01works.
- 23:03It's just um
- 23:06just the basic and yeah.
- 23:11And the second question
- 23:16is add the two arrays that you created
- 23:19in number one and confirm that it
- 23:22results in element wise addition of the
- 23:26two arrays.
- 23:29So just we just need to add these.
- 23:34Yeah. Print
- 23:36try even plus odd and see how it goes.
- 23:44So yeah, it seems like it's working as a
- 23:50element wise edition.
- 23:52If we um try a Python list
- 23:58even
- 24:032 4 six
- 24:08I think they are talking about an N too
- 24:13right?
- 24:17Uh yeah.
- 24:20Um if you're curious, just try every
- 24:23time even and odd.
- 24:28I'm going to see how it goes in Python
- 24:32list
- 24:33that
- 24:36even plus odd.
- 24:41Yeah. So it works differently. Um you
- 24:44can see easily see the difference that
- 24:47for Python list which is above
- 24:52that one
- 24:54it's just adding the entire list.
- 24:57However, for numpy arrays, um it is
- 25:02adding the elements by elements 1 + 2 3
- 25:09and 4 + 3 7 and 6 + 5
- 25:1411. So yeah.
- 25:19And for number three,
- 25:22it is about um asking me to normalize
- 25:28um
- 25:32and
- 25:33output a one mention array that inputs
- 25:37normalized. And this question is a
- 25:40little bit um more complicated and I
- 25:43don't think I have a time to talk about
- 25:46in deeps but um
- 25:51what I can tell is um yeah you can give
- 25:54it a go if you curious always and
- 25:59it what it's asking me us is
- 26:05it's about a vector and
- 26:09you want to um
- 26:13you want to change the length of the
- 26:16vector, but you don't want to change its
- 26:20direction. So yeah,
- 26:23that's basically what it's um doing.
- 26:29Yeah. Oopsie.
- 26:35And then
- 26:37next section is about indexing. So I'm
- 26:40going to go back to texture video again.
- 26:46Also universal. For example, computing A
- 26:49equals to A will compare each element of
- 26:52A with itself resulting in true true.
- 26:55Similarly, A equals to B compares each
- 26:58corresponding pair of elements from A
- 27:01and B giving true false. A greater than
- 27:04B will return false true. These
- 27:07conditional operations are not only
- 27:09intuitive but also highly efficient,
- 27:12especially for quick filtering and
- 27:14analyzing data based on specific
- 27:17criteria.
- 27:19Again, compare this with conditional
- 27:21operation using Python lists. When using
- 27:24Python lists, the operation will
- 27:27evaluate the entire list and return a
- 27:30single boolean value.
- 27:33Next, I will explain broadcasting.
- 27:35Broadcasting is a powerful feature in
- 27:38NumPy to perform operations on arrays of
- 27:41different shapes by automatically
- 27:43adjusting their dimensions to be
- 27:44compatible. This means that when you add
- 27:47a scaler to an array, the scaler is
- 27:50extended to match the array's shape,
- 27:52enabling element-wise operations without
- 27:55the need for explicit loops. For
- 27:57instance, if you have an array A equals
- 28:0024 and you add a scalar 3, the result is
- 28:0357. This is because scalar 3 is extended
- 28:07into array 33. The same concept applies
- 28:11to multiplication as well as other
- 28:13operations.
- 28:15Aggregate functions are essential tools
- 28:18in NumPy to summarize the data within an
- 28:20array and to a single meaningful value.
- 28:24For example, if you have an array 1 2 3
- 28:27using npm max will return the maximum
- 28:29value of three. If you're looking to
- 28:32find the sum of all elements and pome
- 28:34will give you six. These aggregate
- 28:36functions enable you to quickly gain
- 28:39insights into your data.
- 28:44Now let's look at broadcasting in
- 28:46practice. If you try multiplying a
- 28:48scaler with a Python list, it will
- 28:51return an error. However, in numpy, the
- 28:54multiplication will be applied to each
- 28:57element. For example, if you do 2 * a,
- 29:01you will get the following result.
- 29:06This is equivalent to multiplying an
- 29:08array of twos.
- 29:13Note that broadcasting doesn't apply for
- 29:15all arrays. If you try multiplication
- 29:18between arrays with size five and size
- 29:21two, it will return an error. The length
- 29:24of either array must be one for
- 29:26broadcasting.
- 29:30Next, let's move on to another aspect of
- 29:33using numpy arrays, indexing, which is
- 29:36used to extract specific values from a
- 29:38numpy array. Suppose there is an array
- 29:41containing wind speed at some location
- 29:44measured every 30 minutes. We can ask
- 29:47some questions such as how can we
- 29:49extract data at a certain time? How can
- 29:52we extract data within a range? And how
- 29:54can we extract data every 60 minutes?
- 29:58All these questions can be answered
- 29:59using indexing.
- 30:02Just like in Python lists, each element
- 30:04in an ND array is assigned an index
- 30:07starting from zero on the left. To
- 30:09retrieve a particular value, simply
- 30:12specify its index within square
- 30:15brackets. For example, consider this
- 30:18array.
- 30:19Here the element three is at index one.
- 30:22By accessing A1, you obtain the value
- 30:25three. Like Python lists, NumPy supports
- 30:29negative indexing. Starting from the
- 30:32right end of the array, we can count
- 30:34minus1, minus2, and so on. This feature
- 30:39is particularly useful when you need to
- 30:41access elements relative to the end of
- 30:43the array which length may vary.
- 30:46You can also pass a list to specify
- 30:49multiple elements to retrieve from the
- 30:51array.
- 30:53Slicing is a type of indexing using
- 30:55colons. By specifying a slice in the
- 30:58format start colon end colon step, you
- 31:01can retrieve elements at regular
- 31:03intervals from the array. Here the start
- 31:06indicates the beginning index. The end
- 31:08marks the position where the slice stops
- 31:11and the step shows the interval between
- 31:13elements to extract. Note that the end
- 31:16index is exclusive. So that index is not
- 31:19included. In the example here, if you
- 31:22want to extract elements starting from
- 31:24index one up to index 7, taking every
- 31:27other element, you would use the slice 1
- 31:30col 7 2. This slice retrieves the
- 31:33elements three, 9, and 15.
- 31:37You can abbreviate the start, end, and
- 31:40step. For example, abbreviating the step
- 31:44will default to a step equal to one.
- 31:47If you abbreviate the end as well, it
- 31:50will continue slicing up to the last
- 31:52element.
- 31:54If you abbreviate the start, it will
- 31:56start slicing from the first element.
- 31:59If you only specify the step, it will
- 32:02slice the entire array at the specified
- 32:04step. Going back to the first quiz, if
- 32:08you want to extract the data at 1 p.m.,
- 32:10you can index as A2.
- 32:13If you want data between 1 p.m. and 4
- 32:16p.m., but excluding 4 p.m., you can
- 32:19index as A2 8. If you want to extract
- 32:23data every 60 minutes, you can index it
- 32:26as a colon 2.
- 32:31Now let's look at the notebook. Go to
- 32:34section 2.3. In this section, we will
- 32:37use only the data from January 2010.
- 32:44As we just saw, we use square brackets
- 32:47to retrieve elements from an array or
- 32:50slice an array. For example, we can
- 32:53access the t-max from January 1st with
- 32:56Jan Tmax 0. Keep in mind that the
- 32:59indexing in Python starts from zero.
- 33:05Using a negative number will retrieve
- 33:07elements counting from the end. For
- 33:10example, t-max from January 31st would
- 33:14be Jan Tax minus one.
- 33:19You can also retrieve multiple elements
- 33:22regardless of its order as well.
- 33:27Next, let's look at slicing. As we saw
- 33:30in the slides, slicing uses square
- 33:33brackets with three specifications. The
- 33:36start, end, and step index. For example,
- 33:40we can access the T-max from the first
- 33:43seven days using Jan Tmax 07
- 33:47and calculate the average maximum
- 33:49temperature.
- 33:53Also, we can look at precipitation on
- 33:56every Friday using genp0
- 33:59col 7.
- 34:04Let's move on to 2D arrays. We can
- 34:07actually reuse most of the concepts that
- 34:09we covered in 1D arrays.
- 34:12Just as we created a one-dimensional
- 34:14array earlier, a two-dimensional array
- 34:17can be formed by nesting lists within
- 34:19another list. This structure organizes
- 34:22data into rows and columns or in
- 34:25mathematical terms, matrices. Each row
- 34:28can represent a data entry and each
- 34:31column can represent an attribute. Note
- 34:34that in Python you can write with line
- 34:37breaks which helps improve readability
- 34:40especially with larger data sets.
- 34:43As you may have imagined universal
- 34:45functions in numpy are extended to
- 34:48n-dimensional arrays. This means that
- 34:50you can perform elementwise arithmetic
- 34:53operations on multi-dimensional data
- 34:55just as effortlessly as you do with
- 34:58one-dimensional arrays. For example,
- 35:00consider two 2D arrays A and B. When you
- 35:04add them together using A + B, NPI
- 35:07automatically adds each corresponding
- 35:09pair of elements, resulting in a new
- 35:12array where each element is the sum of
- 35:14the elements from A and B. This is the
- 35:18same for subtraction as well as other
- 35:20operations.
- 35:22Broadcasting is also applied just like
- 35:251D arrays. This means that when
- 35:27operating a scaler with a 2D array, the
- 35:31scaler is extended to match the array's
- 35:33shape. When you operate with a 2D array
- 35:36with one dimension of size one, it will
- 35:39extend in that direction to match the
- 35:41dimensions.
- 35:42If array is size 3x 1 and array B is
- 35:46size 2, then both arrays will be
- 35:48extended to match the dimensions of each
- 35:50other. This results in both arrays
- 35:53extended to shape 3x two. Understanding
- 35:56the concept of axis is essential when
- 35:59working with multi-dimensional arrays in
- 36:01numpy. An axis defines the direction
- 36:04along which operations are performed
- 36:06within an nd array.
- 36:09For example, consider a one-dimensional
- 36:12array like the one on the left. Since
- 36:14there's only one dimension, the axis is
- 36:17zero and any operation you perform will
- 36:20apply to the entire array. For a
- 36:23two-dimensional array like the one on
- 36:25the right, specifying axis zero means
- 36:28you are focusing on the rows.
- 36:31On the other hand, specifying axis one
- 36:34means the columns.
- 36:36This might be easier to understand when
- 36:38the arrays are written flat, especially
- 36:41for arrays with more than three
- 36:43dimensions.
- 36:47Aggregation is also almost the same for
- 36:50multi-dimensional arrays. The biggest
- 36:52difference is that arrays have more than
- 36:55one dimension. So you need to specify
- 36:57which dimension you want to apply the
- 36:59function. Imagine having a table where
- 37:01the rows represent subjects and the
- 37:04columns represent different students.
- 37:07You might wonder what are A and D's
- 37:09highest scores or what are the highest
- 37:11scores for each subject. Using
- 37:14aggregation functions, you can easily
- 37:16answer these questions if the dimension
- 37:19is specified correctly.
- 37:21If you just apply epmax, you will get
- 37:24the maximum score of all students in all
- 37:26subjects.
- 37:28With multi-dimensional arrays, you can
- 37:31specify the axis to operate the function
- 37:34using the axis argument by applying epax
- 37:37axis zero. Numpy performs aggregation
- 37:41column-wise and returns the highest
- 37:43score for each student.
- 37:46If you change to epimeax axis one, numpy
- 37:49examines each row individually and
- 37:51returns the highest score in each
- 37:54subject.
- 37:55This applies to other functions as well,
- 37:57such as calculating the minimum, sum,
- 38:00mean, standard deviation, and more.
- 38:07Now, let's return to the notebook. Go to
- 38:10section 2.4.
- 38:17Um, so I'm going to pause post a video
- 38:22moment and I want us to do a practice
- 38:27question again
- 38:29for question 2.2.
- 38:32Yeah. Before we further into the 2D
- 38:36arrays
- 38:38and the notebooks. Um and I think when
- 38:43I'm looking at these questions I think a
- 38:48lot of you
- 38:50finding little bit difficult and like
- 38:53fast and there I know there are a lot of
- 38:57information and I think you don't need
- 39:01to understand everything
- 39:04for now. Um I think it is very hard to
- 39:09just understand everything um just um
- 39:13listening one time. So um I recommend
- 39:17you to maybe try to understand
- 39:22the
- 39:25question number one level which is not
- 39:30that hard I find.
- 39:34And then you can um
- 39:37your buys um by watching recordings or
- 39:42um the slides.
- 39:46So yeah and in this question
- 39:50it's asking me to using the array
- 39:54January t-max that we built on the above
- 40:01create a new array
- 40:03T-max Monday month that contains the
- 40:07maximum temperatures of all Mondays in
- 40:11January.
- 40:12As mentioned previously, January 1st is
- 40:15a Friday, so um we don't really have a
- 40:20lot of time, so I'll keep it um
- 40:25fast. But
- 40:28assuming to create t-max,
- 40:31then what you need to
- 40:36think about in here is um this um AI is
- 40:40going to do it for me anyways. But um
- 40:43no, it's not what I want to do.
- 40:50You need to note that um in numpy I
- 40:55think it's in Python 2 um the index
- 40:58doesn't start from one. Yeah, that's
- 41:02step one point. So in in this case we
- 41:06want all the Mondays and then the f
- 41:10January 1st is start um starting from
- 41:14Friday. So it's Friday,
- 41:17Saturday, Sunday and Monday. So there is
- 41:20four
- 41:22days.
- 41:24So that um but you don't want to put
- 41:27four in here since it starts from zero.
- 41:30So 0 1 2 3. So you need to put three at
- 41:34first.
- 41:36then
- 41:38these um
- 41:44think you've covered in here is slicing.
- 41:47So for the rows um the next slicing row
- 41:51is for the end. So we don't need it for
- 41:55now.
- 41:56Then for the steps, the um last
- 42:01element, the steps, you want
- 42:06all the max temperatures of all Mondays.
- 42:09Therefore, you want it to make it to
- 42:14every week, right? So, left seven days.
- 42:19That's it. Oh, I think you need to run
- 42:23these all all these cells
- 42:27now, but it should be working.
- 42:35The next question, it's a little bit
- 42:37seems like a little bit difficult, but
- 42:41um it's not [clears throat] that hard.
- 42:44It's just asking me to calculate the
- 42:46average max temperature of all Monday in
- 42:49January and standard deviation of all
- 42:54max temperature of all Mondays in
- 42:56January. So, and that these hints um are
- 43:01talking about the standard deviation how
- 43:04do you the formula and how do you
- 43:07how what what is going on and something
- 43:09like that. But like when you're using
- 43:12Numpy, you can [snorts] skip these um
- 43:17steps by just using
- 43:20a
- 43:25I'm not sure where it covers but anyways
- 43:32you just you can just use the
- 43:37let me so for standard deviation.
- 43:40Standard deviation
- 43:43temperature
- 43:45max uh Monday
- 43:50Monday
- 43:53you can just use standard deviation
- 43:58then t-max Monday. So that that just
- 44:03gives an answer for this question.
- 44:08Um and that um works for
- 44:15average mean which is mean too. Um yeah
- 44:20so I'll move on to the lecture video.
- 44:25Here we will look at 2D arrays but the
- 44:28same thing applies to higher dimensional
- 44:30arrays as well.
- 44:32As explained, creating a 2D array in
- 44:35numpy is almost identical to creating a
- 44:381D array by using epi.array function.
- 44:42However, there is also a reshape
- 44:44function which allows you to resize your
- 44:46array in the way you like. For example,
- 44:49instead of defining a 3x 2 matrix from
- 44:52the beginning, you can first create a 1D
- 44:54array and then reshape into 3x two. Note
- 44:58that the total number of elements have
- 45:00to match before and after the reshaping.
- 45:06Another useful aspect of reshaping is
- 45:09that you do not need to specify all
- 45:10axes. For example, in the same example,
- 45:14if you specify the first axis to have
- 45:16three rows, you already know the other
- 45:19axis will have two columns because the
- 45:21total number of elements is six. So you
- 45:24can just write minus one and numpy will
- 45:27do the rest for you. Of course, if you
- 45:29try to specify the first axis to have
- 45:32four rows, it will return an error
- 45:34because six is not divisible by four.
- 45:39Let's move on to section 2.4.2.
- 45:42Here we will treat the entire Noah data
- 45:44set as a 2D array.
- 45:47Math operations are the same for 2D
- 45:50arrays. It will be applied elementwise.
- 45:53If we divide the data set by 10, this
- 45:56operation will be done universally.
- 46:02Move on to section 2.4.3.
- 46:08For multi-dimensional arrays, we call
- 46:11each dimension an axis. If the array is
- 46:142D, the first axis is the row and the
- 46:17second axis is the column.
- 46:20For different kinds of operations, you
- 46:23can specify the axis number to calculate
- 46:25that operation along a certain axis. For
- 46:28example, just ep.mminx will return the
- 46:32smallest number in the data set.
- 46:34However, writing np.mminxis0,
- 46:37we can access the lowest value for
- 46:40t-max, t min and prcp respectively.
- 46:47In such cases, the dimension of the
- 46:49array will be reduced. If you want to
- 46:52keep the dimensions as a 2D array, you
- 46:54set the keep dims argument to true.
- 46:58Now, let's return back to the slides.
- 47:02Continuing
- 47:04with this same example from before of a
- 47:072D array representing students scores in
- 47:10different subjects, let's explore how to
- 47:12extract specific elements using
- 47:14indexing. You might wonder what is
- 47:17student D's math grades or what are AB
- 47:20and C's grades in English and Spanish.
- 47:24For better understanding of indexing in
- 47:272D arrays, recall the flattened version
- 47:29of writing 2D arrays.
- 47:32To extract a specific element, you need
- 47:35to specify the index for all the axis.
- 47:38For 2D arrays, you need to pass two
- 47:41indexes separated by a comma. The first
- 47:44index will be used for the first axis or
- 47:47the row and the second index will be
- 47:49used for the second axis, the column.
- 47:52You can also use negative indexing as
- 47:55well. You can extract multiple elements
- 47:58along an axis by using slicing. For
- 48:01example, you can write a zero which is
- 48:04short for a zon to extract only the
- 48:07first row of the array. Recall how the
- 48:10colon was used for slicing 1D arrays. So
- 48:14a colon without any specification of
- 48:16start or end index means to take all the
- 48:19elements along that axis.
- 48:22Similarly, if you want to extract column
- 48:24one from the array, you would write a
- 48:26colon one. For axis one and more, you
- 48:29cannot abbreviate the colon.
- 48:32Finally, you can slice along multiple
- 48:35axis to extract consecutive elements. In
- 48:38this example, we are extracting elements
- 48:41up to index one along axis zero and all
- 48:44elements after index one along axis one
- 48:48resulting in 2x3 array.
- 48:51While basic indexing allows us to access
- 48:53elements using simple row and column
- 48:56numbers, there are times when we need to
- 48:58extract elements that are scattered
- 49:00across different positions within an
- 49:02array.
- 49:04Suppose we want to extract the elements
- 49:062 and 12 from this array. These elements
- 49:09are located at distinct positions. Two
- 49:12is in the first row, second column while
- 49:1512 is in the third row, fourth column.
- 49:18To achieve this, we can use advanced
- 49:20indexing.
- 49:22Depending on how we specify the indices,
- 49:24we could either extract them in a single
- 49:27row or column.
- 49:29When working with multi-dimensional
- 49:31data, there are specific rules that help
- 49:34us extract elements located at various
- 49:36positions efficiently.
- 49:39The first rule is axis- wise indexing,
- 49:42meaning we specify indices for each axis
- 49:45separately. In a two-dimensional array,
- 49:48this means providing one set of indices
- 49:50for the rows axis zero and another set
- 49:54for the columns axis one. The second
- 49:57rule shape and position preservation
- 50:00ensures that the structure of the output
- 50:03array matches the arrangement of the
- 50:05specified indices. This means that the
- 50:08extracted elements retain their relative
- 50:10positions based on how the indices were
- 50:12provided, maintaining the integrity of
- 50:15the data structure.
- 50:17Lastly, the rule of omitting notation
- 50:20encourages us to simplify our indexing
- 50:22expressions by leveraging NPI's
- 50:25broadcasting feature whenever possible.
- 50:28This allows for more concise and
- 50:30readable code. Let's look at each rule
- 50:33in more detail.
- 50:35The first rule axis wise indexing means
- 50:38that we pass the indexes as lists. One
- 50:41for axis 0 and another for index one and
- 50:45so forth. In the example 2 and 12 are
- 50:49located at index 01 and index 2 3
- 50:53respectively. We organize them so that
- 50:55the indexes are stored in separate lists
- 50:58for each access. As a result, a0213
- 51:03returns an array 212.
- 51:06The second rule ensures that the shape
- 51:08and positions of the elements in the
- 51:10output array align with the indices you
- 51:13provide. For example, consider changing
- 51:16the indexing of the previous example so
- 51:18that its shape is one by two array.
- 51:22Now the output shape is also size one
- 51:24two retaining the shape of the indexing.
- 51:28The third rule is omitting notation.
- 51:31Broadcasting applies to advanced
- 51:33indexing as well. So you can abbreviate
- 51:35the indexing to improve readability. For
- 51:38example, suppose we want to extract two
- 51:41non-consecutive elements along axis
- 51:43zero. We could specify axis zero for
- 51:46both elements, but we can just replace
- 51:49it with a single zero.
- 51:52Similarly, if we want to extract four
- 51:54elements as described in the slide, we
- 51:57can abbreviate both the indexes of axis
- 51:590 and one by shaping the index of axis 0
- 52:03as a single column and the index of axis
- 52:06one as a single row. Broadcasting will
- 52:09handle this so that it automatically
- 52:11aligns their shapes and returns an array
- 52:14and 2x two shape.
- 52:16Finally, of course, you can combine
- 52:18basic indexing and advanced indexing to
- 52:21extract elements from arrays in whatever
- 52:24way you wish.
- 52:26So, if we go back to the quick quiz from
- 52:28earlier, the answer to the first
- 52:30question will be indexing the array at
- 52:3213. To extract A, B, and C's grades in
- 52:37English and Spanish, we would combine
- 52:40basic and advanced indexing as shown in
- 52:43the slide.
- 52:47Let's take a look at the notebook again.
- 52:49Go to section 2.5.1.
- 52:53Indexing for 2D arrays is the same as
- 52:56indexing for 2D lists by using two
- 52:59square brackets.
- 53:03However, what's different is that you
- 53:06can use just one bracket with commas to
- 53:08extract a single element. Also, if you
- 53:11want to extract several rows along a
- 53:13certain axis, you need to use two square
- 53:16brackets and a nested structure.
- 53:22Next, let's move to section 2.5.2 to
- 53:25look at some examples of advanced
- 53:27indexing.
- 53:31If we want to extract elements at 0 1
- 53:33and 2 1, we can specify 0 and two along
- 53:37axis 0 and one along axis one. Depending
- 53:41how you write 0 and two, the output will
- 53:44be either 2x one array or one by two
- 53:46array.
- 53:51Next, if we want to extract elements at
- 53:532 1 and 2 3, we specify two along axis 0
- 53:57and 1 and 3 along axis 1.
- 54:04Now, if we want to extract elements at 0
- 54:061, 03, 2, 1, and 2 3, we write 02 along
- 54:12axis 0 and 1 3 along one. Here we will
- 54:16need to be careful to write 02 as 2x 1
- 54:20array and 1 3 as 1 by 2 array. If we for
- 54:24example write both as a simple list it
- 54:27will extract elements at 01 and 2 3.
- 54:33Since this way of writing is complex
- 54:36especially for higher dimensional arrays
- 54:39there is a function called numpy x.
- 54:45Next, let's move on to slicing. Again,
- 54:48slicing in 2D is also similar to slicing
- 54:51in 1D. For example, X 5 will return the
- 54:56first five rows of X, the Noah data set.
- 55:02We can retrieve T-max and T-min on a
- 55:053day span using X col 32.
- 55:12In multi-dimensional arrays, we can use
- 55:15a single colon to skip a certain axis.
- 55:18For example, to retrieve t min and prcp
- 55:22for all rows, we write x col minus 2
- 55:26colon.
- 55:30Finally, let's talk about boolean
- 55:32indexing. Boolean indexing allows us to
- 55:35filter elements based on a condition.
- 55:38For example, imagine we have a
- 55:40two-dimensional array A that looks like
- 55:42this. If we want to select elements that
- 55:45are divisible by three, we can apply the
- 55:48condition a percent 3 equals to zero.
- 55:51This condition creates a boolean array
- 55:54where each element is true if it meets
- 55:56the condition and false otherwise.
- 55:59By using this boolean array to index a,
- 56:03numpy will return a one-dimensional
- 56:05array containing only the elements that
- 56:07satisfy the condition. This is helpful
- 56:10when you want to filter the elements
- 56:12that fulfill a certain condition from a
- 56:14large data set.
- 56:19Let's look at an example of this using
- 56:21the notebook. Go to section 2.5.3.
- 56:26In the Noah data set, there are certain
- 56:28rows where PRCP is equal to 9,999,
- 56:34which represents data that was not
- 56:36recorded correctly. We can find out by
- 56:38creating a boolean array X= to 999.9.
- 56:47We can also use Boolean indexing to
- 56:50extract only specific axes.
- 56:55You can also apply boolean indexing
- 56:57along only a certain axis or use it for
- 57:00other functions such as ex function.
- 57:06Let's take a moment to recap what we've
- 57:08learned about numpy this week. Numpy is
- 57:11an essential library for performing data
- 57:13manipulation efficiently in Python. We
- 57:16started by learning how to create an ND
- 57:19array using the np.array function.
- 57:23We then explored how NPI's functions are
- 57:25universal, letting us perform
- 57:28elementwise calculations without the
- 57:30need for loops. Broadcasting was another
- 57:33key topic, enabling us to operate on
- 57:35arrays of different sizes by
- 57:37automatically adjusting their shapes to
- 57:40match. Understanding indexing and the
- 57:43concept of axes in n-dimensional arrays
- 57:45was crucial as it helped us extract
- 57:48specific elements or subarrays with
- 57:50ease.
- 57:52Aggregation functions are perfect for
- 57:54calculating statistics along each axis
- 57:57of our data. Whether it's finding the
- 57:59maximum value, calculating the sum, or
- 58:03determining the average, these functions
- 58:05help us summarize and interpret our data
- 58:08effectively.
- 58:10That concludes today's lecture.
- 58:16>> All right, everyone.
- 58:20So again I think it might be a lot of
- 58:24information for
- 58:27I think most of you if you are not you
- 58:31have not any bit of science or
- 58:33mathematics background but [sighs] yeah
- 58:36again you just um don't worry about it
- 58:41just keep um revising and practicing
- 58:44that would really help and
- 58:49for Now I there were actually two more
- 58:54practice questions but
- 58:56only have about 15
- 59:00Oops that's not I want to show um have
- 59:0515 minutes left. So, and I've seen that
- 59:11quite some of you are
- 59:19thinking like curious that how you can
- 59:22use these numpy
- 59:27into actual
- 59:32um like a data science skills and what
- 59:35what I can do with the Python and
- 59:40all that type of question. And I really
- 59:44think
- 59:46that
- 59:48um for example, if we look at this
- 59:51question, it in this question is asking
- 59:54me to standardize
- 59:58um
- 59:59the data in here. and
- 1:00:05what what why why um the thing is um
- 1:00:11standardize um these
- 1:00:15data helps in terms of machine learning
- 1:00:20um
- 1:00:21for example in these data we have
- 1:00:24temperature maximum temperature minimum
- 1:00:28temperature and also pre participation
- 1:00:32um which is in like mill or degrees. And
- 1:00:36they are the these value are really like
- 1:00:40different and those the the value
- 1:00:44difference
- 1:00:46like high or low really um
- 1:00:53effect um in the in terms of machine
- 1:00:56learning. So that um standardize allow
- 1:01:00you to make all the
- 1:01:04scale or value to make it um really easy
- 1:01:09to compare and yeah.
- 1:01:14So yeah and actually in this case you
- 1:01:18standardize um
- 1:01:22this data um all the rows going to
- 1:01:25become for example in the mean would
- 1:01:27become zero and for the standard
- 1:01:31deviation would become one. So it really
- 1:01:35easy to
- 1:01:37compare that and
- 1:01:42we're talking about numpy and what numpy
- 1:01:45does is
- 1:01:47it helps to um calculate all all those
- 1:01:51values and all those data really fastly
- 1:01:55than um the python coding. So that is um
- 1:02:01why we are learning um numpy today and
- 1:02:05we're going to cover um about more
- 1:02:09depths about data science and machine
- 1:02:11learning skills um later sessions but
- 1:02:16yeah it is just the beginning of these
- 1:02:19those
- 1:02:20um road maps. So yeah, I hope you stay
- 1:02:28in tune and
- 1:02:31I really recommend you to revise these
- 1:02:35um and lectures.
About this transcript
This page contains the full transcript of GCI World 2026 September Session2 During Lecture by GCI World (Matsuo-Iwasawa Lab), generated from the public captions YouTube serves with the video. The transcript has 7,188 words across 1,187 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.