Processing of data — Transcript
Full transcript
- 0:00[Music]
- 0:14hello learners i am dr subhad keshwani
- 0:16working with india gandhi national open
- 0:18university in school of management
- 0:19studies the topic which i am going to
- 0:21talk today is processing of data i think
- 0:24in our preceding sessions we have talked
- 0:25a lot about data and this today's
- 0:28session you know revolves around
- 0:29research methodology and statistical
- 0:30analysis which is part of our course
- 0:32called uh
- 0:34mco3 and it's related to the program
- 0:36called amcom so prior to going into the
- 0:39depth of this topic i just want to throw
- 0:41a light what exactly we have
- 0:43recapitulate what exactly we have done
- 0:44in our preceding sessions because
- 0:46when we talk about the research
- 0:48methodology when we talk about you know
- 0:49the statistical analysis
- 0:51there is there is a great use of you
- 0:52know the today's session because the
- 0:55in our preceding sessions we have talked
- 0:56a lot the whole block was talking about
- 0:58you know the research and data
- 0:59collections and if we go more into the
- 1:01depth of this research and data
- 1:03collection we we realized that you know
- 1:05we have already covered you know the
- 1:07introduction to research research plan
- 1:10collection of data how the sampling is
- 1:12going to be done either it could be a
- 1:13random sampling or non random sampling
- 1:15or measurement of scaling techniques
- 1:17which could be you know the comparative
- 1:18in nature or non-comparative in nature
- 1:20so this this stuff this particular the
- 1:22first block which which emphasize on on
- 1:25on the collection of on modus operandi
- 1:27of collecting the data and we have we
- 1:29have seen that you know how the
- 1:31how the data collection is going to play
- 1:33a very important role and now the second
- 1:35block which revolves around you know the
- 1:37processing and preservation of data
- 1:39because how we are going to process the
- 1:41data because in the in the first
- 1:44first you know the block we have uh
- 1:46devoted somewhere around 10 to 12
- 1:48lectures which was talking on you know
- 1:50the data collection and the second block
- 1:52was purely emphasizing on processing and
- 1:55preserving of data so
- 1:57as far as you know this particular block
- 1:58is concerned we have got you know
- 2:00certain certain chapters certain units
- 2:02which are going to talk about you know
- 2:03certain parameters but the today session
- 2:06revolves around you know the processing
- 2:08of data so if you go more into the depth
- 2:10of processing of data you will find out
- 2:12that
- 2:14that this data process is very important
- 2:16because
- 2:17if you if you see this particular steps
- 2:19in quantitative research starts with the
- 2:21theory then you have the hypothesis and
- 2:23we have already have a very elaborative
- 2:25session which talks about you know what
- 2:27exactly the hypothesis is how this
- 2:29alternate hypothesis you know differs
- 2:31from null hypothesis and when these when
- 2:34we are going to accept the hypothesis
- 2:36when we are going to reject the
- 2:37hypothesis so
- 2:39this particular session was talking
- 2:41about that factors and now you know the
- 2:44third part was the research design that
- 2:46how we are going to design our research
- 2:48it's not just you know a lot of planning
- 2:50is involved because there's certain
- 2:52problems which need to be taken care so
- 2:54research design talks about you know
- 2:56something which could not be done at the
- 2:57mid or at the end but at the preamble
- 3:00stage or at the preliminary stage so
- 3:02operationalizing concepts and you know
- 3:04the finally select selecting a research
- 3:06site that is more important and then
- 3:08selecting respondents then data
- 3:10collection and now already we have
- 3:12applied talk about in the collection of
- 3:14data which talks about you know the data
- 3:16collection
- 3:18as far as the secondary data is
- 3:19concerned as far as the primary data is
- 3:20concerned so we have a very elaborative
- 3:22sessions then the today discussion
- 3:24revolves around the data processing and
- 3:27you see that after you process the data
- 3:29there are certain analysis which need to
- 3:30be done then finding and conclusion then
- 3:32how you are going to publish your
- 3:34results so this is all about you know
- 3:36the steps which you are going to follow
- 3:38and
- 3:39this data processing is very important
- 3:40because when you go more into the depth
- 3:42of processing of data you will find out
- 3:45that there are certain ingredients which
- 3:46are you know very important as far as
- 3:48you know the process of data is
- 3:50concerned so one is editing of data then
- 3:52you have coding of data editing means
- 3:55like whatever the data you have is in a
- 3:57raw format now you are going to
- 4:00convert into a finished course for that
- 4:02you know the editing is going to play a
- 4:03very important role or already you know
- 4:06the
- 4:07the finished data is there which need to
- 4:09be revamped so editing of data we will
- 4:12definitely throw a light on elaborately
- 4:14go into the depth of this editing of
- 4:15data then coding of data is concerned so
- 4:18how you are going to decode a code the
- 4:19data so that it can be you know
- 4:22encrypted or decrypted and used for the
- 4:24particular purpose then classification
- 4:26of data is there where we are going to
- 4:28talk about types of classification
- 4:30classification according to external
- 4:32characteristics classification according
- 4:34to internal characteristics and
- 4:35preparation of frequency distribution so
- 4:38this is you know
- 4:40a part which is which is you know
- 4:42dedicatedly talking or with respect to
- 4:44classification of data and then we have
- 4:46tabulation of data like types of tables
- 4:49parts of statistical tables and
- 4:50requisites of good statistical tables
- 4:53because when you are going to tabulate
- 4:54any things i think its going to become
- 4:57quite easier and
- 4:58it could be in a very competitive format
- 5:00or in a tabulated format and which could
- 5:02be quite
- 5:04conducive as far as you know the data
- 5:06analysis is concerned so anyway this
- 5:08processing of data we have already
- 5:10talked about and if you if you go more
- 5:12into the depth of that it's talk about
- 5:13you know the uh the many things now the
- 5:16types of data is going to be bifurcated
- 5:18into qualitative data and quantitative
- 5:20data
- 5:22as far as you know the scaling
- 5:23techniques or measurement is concerned
- 5:24we have already talked about you know
- 5:26the four important ingredients that is
- 5:28nominal ordinal ratio
- 5:30and we over there we have seen that how
- 5:33this
- 5:34qualitative and quantitative data is is
- 5:36going to be
- 5:37measured but here we are talking more
- 5:40about you know the qualitative and
- 5:41quantitative data because these
- 5:43data is going to be types of data is
- 5:45going to be bifurcated into qualitative
- 5:47and quantitative so
- 5:49when you are going to talk about the
- 5:50qualitative it is again bifurcated into
- 5:53nominal and ordinal whereas quantitative
- 5:56is going to be bifurcating to discrete
- 5:57and continuous so nominal two or more
- 6:00categories are not in any particular
- 6:02order or rank
- 6:04example given blood groups a b
- 6:06a positive or b positive or a b or o
- 6:09positive or area of residence nor south
- 6:12east west and center so this is
- 6:13considered to be the nominal whereas
- 6:16ordinals category are in order of ranks
- 6:18that is moderate and severe and on the
- 6:21other hand when you are going to talk
- 6:23about the quantitative data it is going
- 6:25to be bifurcated into discrete and
- 6:26continuous so counted in whole numbers
- 6:29that is example number of family members
- 6:30or continuous can have factors
- 6:33ah example given height of
- 6:35161.5 centimeters of blood glucose level
- 6:38so this is going to be considered and
- 6:41there are certain more examples which
- 6:42can give what exactly the quantitative
- 6:44variables are and what exactly the
- 6:46qualitative variables are so
- 6:49when you are going to talk about the
- 6:50quantitative variables one that can be
- 6:52measured and expressed numerically
- 6:54that need to be considered as a
- 6:56quantitative variable the measurement
- 6:58convey information regarding amount
- 7:00and there are certain examples like
- 7:02blood pressure heart rate the heights of
- 7:04wedding males the weights of preschool
- 7:06children and the ages of patients seen
- 7:08in a dental clinic these are considered
- 7:10to be the example of quantitative
- 7:12variables on the other hand when you are
- 7:14going to talk about the qualitative
- 7:15variables the characteristics that can
- 7:17be measured quantitatively but can be
- 7:20categorized the measurement convey
- 7:22information regarding the attribute the
- 7:24measurement in real sense can't be
- 7:26achieved but person place or thing
- 7:28belonging to different categories can be
- 7:30counted
- 7:32example given you know the sex of the
- 7:34patient color and order of stool and
- 7:36urine samples etcetera are are
- 7:38considered to be the qualitative
- 7:39variables now we have seen that you know
- 7:42when we have bifurcated the data i think
- 7:44the data processing is generally if we
- 7:46more talk go and talk about the data
- 7:48data processing generally the correction
- 7:50and manipulation of items of data to
- 7:52produce meaningful information the
- 7:54intention is very clear that whatever
- 7:55the data is there we are going to make
- 7:57it meaningful we are going to
- 7:59use it for the particular purpose or we
- 8:01are going to customize those data so for
- 8:03customization you know when you are
- 8:04going to customize the data for the
- 8:06particular purpose or for or for the
- 8:10taylormade use i think
- 8:12there you know the data processing is is
- 8:14going to be a very important aspect and
- 8:17in this sense it is can be considered as
- 8:18subset of information processing so
- 8:21uh
- 8:22the change of information any manner
- 8:24detectable by an observer and we have
- 8:26seen that there are certain tools right
- 8:28now there are certain technologies there
- 8:29are certain computers or you know the
- 8:31gadgets are there which are all the
- 8:34customized softwares are there which we
- 8:36are going to talk more in our preceding
- 8:38coming sessions or in the preceding
- 8:39session we have thrown a light
- 8:41that what is data processing so data
- 8:43processing is basically you know is a
- 8:46method of processing the data it could
- 8:48be in the format of table it could be in
- 8:50the form of editing or it could be in
- 8:52the form of coding so anyway we see that
- 8:54you know the data collection is the
- 8:56first step then data preparation is
- 8:57there then data entry is there then data
- 8:59processing is there so
- 9:01while talking about the data processing
- 9:02i think we have to follow the first
- 9:05three steps that is collecting the data
- 9:07data preparation data entry and then
- 9:09data interpretation and finally you know
- 9:11the data storage is there so if you are
- 9:13going to talk about the data storage i
- 9:14think data storage is is a very
- 9:16important aspect because because this
- 9:18data storage is
- 9:20is going to store whatever the data you
- 9:23have processed or you know the finished
- 9:24data which can be used in a real-time
- 9:26manner or in a different manner so
- 9:28anyway
- 9:29if you see this particular
- 9:31steps the first
- 9:33part talks about acquisition the second
- 9:36is going to talk about the publishing
- 9:37third is curation four is processing
- 9:40five is you know how the movement is
- 9:42going to be done and six is also talking
- 9:43more about that and seven is the use so
- 9:47anyway what we have observed that if you
- 9:49go more into the depth of research data
- 9:51cycle
- 9:52this research data cycle is
- 9:54is more talking about you know the data
- 9:57planning and design data collection data
- 10:00processing data study and analysis
- 10:02and
- 10:03data preservation and data reuse we have
- 10:06already talked about and all those you
- 10:08know things in require lot of literature
- 10:10review a lot of you know the
- 10:13the secondary mode of collecting the
- 10:14data the primary mode of collecting the
- 10:16data so anyway
- 10:17we have thrown a light on the glimpse of
- 10:20of certain data how we are going to move
- 10:22into the process of data and what could
- 10:24be the do's and don'ts which we have to
- 10:26follow it's not like that you can
- 10:28process the data at the very beginning
- 10:29stage the modus operandi is very clear
- 10:31you start with the
- 10:33collection of data then you can
- 10:36do certain methodologies and then you
- 10:37can process the data so the collecting
- 10:40data in research is processed whatever
- 10:41the data we have collected in research
- 10:43is processed and analyzed
- 10:46to come to some conclusion or to verify
- 10:49the hypothesis made so what we observe
- 10:52that when we when we develop the
- 10:53hypothesis it could be either alternate
- 10:55hypothesis or null hypothesis or it
- 10:58could be you know they accepted or
- 11:01rejected so in that case
- 11:03you know whatever the data we use to
- 11:05collect uh it whether it could be in a
- 11:08in a primary mode or a secondary mode we
- 11:10see that processing of data is important
- 11:12as it makes further analysis of data
- 11:15easier and efficient because when you
- 11:17process the data
- 11:18you do lot of further analysis and
- 11:20processing of data technically means if
- 11:23we go more into the backdrop of
- 11:25processing of data it is basically you
- 11:26know editing of the data
- 11:28coding of the data classification of
- 11:31data and then tabulation of data so we
- 11:34start with editing we see
- 11:36what are the you know the rectification
- 11:38need to be required then we code or
- 11:40decode it so that it can be and then
- 11:42classify of data and then finally
- 11:44tabulation of data so we are going to
- 11:45cover these four points in a more
- 11:47elaborative manner and if we start with
- 11:50purpose of editing the process of
- 11:51checking and adjusting responses in the
- 11:54com completed question is
- 11:56for omission
- 11:58legibility and consistency and reading
- 12:00them for coding and storage this is one
- 12:02of the important purpose of editing and
- 12:05if you go more into the depth of purpose
- 12:07of editing you will find out accuracy of
- 12:09data collected whatever the data we have
- 12:11collected
- 12:12need to be accurate
- 12:14for consistency between responses there
- 12:16must be a synchronization there must be
- 12:18a
- 12:20homogeneity between the responses which
- 12:22need to be come by the by the
- 12:25respondents and uniformity is there it's
- 12:27not like that whatever the question
- 12:29there must be some questions which can
- 12:31given to some of the respondents and
- 12:33some are not so while you know doing all
- 12:35those things uniformity or homogeneity
- 12:38need to be maintained so for
- 12:40completeness in response to reduce
- 12:42effects of item non-responses to
- 12:45facilitate and simplify coding and
- 12:47tabulation this is
- 12:49one of the way by which you know you can
- 12:52you can do the editing and then to
- 12:53better utilize question answered out of
- 12:55order so whatever the questions near we
- 12:57have floated to the respondents and when
- 12:59they reciprocate i think
- 13:02you are going to utilize that that
- 13:04answer in a more systematic manner this
- 13:06could be the purpose of
- 13:07editing so now when you are going to
- 13:09talk about the coding the process of
- 13:11identifying and classifying each answer
- 13:14with with a numerical score or other
- 13:16character symbol so that we can
- 13:19we can have some coding in between and
- 13:22this can somewhat make the job quite
- 13:25easier or it can
- 13:27make the things in a in a more different
- 13:30manner so the numerical score symbol is
- 13:32called a code and serves as a rule for
- 13:35interpreting classifying and recording
- 13:37data so what we have observed that when
- 13:40you are going to record the data i think
- 13:42the numerical score symbol is called a
- 13:44code and this can somewhat do make the
- 13:47things quite easier so identifying
- 13:49responses with course is necessary if
- 13:52data is to be processed by
- 13:54computer so now what we observe that
- 13:56tabulation is the process of summarizing
- 13:58raw data and displaying the same in
- 14:01compact form
- 14:02that is in the form of statistical table
- 14:04for further analysis and when mass data
- 14:07has been assembled it becomes necessary
- 14:09because what we observe that when we are
- 14:11going to talk about the data i think the
- 14:12data is now gigantic in nature that is
- 14:15that is the reason you know in in
- 14:17technology we used to talk about big
- 14:19data analytics because big data talks
- 14:21about you know the data which are in
- 14:23uh which are known as a mass data and it
- 14:26becomes necessary for the researchers to
- 14:28arrange the same in some kind of concise
- 14:30logical order which may be called
- 14:32tabulation because if you tabulate the
- 14:34data i think you can very easily make it
- 14:36in a table format or in a logical order
- 14:39which can be quite useful for
- 14:40understanding the
- 14:42understanding the data analysis
- 14:44there are rules for tabulation the table
- 14:47should suit the size of the paper and
- 14:49therefore the width of the column should
- 14:51be decided before hand
- 14:53number of columns and rows should
- 14:55neither be too large nor too small
- 14:58as far as possible figure should be
- 15:00approximated before tabulation this
- 15:03would reduce unnecessary details so what
- 15:05we observe that when you are going to
- 15:06talk about rule for tabulation i think
- 15:09the uh there are certain things which
- 15:11need to be taken care and item should be
- 15:14arranged either in alphabetical
- 15:15chronological or geographical order or
- 15:17according to size because when you are
- 15:20going to
- 15:22make a table i think there are certain
- 15:24fields which need to be pre-decided or
- 15:26there are certain columns which need to
- 15:28be made so if you make a table in that
- 15:30format
- 15:31i think somewhere you know the things
- 15:33are going to be solved so there are no
- 15:36hard and fast rules for the tabulation
- 15:37of data but for constructing good table
- 15:40following general rules should be
- 15:42observed while tabulating statistical
- 15:44data
- 15:45so what we observe that
- 15:47that when you are making a
- 15:49the classification of data leads to the
- 15:51problem or presentation of data and the
- 15:53presentation of data means exhibition of
- 15:55the data in such a clear clear and
- 15:58attractive manner that these are easily
- 16:00understood and analyzed
- 16:02so there are many forms of presentation
- 16:04of data of which the following three are
- 16:06well known that is textual presentation
- 16:08tabular presentation diagrammatic
- 16:10presentation
- 16:11video presentation of data so what we
- 16:14have observed that when you are going to
- 16:15talk about you know the diagrammatic
- 16:17presentation we have a full fledged
- 16:19session which is going to talk about
- 16:20diagrammatic presentation of data it
- 16:23could be you know the one
- 16:24[Music]
- 16:27two layer diagram or one layer diagram
- 16:29or you know there are certain parameters
- 16:31which need to be followed so we are
- 16:32going to have a full fledge session
- 16:35which talks about diagrammatic
- 16:36presentation
- 16:37now there are certain advantages of
- 16:39tabulation so it simplifies complex data
- 16:42it facilitates comparison
- 16:44it facilitates computation so what we
- 16:46observe that whenever you have got a
- 16:48complex data there are certain data
- 16:49which need to be quite complex in nature
- 16:52so how you are going to simplify that
- 16:54data how you are going to put the data
- 16:56in that format for that we have observed
- 16:58that this tabulation is very important
- 17:00it facilitates comparison it facilitates
- 17:03computation also how you are going to
- 17:05compute the data it present facts in
- 17:07minimum possible space
- 17:09so tabulated data are good for
- 17:11references and they make it easier to
- 17:13present the information in the form of
- 17:15graphs and diagrams so what we have
- 17:18observed that
- 17:20when you are going to tabulate the data
- 17:21these are the certain things which need
- 17:23to be taken care and classification of
- 17:26data
- 17:27is is again talking about the data
- 17:29classification
- 17:30which which emphasize more on sorting
- 17:33and categorizing data into various types
- 17:35forms or any other distinct class so
- 17:37data classification enables the
- 17:40separation and classification of data
- 17:41according to data set requirements for
- 17:44various business or personal objectives
- 17:47it is mainly a data management process
- 17:49so what we observe that you know when
- 17:50you are going to classify the data i
- 17:52think
- 17:53there are certain data management
- 17:54process which need to be followed
- 17:55because data management process
- 17:58is is going to what you have seen that
- 18:00that these data are unscattered in
- 18:01nature or unstructured so now if you if
- 18:05you see this particular image you will
- 18:06find out that it is going to be you know
- 18:08classify in terms of circles or you know
- 18:11the squares or recta or this triangles
- 18:14so classification of data can be
- 18:17bifurcated in either in geographical
- 18:20classification or chronological
- 18:22classification qualitative
- 18:24classification quantitative
- 18:26classification alphabetical
- 18:28classification so these are these are
- 18:30the classification of data and four
- 18:32steps data classification processes are
- 18:34there
- 18:35define the objectives of the data
- 18:36classification process create workflows
- 18:39based on the selected classification
- 18:40tools define the categories and
- 18:43classification criteria
- 18:44define outcomes and usage of classified
- 18:47data so
- 18:48what we observe that when we are going
- 18:50to classify or the data they are these
- 18:52are the four process which are which
- 18:54need to be considered so and they are
- 18:57certain comparison between the
- 18:58classification and tabulation so
- 19:00classification is basically talking
- 19:03about arranging the data into different
- 19:04groups based on their characteristics
- 19:07whereas tabulation talks about
- 19:09representing the data in more organized
- 19:11way for example in rows and columns it
- 19:13happens after it takes a places after
- 19:15the data has been collected and happens
- 19:17after the classification so tabulation
- 19:20is something which can be governed after
- 19:22we have classified and methods of
- 19:24arranging the data they arrange
- 19:26data based on their characteristics and
- 19:27behavior arranging rows and columns
- 19:29which we have already talked about and
- 19:31the best example is the spreadsheet in
- 19:34in excel spreadsheet we see that we make
- 19:36the we tabulate the data in in rows and
- 19:38column and the fields are quite
- 19:41big in numbers so
- 19:43and analyze the data much easier to help
- 19:45represent data this is what we have we
- 19:48have covered in this in this particular
- 19:49session and we have seen that you know
- 19:52how the how this
- 19:54how the things are going to be done with
- 19:56the help of that
- 19:58so anyway i think we have we have talked
- 20:00a lot about this this particular thing
- 20:02and we have seen that how this data
- 20:06processing of data is is quite important
- 20:08and when you are going to process the
- 20:09data
- 20:10i think
- 20:11up to some extent you you lead to a
- 20:14conclusion and which can help you in in
- 20:17interpreting the results in a more
- 20:19systematic manner so in our next session
- 20:22we are going to talk about something
- 20:24which is over neighbor to that that is
- 20:26diagrammatic presentation of data and
- 20:28this diagrammatic presentation of data
- 20:30is is more talking about the the way of
- 20:33of you know
- 20:35the data which you need to be you know
- 20:38put in that format thank you very much
- 20:43[Music]
- 20:56you
About this transcript
This page contains the full transcript of Processing of data by MCO-3 [RM&SA], generated from the public captions YouTube serves with the video. The transcript has 3,769 words across 602 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.