[00:02] from turning that data into intelligent decision. From recommendation systems to systems, machine learning is already With that being said, I welcome you all to this session on machine learning with [00:15] Python full course. Before I begin our session, just a quick info guys. If you want to build a strong career in AI and machine learning then simply learn offers a advanced professional certificate course in generative AI and [00:27] machine learning design in collaboration with EICA consortium of IIT Kpur and powered by industry partners like Microsoft Azure. This program includes live online classes, hands-on projects, expert mentorship and career support [00:41] helping you move from fundamentals to real world AI applications. Now coming will learn how to use Python for machine learning from scratch to advanced Python refresher and maths and [00:54] into data processing and exploratory data analysis. After that we are going to explore core machine learning concepts like supervised, unsupervised, regression, classification, clustering, model evaluation techniques. We will [01:08] also understand advanced areas like deep learning with TensorFlow and KAS, NLP, reinforcement learning and real world ML project building. And finally, we are hands-on industry projects like sales prediction, customer analysis, [01:21] modeling. By the end of this course, you will not just understand machine learning, but you'll also be able to build real intelligent systems using Python. So let's let's begin with this journey and try to understand what is [01:36] machine learning and why is there so much buzz around it. Let me first uh let me go through what all we would be covering in this particular course and then we begin the discussion. [01:50] So if we talk about a learning path you know the first uh you know uh topic that we will be covering today is introduction to machine learning which focuses on the basics of machine learning. Second is supervised learning [02:04] regression and application which focuses on supervised learning with an emphasis on understanding and implementing different types of regression models. Third is supervised learning classification applications like [02:18] what are the different learning techniques available in the machine learning. So first technique that we would be covering up is supervised learning. Under supervised learning we would be covering up regression as well [02:31] as classification. Then we would be moving on to the ensemble learning method which focuses on advanced ensemble methods to enhance the performance and robustness of the models. Then we would be moving on to [02:45] unsupervised learning and finally recommener system along with the application. What is machine learning and why is there so much of buzz? Why are you here to learn machine learning? Let's we let me put this in another way. [03:00] So basically we want the machine to get trained with our data. We want the machine to learn from the data so that it can predict data. What kind of data that we want to predict? Why do we want to have want these predictions? Right? [03:15] Because now are we living in the digital technology where data is all around us? Even this you know uh you know session is data right. When all the material [03:29] that has been sent to you is data, anything on the news which is coming is data. Anything doing for entertainment is data. E-commerce is data. You know your work profile is data because we are living in huge amount of [03:45] data. Data is all around us. Do you think is there any escape from data now? No, not now. uh when I used to take sessions 5 years you know and 7 years uh [03:59] before you know uh the scenario was little trying to adapt but I don't think little trying to adapt but I don't think so there is now any survival without data can you survive without this data not moving on social media [04:15] and you know one day you know as it says you know if the internet stops do you think your life also stops your your mobile phone is lost everything is lost mobile phone is lost everything is lost Isn't it the data has become a lifeline [04:27] Isn't it the data has become a lifeline you know and now we see several applications which are working or the concepts that are being bas you or the concepts that are being bas you know based on data now it automatic [04:41] translation translation has not become difficult difficult if you want to convert something from English to German to Spanish any language translation is right there you We speak to the machine and we get the translation. Virtual [04:56] personal assistants are there. Image recognition. Email spam filtering that's that's actually comes under the domain of machine learning that it you know based on the algorithm or the pattern or the text or the words which are there. [05:10] the text or the words which are there. It is able to filter out whether the mail is a spam or not. So what do you think would be the criteria? Generally the males with spam are saying that you have a lottery system other system or [05:24] there is a bonus. So that becomes your email spam filtering. Then we have the text and the speech recognition. Medical diagnosis, online fraud detection [05:36] is there you know where we want to detect uh the how is the online fraud happening. Web search uh search and recommendation engines and of course what is going to be the traffic prediction. Not only the traffic [05:49] prediction on the roads but it also relates to the traffic going onto a particular website whether that website is going to get crashed or not. Data becoming now difficult to handle because it's digital data. You know [06:03] numerically is difficult. You know earlier the data was this much it could much and this much and this much and it's increasing. So we need certain algorithm we need technologies which can help to analyze so that we can improve [06:19] our performance. But on if you if you look at on the overall scenario still there is a lot of confusion about the AI the machine learning and the deep learning. Yeah. What is the you know the subset or the [06:35] issues like what is the artificial what is artificial intelligence machine learning and deep learning. Are you able to distinguish between the three? So this is one of the initial chess you know developed by the computers. IBMD [06:49] know developed by the computers. IBMD blue chess program developed in 1997. When was it developed? It was developed in 1997 in 1997 by IBM. And this particular program was [07:04] by IBM. And this particular program was strong enough to defeat the uh world chess champion at that particular point. Uh his name is Gary Kasparov. Okay. He Uh his name is Gary Kasparov. Okay. He was capable of defeating the world chess [07:18] champion. Getting my point right and [snorts] then we have this IBM right and [snorts] then we have this IBM Watson under machine learning. So AI is Watson under machine learning. So AI is the bigger branch. All right. Artificial [07:32] the bigger branch. All right. Artificial intelligence is the bigger branch right intelligence is the bigger branch right under which we have the machine learning subset right which we are going to study. In machine learning basically it [07:46] study. In machine learning basically it works on statistical algorithm. That is why we say that before learning machine learning it is important to have a good concepts and knowledge of statistics that helps to understand. So based on [08:01] statistical foundation, statistical algorithm uh you know the machine is capable of doing Google search algorithm, Amazon recommendation and email spam filtering. How like we just saw that it is capable of filtering the [08:17] email. And then finally we have the deep learning which is also under machine learning in which we have the alpho the natural speech recognition and the level [08:29] natural speech recognition and the level four automated driving system. the AI revolution. So the bigger branch is still AI. We are seeing we are the [08:45] you know witnessing this revolution in front of us. Under AI we have the machine learning under machine learning we have the deep learning and under deep we have the deep learning and under deep gen AI task. Right? So when we talk [09:00] about the gen AI task generation fine-tuning we have the agents fine-tuning we have the agents automation and the virtual assistants automation and the virtual assistants and this gen AI is now capable of even [09:13] doing lot of text generation that we see we ask a lot of things to the chat GPT other uh models like Gemini copilot also we want to image uh generate images [09:25] video generation all these things are being possible AI that we are talking at the moment is only and only related to the software. only and only related to the software. Through software we are able to [09:38] or give intelligent answers. But what is basically the difference between a traditional programming and a machine learning programming. So let me explain you with this particular con uh example. So in traditional programming again we [09:54] have this data right that this is my data which is 1 2 3 and 4 right and over here we have machine right and over here we have machine learning algorithm where we have 1 2 3 [10:10] and four okay the data has not changed I'm giving you a very simple uh you know example now if we talk about the program initially ally the program when we talk about C, Pascal, [10:25] Photron, even C++ and Java and even Python they are capable of logic. Now if I want to distinguish that what are the numbers what is the logic behind that these uh you know numbers are even or odd. So simply we understand the logic [10:41] odd. So simply we understand the logic that if I if I talk about uh Python I that if I if I talk about uh Python I percentage 2 is equal equal to zero that means if the remainder divided by two is zero then the number [10:57] zero then the number it is odd right? it is odd right? else it is odd. Getting my point? [11:12] So now even if a number 10, it would automatically give me that this is going to be an even number. And if I give the number 57, it is going to be odd. But now machine learning how things happen. I give the output along with it. I say [11:28] I give the output along with it. I say that one is odd, two is even, that one is odd, two is even, three is odd and four is even. Got it? Now the computer based on certain statistical concepts algorithm [11:44] certain statistical concepts algorithm will try to detect that when I feed the number 10 over that it is going to be even. Can it predict me as odd also? even. Can it predict me as odd also? Yes, the prediction can be wrong also. [11:58] Yes, the prediction can be wrong also. Clear? But if the algorithm has to be good enough that it is predicting that 10 is even and 57 is odd. 10 is even and 57 is odd. So coming on to the concept the first [12:13] So coming on to the concept the first kind of learning is known as supervised learning. What is it known as? Supervised learning. that if this is my known data, this is already now images [12:28] that this is my input that this is my input and IO feed the output. When I feed the in as well as output to the machine, it becomes my labeled data, right? That [12:44] means this is an image of an apple. I feed it into the machine and when I feed feed it into the machine and when I feed this apple it says predicts that is this apple or not and it predicts it's an apple but it can predict wrong also. So [13:01] might have seen sometimes the chart GPT also predicts gives wrong answer the also predicts gives wrong answer the image uh development jibli and all all give wrong answers. It's the biggest fear you know that these if the if a [13:16] fear you know that these if the if a wrong or a incapable data is uh set to uh them then it can give false and nonsensical and fabricated information also. So this is the f fear that we are now moving ahead right AJ Charardik [13:33] but if we talk about system if we talk about the traditional systems expert systems expert systems were working that we have a user we are trying to give the [13:45] we have a user we are trying to give the query and then try to get the output out of it right if this is my query and this is how I get the output out of it and then I try to infer what is it but inference is not coming from [14:02] is it but inference is not coming from the data but this data it's coming from the data but this data it's coming from data but this data is created from an expert now what do I mean by that let me explain this to you expert can be a [14:16] cardiologist a lawyer in different domains right a cardiologist a lawyer somebody from in the finance domain maybe for 30 years 30 plus years of maybe for 30 years 30 plus years of experience and weated [14:30] if else knowledge you know inference knowledge that whether you are capable of getting a loan or not and if I have a cardiologist some information has been fed that maybe your BP rating is this much or your um terms if that is [14:47] matching to the inference engine that's how it will give the output so what happens for example if you are a user and you enter to the user interface. Maybe all your uh you know blood reports, your BPS and your test reports, [15:02] all things are given as the user interface right and then we do the inference engine and based on this knowledge base engine and based on this knowledge base we get the output right. So what has [15:16] been replaced now rather than knowledge base it is completely based on the original data and things have become complicated on images on you know complicated on images on you know nonstructured data. So uh do we [15:29] nonstructured data. So uh do we understand supervised learning? Now understand supervised learning? Now can I say more the data more my system can I say more the data more my system becomes intelligent [15:44] triangle and square and if I add a shape of a circle maybe a parallelogram maybe a rectangle it's capable of analyzing that. So that is where you know the [15:56] systems are getting modeled because the data is becoming huge day by day right. So this is supervised learning. So under supervised learning we have the label data. What is label data? We have the input as well as the output where we try [16:12] to train the model. After the model has been trained based on the test data we try to do the prediction whether it is a square or a triangle. Clear? [16:26] We will initially start with supervised learning that we have the input the output. We will try the model. So where will we do the model? That means now the will we do the model? That means now the data the 70% of the data will be used [16:42] for training and 20 to 30 uh you know percent will be used for testing. Then it will predict the output. [16:54] Okay. And as I've been telling you, there is a very strong relationship between machine learning as well as statistics part of it. [17:10] mathematics but when it is combined with computer science you know the machine learning it becomes statistics and machine learning. The idea of statistics is that it helps us to draw inferences, relationship between variables. Whereas [17:27] machine learning gives optimization, prediction, accuracy etc. Right? And then we have prior assumptions about the data. Some knowledge about the population usually required. This is none. Dimensionality of the data usually [17:42] applied to the low uh dimensional data. And knowledge overlap. There's no ML knowledge required when we study statistics. But in machine learning, some statistics knowledge is usually needed as it is becomes the foundation [17:57] needed as it is becomes the foundation for few algorithms. So if you do not know much about statistics, there is nothing to worry. It's not very difficult. Definitely concepts of probability would be required. So I [18:09] would just require request you the learners to get familiar with the concepts of probability conditional probability basian theorem and probability distribution. So this is what I expect from you all. Got it? [18:26] Right. And to make the picture a little more clear, you know, the boundaries are not very crisp now because there is data everywhere. But to make the picture a little more clear that you know when we start with initially Python course you [18:41] know we are doing visualization exploratory data analysis maths and statistics and when we try to overlap do overlap between AI machine learning deep learning we get this data science. So data science becomes the foundation for [18:57] data science becomes the foundation for AI ML and deep learning. If I technically ask you what is learning and now you might have been hearing this word agent you know it could be a human agent it could be a [19:11] robotic agent it could be an AI agent this so if we talk about a little more technical definition of learning basically we are trying to improve the behavior based on the experience when we [19:26] say you are a very learner thing you experience that means means uh different types of knowledge. The range of behaviors is expanded. The agent can do more. Right? The range of behavior is expanded and [19:42] agent can do more. The accuracy on the task is improved. The agent can do things better. And the speed is improved. The agent can And the speed is improved. The agent can do things faster. [19:59] see this means that DS okay so if we talk about learning from the machine as well as from the human point of view this definition is valid that it is the ability to improve our behavior based on [20:14] our experience right it's not always about acquiring new skills that's one of the thing that range of behavior that you are you know driving you know swimming you know stitching you know cooking of course the range is increased [20:28] but also So accuracy by doing the thing again and again doing learning increases again and again doing learning increases my accuracy as well as speed. And if we my accuracy as well as speed. And if we now see technically in machine learning [20:41] the different types of learning techniques are supervised learning, unsupervised learning, semi-supervised learning and reinforcement learning. So [20:54] all these of actually learning come from the fact that how we as humans also the fact that how we as humans also learn. So do you also agree [21:07] only agree through training and testing? Training and testing happens when we are going to school we are trained and then we are tested in universities, colleges or even this session. Is there any other way we learn also? Do we learn through [21:22] observations? Do we learn through our mistakes? Do we do we learn? There are different kinds of learning also possible. And exactly the same thing we possible. And exactly the same thing we also try to replicate in our machines [21:38] also. So one of the learning method is observation. So one uh basic difference between supervised learning and unsupervised learning. What do we [21:51] mean by supervised and unsupervised learning? Supervised learning says that this is my data right are my apples right and I tell them that [22:05] are my apples right and I tell them that this is my image these are apples I this is my image these are apples I 70% of my data is used for training and 70% of my data is used for training and 30% of the data is used for testing [22:20] label data is given in absolutely correct mega label data means supervised is learning and then the model predicts me that it is an apple and what is unsupervised learning? Have [22:35] I have do I have a label data? No, I have not told that this is an apple or this is a banana or this is a peach. But the model or the algorithm, [22:47] the model or the algorithm, sorry guys, is capable enough to distinguish between the three of them that this is my apple, this is my peach that this is my apple, this is my peach and this is my banana. Right? So [23:01] initially we will you know build concepts on this supervised as well as unsupervised learning. The machine is given huge sets of data that are not [23:13] given huge sets of data that are not labeled as inputs to analyze. The machine needs to figure out the output its own where it identifies the patterns. And the two types of algorithm which come under unsupervised learning [23:27] are association and clustering and K means for clustering problems and a priority algorithm for association role learning problems. Right? And then we have supervised learning. The input is in the form of [23:42] raw data that is labeled. The machine is already fed with the required feature set to classify inputs divided into two classification. And then we will try to understand [23:57] different regression algorithm. This is the first stage that we are going to the first stage that we are going to work on. Clear? So now let's understand what is selfsupervised learning. You know here [24:11] we have the input data. Is it labeled? No. This is not labeled data. But we No. This is not labeled data. But we have partial label data that this is an orange and this is a banana and its quantity is less. And here both the [24:26] quantity is less. And here both the types of data are mixed machine learning model and this is my unlabelled data to predict the output. So it's an apple. So predict the output. So it's an apple. So this is my input data, right? This is my [24:42] partial data. And when I combine them together, prepare the model, I get the output, clear [24:57] effective way of learning from which we as humans learn a lot. That means from our mistakes, from our feedbacks, from our punishments, from our rewards. Right? So if this is the input given to the machine and the machine predicts [25:13] that it is a mango and I give a feedback wrong, it's an apple, it notes it down. And now when I feed the apple again to the machine, it says that it is an [25:26] apple. So take making the overall picture a quick recap that basically you know if we talk about machines there are three types of learnings we have the supervised learning where we have the input and output and this supervised [25:41] learning is capable of calculating error that is output minus the input uh uh you know and what is the because why are we capable of calculating error in [25:53] supervised learning because we have the actual output over here, right? And based on the actual output, is my machine predicting the correct result that if it is an apple, is it actually telling me an apple or not? Or is it [26:09] telling me this is a cherry? So I can calculate my error. So the biggest advantage or the simplest way to learn is supervised learning where we are is supervised learning where we are capable of calculating the [26:24] errors. Right? Here we have unsupervised learning where we do not know the output. It's just trying to do the clustering and association between the objects. And the other type is reinforcement learning that is not going [26:38] to be part of this journey that's generally taken into deep learning concepts that where the machine learns from its you know uh punishments and rewards right and of course that we are capable of calculating error. Again I'm [26:53] telling you in this particular course we are going to try to cover supervised and unsupervised learning uh you know concepts algorithms in detail. Now concepts algorithms in detail. Now moving ahead to supervised learning. [27:07] There are two types of supervised learning. We have regression learning. We have regression and then we have classifification. [27:32] classification. Now what is the difference between the two? Please try to understand. Under supervised learning, if the output, it's all about learning, if the output, it's all about output. If the output is numerical [27:55] then this is known as regression and if the output is categorical then it is the output is categorical then it is known as classification. Getting my point learners? What we are trying to achieve that what [28:09] is going to be the temperature tomorrow. So if it gives me that tomorrow it is So if it gives me that tomorrow it is going to be 84° Fahrenheit or any other going to be 84° Fahrenheit or any other value then this kind of algorithm is [28:23] regression. But if I want to predict whether the temperature is going to be whether the temperature is going to be cold or hot this is known as categorical data. Clear? [28:36] So the regression works on numeric and classification works on variable and categoric. Two types of numerical data. One is Two types of numerical data. One is discrete and other one is [28:48] discrete and other one is continuous. Do we understand that? And if we talk about categorical data, do we understand nominal and ordinal data? This is what is covered in data science [29:02] class. So basically you know data types are divided as qualitative and quantitative very very important when you um divide the data as qualitative it's categorical order data is something you know like [29:17] rating ranking they come under order feedbacks of uh like movies nominal is that there is no order as such the color of the eyes or nationality and [29:29] quantitative is numerical values continuous which can be divided ed such as distance, salary, price and something which cannot be uh divided is example cats etc. So now I hope the concept of regression and classification is clear [29:44] to everybody right. So uh making your concept more clear that regression is something when the task of predicting a continuous uh you know quantity that I [29:56] want to predict the price of the house that price of the house in 2014 that price of the house in 2014 was this much then in 2024 it is this much right and then what will happen what will be the price in 2034 [30:14] Clear? So since price is a numerical quantity, it comes under regression and classification as I told you that if I want to separate whether the male is a spam or not, that comes under a [30:29] spam or not, that comes under a classification problem. Clear? So now we classification problem. Clear? So now we can we will begin our journey in this machine learning through supervised learning. First we will try to complete [30:42] algorithms which are like simple linear regression, multiple linear regression, polomial, support vector, decision tree, random forest. We are not covering neural network. All the others will be covered. Similarly in classification we [30:55] will cover logistic K nearest neighbor support vector machine nave bias decision tree random forest. Again neural networks are not covered. Yeah. neural networks are not covered. Yeah. So with this uh we come to the end of [31:09] the introduction and if you will now look at your u slide the ebooks that is lesson number two now we can begin with lesson number two so just I've prepared [31:21] the foundation for that so let me show you the lesson number two all right so we are now starting with this I hope now you will be able to locate this particular file in your LMS in your material Please look at that. So analyze [31:37] the distinctions and applications of machine learning, deep learning, artificial intelligence through the real world examples of various technical applications. Differentiate among various machine learning models and [31:49] explore each model learns from data to predict outcomes. Explore Python libraries for effective data manipulation, visualization, implementation, and machine learning algorithm. So the business scenario says [32:03] that ABC is an e-commerce company which is struggling with a surge in fraudulent transactions on its website. The manual review process for transaction has caused delays in order processing and led to negative customer experience. To [32:18] address this ABC will use machine learning algorithms to detect realtime fraudinal transactions. So machine learning algorithms are capable for detecting the fraud transaction and these algorithms will be integrated into [32:32] the company's transaction processing systems to flag suspicious transactions and prevent fraud. Additionally, the company will use these algorithm to predict the customer behavior on the past purchase history thereby improving [32:45] the recommendation engine's performance. Now everybody is using mixture. Everybody wants the best. You know the more you learn the best output you get and when I say best output you want the prediction to be highly accurate. So [33:01] when your report goes to a machine learning or a AI machine, it it has to give accurate result that whether you have whether you are you know that accurate enough to predict that you know yes you know your you are you know uh [33:19] yes you know your you are you know uh cap you have a tendency to have cancer or not you know so no so not only one technique will be used it will try to use mixture of techniques to get the best results. Got it? Now, so now are we [33:33] clear? What is machine learning? So, machine learning is a subset of AI that machine learning is a subset of AI that assist systems to learn and improve automatically from the experience without being explicitly programmed. [33:49] Arthur Samuel coined the term machine learning in 1959. It enables programs to learn automatically making computers more automatically making computers more intelligent without human intervention. [34:04] But human feedback is extremely extremely important because see ultimately machines are not genius. We have to tell them that this is you know have to tell them that this is you know a an acceptable result or not. So who is [34:18] known as the father of machine learning? It's Arthur Samuel and he coined the It's Arthur Samuel and he coined the term machine learning in 1959. Father of term machine learning in 1959. Father of AI that's John Mcathi in 1956. [34:32] AI that's John Mcathi in 1956. Yes, John Mcathi in 1956. difference between the traditional approach where we had the data and where we were programming explicitly the output right and the machine learning [34:48] approach based on the algorithm based on the data the patterns it understands it predicts the output that is why statistical techniques algorithms are the foundation for machine learning it automatically learns the features and [35:05] automatically learns the features and reduces the need of manual featuring clear. It handles complex and unstructured data such as images, text, audio without requiring extensive [35:19] pre-processing. Performance improves with more data and learning iterations with more data and learning iterations enhancing accuracy and generalization. Right? And now do we understand this ven diagram also that machine learning, deep [35:34] learning, AI are often used interchangeably. So AI encompass encompasses simulation of human intelligence and machines. So self-driving cars are you know AI all the robotics come under the category of [35:47] the robotics come under the category of AI but machine learning Amazon Alexa where we are giving it specific instructions and it gives us output and when I talk about deep learning deep learning involves neural network which [36:00] is going to be your next stage after uh you know machine learning to understand how neural network uh you know work how are they capable of understanding complex pattern recognition such as recognizing patterns in images, speech, [36:15] recognizing patterns in images, speech, text etc. So where are what are the examples of machine learning in a chess game between a computer and a person? Why do you think that why is this chess game always coming into the picture? Why [36:29] game always coming into the picture? Why do you think is the chess game always coming into the picture? Well, when we talk about human beings, a person who plays chess well is said to be intelligent. It's an intelligent game. [36:42] Agreed. It's said to be intelligent and even if the results go wrong, there's no harm. There are no catastrophic results. You know, even if the person is winning and says the machine is winning, it is not [36:55] that harmful. Yeah. So that is why you know that that is was one of the way where AI exploded in you know intelligent gaming systems theorem solving that is why chess alpha go these are the games which you will you know [37:11] read or see when we talk about a little history of AI. So in a chess game uh between a computer and a person, the computer uses AI to analyze the game, predict moves, decide its decision. AI [37:27] uses machine learning to figure out the opponent is a beginner, intermediate, opponent is a beginner, intermediate, and an advanced level. How the AI uses machine learning to identify whether you are a beginner, intermediate or an [37:41] advanced player? By predicting our moves, you know, based on our moves, it will immediately judge, you know, immediately judge that you are a beginner or a intermediate or an advanced level, right? And you might be [37:55] playing a lot of games where there is AI and other you know graphics games over there, you know, especially the young generation, right? Well, I don't, but you can, you know, how smartly, you know, they are capable of hiding things [38:10] and they become smarter. The level changes you know as you are also becoming smart the level of the game becomes smarter right we all observe becomes smarter right we all observe that so AI decides its next move against [38:23] the oppo opponent using a complex neural network that learns various features network that learns various features patterns from the data right then of course many applications are [38:36] there we see machine learning all around us in spam filtering spam filtering actually use Drive based classifier, social media analysis, customer service, now a lot of available online recommendation. We all see that [38:51] sentiment analysis. How does the sentiment analysis happen? Anybody who has an idea based on the words and especially the emojis. Is it a words and especially the emojis. Is it a smiley? Is it a angry face? Is it a st [39:05] sad face? Crying face. All these things are taken into account. So these patterns from the data, make predictions, predict outcomes, classify [39:18] predictions, predict outcomes, classify target feature and improve performance. So that's what uh we want machine learning that they help to predict the outcome, what is going to be the price of the house or any other thing after 10 [39:30] years down the line. classify target features based on the similarities and improve the overall performance of the system. And we have seen is there a very very strong relationship between the data and [39:44] the output. Yes, the amount of data if it is more of course the output or the quality or the prediction also increases where the red line is uh the quantity [39:57] and if if it is high quality data we get better results better insights and better predictions of the output. Of course maintaining the quality, authenticity and removing the errors. All these [40:12] points have to be taken into account when we are talking about the data and the machine learning algorithms. Clear? Right. So when we talk about types of machine learning, ML can be divided into [40:26] four main categories each characterized by its capacity to predict the our conditions or identify the patterns to produce outcomes such as what is supervised learning, unsupervised learning, semi-supervised [40:41] learning and reinforcement learning. Now I think the distinction is clear. What are the three main points under supervised learning? [40:53] supervised learning? First is the label data. What do we mean by label data? It will have the input as well as the output. Second point is well as the output. Second point is splitting of the data into training and [41:08] testing. Right? Generally training happens on most of the data and testing of on the rest of the data. And third is calculation of the error because in [41:20] supervised learning we know the actual output also and whether it is predicting output also and whether it is predicting it right or wrong. Clear? And some commonly known uh you know uh supervised learning algorithms are linear [41:33] regression, decision trees, logistic regression, support vector machines etc. regression, support vector machines etc. So some examples of supervised learning are predicting temperature rise based on the yearly temperature trends. [41:48] Predicting why supervised learning again predicting temperature rise it comes under regression because it's a numerical problem right numerical output predicting crop yield based on the [42:02] seasonal crop quality changes. Again regression sorting waste based on the known waste items corresponding to the waste type. This types comes under waste type. This types comes under classification or filtering. So under [42:15] supervised learning calculation of error is also going to be an important aspect of machine learning because we have to be very clear that the output of machine be very clear that the output of machine learning will always not be 100% [42:27] correct. Clear? Now this example are these example making the picture more clear. And if we talk about unsupervised learning very much used to you know uh identify different parts of the object image [42:42] segmentation for object detection. Identific identification of user groups based on commonalities. Identification of anomalies over geographical landscapes based on the data patterns. [42:56] The unlabelled data set is provided to an unsupervised learning algorithm to discover hidden patterns and to recognize their relationship. So it's not that unsupervised learning is not important. It is equally important to [43:11] analyze different features, different relationships in the data. Right? Rather this is the algorithm which helps us to discover the hidden patterns and discover the hidden patterns and discover the relationships. Clear. And [43:27] now coming on to the unsupervised learning example. It automatically groups the images based on the similarity that this is unlabelled data. Based on that it is capable of distinguishing the middle-aged people, [43:41] old age people, the young, the infants, the teenage etc. Got it. And what is semi-supervised learning? As I've already told you, it uses a combination [43:53] of small amount of labelled data. Sometimes, you know, the data is that's one of the constraints that we see in data is not completely labeled. A large amount of unlabelled data is used for training. Like supervised learning, it [44:07] aims to learn from a function that can accurately predict the output variable from the input variable. It uses the unlabelled input to assist the learning process by collecting more information improving model generalization. So it [44:23] falls between supervised and unsupervised learning. So suppose this is my raw data and I have partial that this is adults and these are kids then the machine automatically distinguishes between babies, teens and tween and also [44:39] distinguishes between the senior citizens, youth and adults. Got it? So it automatically learns the correct So it automatically learns the correct uh you know uh groupings of the kids uh [44:53] or the different people into teens uh tween and babies and the adult ones into tween and babies and the adult ones into into these category and another example of semi-supervised learning which we see it practically [45:09] photos is popular example of semi-supervised learn learning that when a picture is taken it gets stored in the Google cloud platform and from slowly the Google tracks you know whose picture it is at what place it was taken so in [45:24] various instances uploaders label images despite Google's lack of knowledge regarding image names its algorithm can identifying by analyzing visual features and shapes and colors and it does that it does a lot of lot for me it is able [45:39] to it's able to identify my friends my family in which location I was and where and reinforcement learning is a type of machine learning where algorithms learn from the environment by performing actions and receiving either rewards or [45:53] penalties as feedback. If the pro program finds correct solution, the interpreter rewards the algorithm. If the outcome is incorrect, the algorithm is penalized for incorrect predictions. It must I reiterate until it finds a [46:10] It must I reiterate until it finds a better result. Right? So ultimately reinforcement learning involves an agent. It interacts with an environment. Learning from the rewards and states to choose from and then based on the output [46:26] it gives the best action and if it is an error it learns it again. All right. So this is my input raw data based on the environment reward state and action. [46:38] it's capable of detecting them separately. Okay. So the example is this type of learning is seen in YouTube recommendation where the user searches for a particular song. The program shows [46:52] the list of available song. So when a user selects a specific song, the system trains itself to remember and deliver similar results for future s searches based on the user's interaction like views, shares etc. So this is the [47:08] views, shares etc. So this is the concept on which recommendation systems So other examples of reinforcement learning are game where players can play with bots, order correct tools, search recommendation income uh engines, [47:22] self-driving cars and then the Python packages that we would be doing in for machine learning. We are aware about NumPy. It's a very powerful tool for numerical Python computing. Mattplot lib [47:35] for drawing data visualization pandas. So I hope you all are aware about numpy mattplot lib pandas but we will be working more on the scikitlearn uh uh you know uh file which consists of different algorithm but the [47:50] pre-processing the other part is also being taken care before we feed in into the algorithm. So a quick recap. So machine learning refers to the machine's ability to learn from the data and replicate human behavior. AI [48:06] learning, each with unique capabilities four main types of machine learning, supervised, unsupervised, semi-supervised, and reinforcement learning. Python packages are folders [48:19] with modules that organize code for easy reuse and maintenance, improving the de reuse and maintenance, improving the de development efficiency. Now clear. So now let's go in for a knowledge check. Question number one. [48:35] Yes learners are you there? Which of the following best describes the machine following best describes the machine learning? A, B, C, and D? Question number two. Which example illustrates the use of machine learning to enhance [48:50] customer experience in an e-commerce company? What distinguishes between deep learning, machine learning and AI? Yes, it is a subset of ML that uses multiple layers for complex pattern recognition [49:03] such as recognizing patterns in images, speech and text. Right? So with this we come to the end of the very introduction and basics of [49:15] machine learning. Yeah. So what I'm looking forward is this is the overall uh picture of statistics that uh we have the learners [49:29] who have already done data science are aware about it. The ones who are not aware about it. The different types of statistics that we have is descriptive. Under descriptive we have measures of central tendency and measures of [49:41] variability. Under measures of central tendency we have the mean, mode and median. Under measure of variability we have the range, variance and dispersion. And here we have the inferial statistics how we infer the [49:56] results. So this is the important part that we are looking at that that includes confidence interval hypothesis testing. So you can do a lot of search testing. So you can do a lot of search on inferial statistics and uh this is [50:10] what I am saying that uh statistical inference constructing confidence and intervals on population hypothesis testing that is what I'm looking and what is the advantages of these particular uh program because ultimately [50:26] you know probability in data science and AI play a very very major role in understanding uncert certainty predicting outcomes how probable correct [50:38] the output is even the LLM models the chart GPT is predicting on the probability okay this is the next word so if it is 70% above then let's predict the outward how do we model complex system enhancing AI and it is the [50:55] statistics and the probability together which help in exploratory data analysis and give meaningful full insights and features. So as we are going through [51:08] this flow we understand we are going to start with supervised learning. What is supervised learning? Before supervised learning you know uh I would like you to cover we would uh you know the cover the basic concepts of [51:21] machine learning which I have prepared through my PPT and then we will move on to what they have shared. Okay. So the first thing is regression. You know what are the two main algorithms which come [51:35] are the two main algorithms which come under supervised learning? Yeah, regression and classification. What is the difference between regression and classification? Regression happens when the output is [51:47] numerical and classification happens when the output is categorical. All right. So I start with my PPT again so that it helps and it gives you strong foundation clear concepts so that we can [52:03] move along with that. So if we talk about the first knowledge check what is about the first knowledge check what is machine learning? Correct answer it is see that it is an autonomous acquisition of knowledge through the use through the [52:18] use of computer programs. Second question, what is the key difference between supervised and unsupervised learning? A and B are both correct. What's the key benefit of using deep learning for task [52:36] like rec? How do I u Okay, you've written yes. Okay. What's the key benefit of using deep learning for task like recognizing images? Yeah, they can learn from complex details from data on the own. That's the [52:53] idea of deep learning. That's the use of neural network. Absolutely correct. All right. So now we start with the concepts of machine learning under supervised learning which are valid for regression [53:08] and classification algorithms. Whenever you will say you know machine learning these are the basic questions that will be asked right. So if I talk about uh machine learning or supervised learning what are we trying to uh you know do in [53:24] supervised learning what is the main aim what is our main objective of supervised learning if but prediction of the data. Okay, we would call it prediction of the [53:36] data future based on given data. Right? We want to predict and is this prediction always correct or it can be wrong. Can be wrong. They can be errors. Right? Right. And so they there should be a [53:51] limit of accepting the errors or rejecting the result. Right? There should be some way of accepting and rejecting the results. We all understand this very in a conceptual subjective matter. Now let's try to understand it [54:05] matter. Now let's try to understand it on the basis of mathematical functions. on the basis of mathematical functions. Right? So here we have the plus over here and here we have the minus over here. Right? This is my data set. This [54:18] here. Right? This is my data set. This is my x-axis or and this is my yaxis. is my x-axis or and this is my yaxis. These are given as my data. Right? And what are we trying to do over here? We are going we are trying to predict the [54:33] output. So what are we trying to predict the data that what is the value of this question mark? What is the value of this question mark? question mark? All right. And what is the value of this [54:46] question mark? Right? Is this question mark a plus sign or a minus sign? Can you tell me what is this question mark? A plus or a minus sign? What about this? What about this? So this is my first question mark. [55:01] Second question mark. Third question mark and fourth. So based on your observation, can you tell me what is this first? Do you think it is a plus or a minus sign? What does this particular data point [55:16] represent? Is it a plus or a minus? Because this this particular data point is more close to the plus. So there are chances that this is going to be plus sign more chances. Yes, it can be negative also but we can say 70 to 75% [55:33] or 90 to 95% chances are that this is going to be plus. What about the fourth one? This is negative. Now what about the second and the third one? What about the second? [55:47] What about the second and the third? This is going to be a little difficult uh to say that this is going this can be plus this can be minus based on the way the neighbors I am selecting what will be my output agreed but but if [56:03] be my output agreed but but if technically now if I use a straight line function ma mathematically how do we use a straight line function a straight line function that y is equal to mx + c. So any all [56:19] the points which lie on the left side of this line they all are known as plus they all will come under the category of plus and the data points. Now all the [56:32] points which are lying on the right side of the line they will be all negative because most of the points are minus or the red points. Can I say that all the [56:44] points which are lying on the left side are plus and all the points which are lying on the right right side are negative. There are very few now this is there is only one plus sign and if you see all the points belong to the red red [56:59] negative class one way I have to do something but again the question is why something but again the question is why this straight line that how do I decide this straight line that how do I decide the straight line? So one point to [57:13] understand is that when data points are given that is known as a hypothesis space represented by capital H and the line is represented by small H [57:33] line which is going to divide the data points into two different classes. So there can be more than one solutions to a problem or infinite solutions. I can draw infinite straight [57:48] lines. But which one to accept? The line which But which one to accept? The line which will be accepted is going to be one with will be accepted is going to be one with gives me the minimum error. [58:01] Right? The error word will be different for different algorithms. But I will accept that straight line which will give me the minimum error or I can say the maximum accuracy. [58:16] Clear? We will calculate as I told you in supervised learning we keep the track how we can have maximum accuracy and minimum error. So what is the capital H over here? [58:29] Hypothesis space is the set of all possible legal hypothesis. This is the algorithm would determine the best possible only one which would describe possible only one which would describe the target function and small h is a [58:44] hypothesis function that best describes the target in supervised learning algorithm. So now technically what does supervised learning mean that the hypothesis or the small edge there can be several small edge that an algorithm [58:59] would come up depends upon the data also depends upon the restrictions bias that depends upon the restrictions bias that we have imposed on the data. So next thing is we need to find out mathematical functions which give [59:14] mathematical functions which give relationships between the data points. Right? So basically now the idea is that we want to find a mathematical function [59:26] we want to find a mathematical function where you know uh we want the error to where you know uh we want the error to be minimum or the result to have maximum be minimum or the result to have maximum accuracy. Again I'm repeating we want to [59:40] accuracy. Again I'm repeating we want to find a simple mathematical function find a simple mathematical function which gives me minimum error or maximum accuracy. Clear? Thank you at Thank you for [59:53] understanding. Thank you learners. So now what are the different stages that we are looking at? Now suppose data points can be of any kind. So if this is my data points over here and if I want to solve it with with a simple linear [01:00:09] to solve it with with a simple linear function see straight line function is a function see straight line function is a simple mathematical form uh you know simple mathematical form uh you know calculation that y is equal to mx + c. [01:00:23] calculation that y is equal to mx + c. Okay. So when y is equal to mx + c right and if these are my data points now how do I calculate my error? This is [01:00:35] going to be actual minus the predicted. It will have very high error and this particular concept is known as underfitting. Do you think that this [01:00:47] straight line is consistent covering all the data points? No. Right? So the simple straight line equation or function is not capable of [01:00:59] covering all the data points that is known as underfitting. Other way is that if I draw a function which passes through all the data points so it makes it a very very highly complicated mathematical function with [01:01:16] high degree. Right? So it makes a complicated function with highderee mathematical function that also we don't want. Why we don't want because if the want. Why we don't want because if the data point is out of all the data points [01:01:31] then this is going to give me my maximum error. So we want a mathematical function which is simple and which covers all my data points. So something [01:01:44] like this exponential right and this gives me a good fit or a good balance now clear. So AJ says no. So the model even fails to predict labels what it learned right [01:01:57] it will it will predict the labels but with very low accuracy the error will be high. So we we will not accept models with high errors or low accuracy. [01:02:09] with high errors or low accuracy. So overfitting is a situation. How do I know the error is high overfitting? So one way is that error is equal to if I [01:02:21] talk in technical terms. So how do we know that it it is underfitting overfitting? So that's what I'm telling you. Generally the error is I'm telling you. Generally the error is the formula for error is [01:02:37] the formula for error is bias² variance. What is it equal to? It is bias squared [01:02:49] plus variance. But the equation is in good balance not linear. Yeah. It it it's not necessary that it has to be a linear equation. No, we want an equation which is simple and nice. So how do we know? So now if I say that error is [01:03:06] know? So now if I say that error is equal to bias squared plus variance over equal to bias squared plus variance over here right and this is you know so how underfitting it's not about the straight line but when I have variance if if the [01:03:22] variance of the data goes very high then I know that it is an overfitting model I know that it is an overfitting model if the bias of my data goes very high I know it's an underfitting model and to keep my error low I want the bias as [01:03:39] keep my error low I want the bias as well as variance to be low. Getting my well as variance to be low. Getting my point right now what do we mean by bias and variance? A very very typical example. A [01:03:53] variance? A very very typical example. A bias means how far are we away from the bias means how far are we away from the original data points. So bias means we are close to the center. Variance what does variance mean? The spread of the [01:04:05] does variance mean? The spread of the data is also low. So this is the ideal situation we always want the data points to be to be in. This is where we will [01:04:17] to be to be in. This is where we will say okay a good fit has been achieved where we have the low bias and low variance. variance. Okay. And when we have the low bias and [01:04:29] when we have the high variance low bias means that the data points are near to my actual target points but spread out. So what is the [01:04:42] points but spread out. So what is the case happens over here? This is known as case happens over here? This is known as the overfitting case. [01:04:55] And what is high bias and low variance that over here these are my data points and this is high bias. So this is underfitting situation. So in whole of machine learning we don't want underfitting we don't want [01:05:11] case where the bias is more variance is both. The idea is to have low bias as both. The idea is to have low bias as well as low variance. Clear? Okay. Now pre I'll explain the previous example with this explanation after this [01:05:27] pane I think. So this will make things more clear. more clear. So these are my data points over here a linear model is not a good fit. Under fit that means over here what is more my [01:05:42] bias is more. What is bias? That my data points are very far away from the actual points. Bias means we are far away from the [01:05:56] Bias means we are far away from the actual points. Right? This is high bias. And what is variance? Variance is when the spread of the data is more. Right? And when spread of the data is more or mathematical function is too complex [01:06:14] then it is overfitting. When my mathematical function is too simple, it is underfitting and when it is a balance, it gives me the right data. [01:06:26] balance, it gives me the right data. So now when I talk about complexity, now when I say overfitting means more complexity that generally linear models [01:06:38] give me underfitting, right? We start with simple models but generally they do underfitting. Then we will move on to nonlinear models, support vector machines, treebased models, deep learning models. Why are we in doing [01:06:54] this? Because it they give me more results, better results, more accurate result, less of error. But what is the cost that I my model has become [01:07:06] complicated. I'm losing the interpretability of the model. What do I mean by interpretability? What do I mean by interpretability? That if this is my input, right? How do [01:07:20] output? That's possible in linear model because I know y is equal to mx plus c that if this is my input, I will get this output otherwise I will not. Clear? [01:07:33] And to make it more clear you know the overall picture that I was talking about please look at this slide. So ultimately where is the trade-off? Where are we fighting in the whole of machine learning models? If life was so [01:07:49] easier wouldn't have the problems be solved by now but no there is still a trade tradeoff which is going so how will I know it is underfitting though it is showing in diagram is there any yes yes they are mathematical way but first [01:08:04] let's get the grasp of the thing okay that ultimately the idea is to get highest accuracy right and interpretability is also important Why? Because we want to see why am I [01:08:20] getting this output. So we always start with linear regression or logistic regression problems. That is what we do. That is what we are going to do in our uh whole of this journey. And why uh regression linear regression? Because [01:08:36] they are linear and smooth well definfined relationship easier to compute. And as we move along this journey, decision trees provide good accuracy with high interpretability. KNN clustering clustering and KN&N [01:08:51] algorithm are mid-range. Interpretability is okay and accuracy is Interpretability is okay and accuracy is also okay. But if I want really good accurate results that is why new algorithms complex algorithms were built [01:09:04] with kernel based approach for support vector machines onsembled methods and of course neural networks now clear [01:09:16] is this slide getting clear to everybody? That is why complicated neural networks are capable of solving nonlinear relationships, non smooth [01:09:28] nonlinear relationships, non smooth relationships and long computation time. relationships and long computation time. Clear? Okay. If I want interpretability of the models, then I'll might go in for decision trees, but I will have to [01:09:40] compromise on my accuracy. But if I am not interested in interpretability, that we are not. We always want best result. That is why neural networks, deep learning have taken the market. It is we as users who will decide. What is the [01:09:55] difference between accuracy and interpretability? Interpretability is how am I going to find out the relationship of output given a particular input. So if I know a mathematical equation okay this is [01:10:09] mathematical equation okay this is related with beta x1 beta_2 x1 I know why am I getting the result. But if it is some integrable model like neural is some integrable model like neural networks working in layers in different [01:10:22] differentiation integration addition subtraction is happening I will not be able to interpret the result but it is giving me highest result then of course I will use it if I have a very mathematical complication complicated [01:10:35] mathematical complication complicated system. See now try to understand the whole story again. Let me just uh repeat the story whole of machine learning and these concepts are valid for deep learning [01:10:49] concepts are valid for deep learning also. Okay. The story says that we want to predict the output or the data. You're clear. The first point says we want to predict the output of the data with minimum error or maximum accuracy. [01:11:05] with minimum error or maximum accuracy. First two points are clear. accuracy and minimum error. That point is clear, [01:11:19] right? And which is the simplest mathematical function that we have. We mathematical function that we have. We always talk about a straight line. If that straight line is giving me minimum error, is it? So is it that only [01:11:32] we will always apply one algorithm or is it a different algorithms or different hit and trial methods that we have to apply? It is different and trial methods algorithm that we will try to get the output. It's not just you know oh god [01:11:46] regression and decision tree is working. Okay, this is the end of it. No, you might go in for advanced algorithms and that's where the research is going. Why do you think that the problems are not solved? because every time every [01:11:59] algorithm will have its own pros and cons and the research is going at each and every level and it's still going you know we are just trying to improve on the algorithms every time why because we still see the [01:12:13] every time why because we still see the LLMs the chat GPT giving us wrong answers wrong predictions so there is still a lot of scope of incr improvement you know where we have to work [01:12:27] work clear. So life was if the thing was so easy I think so by now everything would have been solved by AI and machine learning algorithm. No why is it not simple? First data is variable uncertain [01:12:43] it keeps on changing. Secondly every algorithm has its pros and cons. So the errors the accuracy the the the my requirements upon the product you know on the data or the algorithm keeps on changing and that is why this change [01:12:58] happens right. So they are pilotra of algorithms which have been built. As in machine learning we will always start with linear regression algorithms because of the simple mathematical function and [01:13:15] high interpretability but they have minimum accuracy. But as we move along with the journey of decision trees clustering kernel base we have we want maximum accuracy. But the trade-off is that we decrease on [01:13:32] the interpretability of the model. Right? So as I have been telling you splitting the data for machine learning in supervised learning, we always and [01:13:44] always want to select the data randomly. So there is a function train test_plit that we would be using to split the data. train [01:13:57] test_split which is part of the skarn library. This is my data and we will be randomly selecting it and and we will be randomly selecting it and dividing into training and testing data. [01:14:11] Now clear and most of you you see this is around 70 to 80% of the data [01:14:25] and this is around 30 to 20% of the data. interpretability is more important as we go to complex problems. Not really. go to complex problems. Not really. Not really. you know ultimately punit uh [01:14:40] not really because we are always interested in the final result it's interested in the final result it's something like as I give my example suppose you know you have a magician and you know the magician turns a pigeon [01:14:52] into a flower are we interested how does how does it does or are we interested in how does it does or are we interested in the output as a flower okay so we are always interested in the output got it now better yeah so if [01:15:07] logist istic regression is having minimum accuracy minimum accuracy then it is a no no [01:15:19] doesn't mean that it's a wrong prediction logistic regression might not give you the best results for the prediction and maybe using a neural network would give you you know it's something like n uh [01:15:35] you know using logistic regression is giving you 70% of accuracy and neural network is giving you 99 97%. So that's the difference but of course neural networks are better that means right got it so yeah there are a lot of [01:15:51] terms which you have to be associate you have to understand it in the correct sense I understand I'm giving you time for doing that yeah now coming on to the part that you all are interested that ma'am what is [01:16:05] the mathematical equation you know how do we know the mathematical term which is used to see the output. Now first of all we are going to [01:16:20] calculate error right in regression the error is known as mean square error. Now let's start understanding this. So we have this mean square error. Okay the [01:16:32] error name is that that I'll explain you technically when we move on to regression. Now in supervised learning whether it's classification or regression we know we we know the data is divided into training and testing. [01:16:45] Now tell me which error is more important. Now let me put it in this way that you go to a class you know and you are pursuing some course you know you [01:16:57] you you do errors you are learning and you do errors while training right and still you don't perform well in the test and there is somebody who doesn't attend [01:17:09] even one single class and performs well in the test. So which error is more important? It is the training error which is important or the testing error. Among the two of course both are important but among the two which one is [01:17:24] So what is more important to have the training error less or the testing error training error less or the testing error less? The training error is less but the testing error is very high. Is it a good thing? It's a even big [01:17:39] it a good thing? It's a even big failure. the training error important or the testing error important? [01:17:53] Of course the final test if you're not performing well on the final day of the performing well on the final day of the test then the whole training is useless. Agreed? So when my training MSSE is high, of [01:18:07] course I can reduce my training by giving in more data, more data and my giving in more data, more data and my testing MSE is less, what is this case? testing MSE is less, what is this case? This is the case of underfitting. Of [01:18:19] course, the training error should not be high than the testing MSE. The model is too simple for it to solve and then I will say the model is underfitting. Now [01:18:33] clear that was the question I think so so now clear mathematically so this is how I will I will judge based on my training and uh and testing that my model is underfitting so I will now not use a straight line to solve my data [01:18:49] use a straight line to solve my data points I will not use okay other is my points I will not use okay other is my training MSE is zero excellent but my training MSE is zero excellent but my testing MSE is so high absolutely [01:19:03] not acceptable case. So this is the case of overfitting which is high variance. And what do you think the data or the model goes through? Is it more of underfitting problem or an overfitting problem? What do you think most of the [01:19:20] models and machine learning face which kind of problem? Reducing the training kind of problem? Reducing the training error is easy but failing on the testing is a total failure and that is what we will will we will look at most of the [01:19:36] algorithms that how to prevent overfitting in the models that is why I was telling you these concepts are very very common to machine learning deep learning models so are you getting a grasp of it and I don't straight away [01:19:52] start with regression come on let's start doing at other things. You have to have good foundation of machine learning concepts so that you know you understand other materials the indepth knowledge how we accept and reject the particular [01:20:07] model. Training and testing that we will decide because when the data is there we will split the data into training and testing and then we will calculate the training error separately and the testing error [01:20:20] separately. Okay. We will have the suppose I said 1,000 rows 700 rows will be used for training then we will have uh 300 rows used for testing then [01:20:32] calculate the errors for training and testing separately and then decide overfitting. So what will be the best case where my training error is also low testing error is also low or almost equal to each [01:20:49] is also low or almost equal to each other. So again looking at the concept other. So again looking at the concept xaxis we have the [01:21:01] predictive error. So over here on the x-axis this is my underfitting x-axis this is my underfitting over here and if what is overfitting that when I start with my model simple model my errors are high both the bias [01:21:19] and the both the test and as well as the training as my training increases the training error or the model complexity increases the training error decreases but the test error will dip. and then increase because of high variance. So we [01:21:37] want to find out a model. How do I decide which model is best for my data where the training and the testing error are minimum or close to each other. So [01:21:51] this is what I'm trying to explain you. Please try to understand that as the model complexity is less here we have the bias more as the complexity increases the bias decreases. When the interpretability is high the variance is [01:22:07] low and as the complexity increases the variance increases but we are looking at variance increases but we are looking at a point where these two inter intersect a point where these two inter intersect to find the optimal model complexity. [01:22:20] The variance is the spread of the data. Variance is a statistical term which means the spread of the data. So as my model complexity is going to increase model complexity is going to increase the the the complexity as the complexity [01:22:35] of the model increases the variance or the spread of the data also increases. Got it? That's the relationship. Again I am telling you we are trying to create a [01:22:47] graph between error or accuracy is going to be opposite of this and this is my model complexity. We always start with minimum complex We always start with minimum complex model. This is the case of underfitting. [01:23:00] Agreed. when my model is simple but my bias is high as the complexity of the model increases my bias decreases but my variance increases so I don't want the [01:23:16] overfitting situation or the underfitting situation I am looking at a situation where my error is minimum so the test error is U-shaped in curve and [01:23:29] where the value is minimum that becomes comes my optimal model complexity. Now clear to everyone? Okay. Now uh okay now since the uh topic [01:23:42] has been taken now let's do one thing. If you have downloaded the uh ebooks material I want everybody to make a folder in the desktop and move under lesson number three. [01:24:01] something like this. See what I would see is let me share my screen in work in data brick. How do I understand that? So once you are into now we starting with lesson number [01:24:15] now we starting with lesson number three. So do you see this 3.1 and 3.2 learners? So how do you go about opening it? Open the Jupyter notebook it? Open the Jupyter notebook environments. install and open the file [01:24:31] along with me. So chapter number three, we have to start with lesson three. Lesson three is supervised learning regression and its application. regression and its application. Yeah. Are we all able to open this? Let [01:24:46] me guide them that how do we go about it? See the prerequisite for the course I've been telling it is Python. So we always we like there are lot of tools like PyCharm, VS Code, you can do it through uh other tools but generally we [01:25:01] through uh other tools but generally we use Anaconda Navigator download. Please go ahead and go to the Anaconda, fill in your uh like [01:25:13] the details, the email id and I want everybody to download the Anaconda right and uh based on your uh system whether it's Mac, Linux or Windows and [01:25:26] go in for a full distribution don't go for mini Google Collab if you're aware about Google Collab. [01:25:40] available you know which helps you to run the code you know which helps you to run the code in the Jupy Jupiter kind of environment. Let me show it to you. Anaconda distribution. Go and download it. [01:25:53] Install it. Or the other one is go to Google Collab over here. Just login through your Google account. That's the another uh way to go about it. And from here I will upload the [01:26:10] it. And from here I will upload the file. Which file? The file which I have already downloaded on my ebooks that's there on my desktop. machine learning there on my desktop. machine learning third chapter 3.1 [01:26:29] Yeah. So there are two three methods whichever one you are comfortable with that's not an issue. If you clear you're comfortable Anush on VS code that's not an issue. Got it? Now everybody is there with me now. [01:26:44] So this is one of the very safest tool. You don't need to uh you know download the Anaconda on your desktop through the Google Collab through your good internet connection. You will be able to run the code. Yeah. So this particular file is [01:26:58] uh you know basically if I talk about 3.1 again I'm repeating go on to your learning management system on the reference material please download the reference material please download the ebooks the Jupyita notebooks [01:27:10] and then try to open. So if we talk about if we are back to 3.1 let's let's quickly go through the file that file is completely theoretical okay nothing great in that file it talks about what supervised learning is and what are the [01:27:25] two types of algorithms in supervised learning it's regression and classification what is the difference between regression and classification and supervised learning in regression the output is always [01:27:38] in regression the output is always numeric IC and classification output is always categorical. So when the target variable is categorical we do classification. So [01:27:50] this example we are very clear. So when it's predicting numerical value right that is regression. When we are predicting a categorical outcome determining the whether tomorrow it is going to be hot or cold. visualization. [01:28:03] going to be hot or cold. visualization. The same thermometer scale divided into The same thermometer scale divided into two colored regions cold and red for hot where the threshold separating is classification. [01:28:15] And if you look at the applications of supervised learning, do you think HR supervised learning, do you think HR people use machine learning to um uh identify the different job profiles uh shortlisting of the rums? HR is also [01:28:30] using machine learning now. Yeah. In finance, what is the use of machine learning and predicting of fraudulent data is are you know ready for that loan data is are you know ready for that loan approval or not. Right. And um this is [01:28:45] similar to how to how a credit card company determines your credit worthiness issuing before issuing a card. Then emails this is about spam filtering whether it's a spam or not. manufacturing. Anybody in manufacturing? [01:28:59] So in manufacturing, supervised learning is also used to inspect the quality, classify the products into different grades. For example, a factory might use a machine learning model to check for defects in products and ensure they meet [01:29:13] quality standards. Much like the quality control inspector, maritime industry, it helps in forecasting the historical events, weather condition. So [01:29:25] precautionary accident it's capable of predicting any kind of you know weather which might get wrong also which we have seen and supervised learning techniques like regression model can be used to predict tidal currents forecast [01:29:40] demand and supply reducing inventory losses think of it as how weather forecast predicts rain based on past weather patterns fraud predictions we use this in agriculture field too I'm into automating parking so under [01:29:55] supervised algorithm you know there are lot of algorithms that we will begin tomorrow again as you understand my style we'll start with regression then style we'll start with regression then linear regression logistic nave bias k&n [01:30:09] under classification they all come under classification so under linear regression we would be working on multiple linear regression and before we move on to classification there are lot of concepts that we will study such as [01:30:24] excuse Okay, excuse me. Cross validation, regularization, hyperparameter tuning, skarn pipelines. So there is lot more to be explored tomorrow. We will be concentrating on [01:30:40] linear regression and its concepts. Tomorrow, right? And linear regression is now very clear where the output is numerical. Then like predicting housing numerical. Then like predicting housing pricing, we will use regression. After [01:30:53] doing supervised uh regression in supervised learning then we will move on to classification algorithm. Under classification pelra of algorithms that need to be done initially we'll start with logistic regression kn bias kn [01:31:08] decision trees random forest and support vector machines. So it's a long journey. Tomorrow we'll do linear regression. Next weekend we'll we will be doing classification and then we will move on [01:31:21] to the next stages of ensemble learning. then it is uh unsupervised learning right so I think so you all are excited about this journey so it's a long journey I don't want to burden you today with all the [01:31:36] concepts so let's go slow and steady that's my rule slow and steady wins the race but I hope you got the crux of today's class right so this is 3.1 right and um we'll be beginning with uh 3.2 two. [01:31:51] >> We started with the very first session of machine learning where we understood different types of learning techniques, what are the different types of m machine learning techniques, supervised learning, unsupervised learning and [01:32:06] reinforcement learning. What is supervised learning? Supervised learning means it consists of label data which consists of input and output. We split calculate the errors. The two main algorithms under supervised learning are [01:32:21] regression and classification. Then we stood unsupervised learning. Then we stood unsupervised learning. Unsupervised learning consist of unsupervised learning consist of unlabelled data where based on the [01:32:34] similarity of the data the groups and clusters are formed. So two main algorithms are clustering and association. And last but not the least is reinforcement learning that based on the feedback positive and negative the [01:32:47] machine learns. Are we good to go? Everybody is clear with these concepts of machine learning the basic concepts. And then we started with the concepts of overfitting, underfitting and a good [01:33:02] fit. Yes learners now you will tell me what do we mean by underfitting fitting in machine learning very very important question it will definitely be [01:33:14] asked if you say that you know the concept of underfitting and overfitting. Yes learners can you tell me what do we mean by underfitting and overfitting? Underfitting is the case where the error [01:33:28] is high. What error is high? Bias is high. We are very far from the actual high. We are very far from the actual data point. But we talk about simple advantage of simple mathematical function? That it gives us more [01:33:42] interpretability of the model. What do we mean by more interpretability? That it gives the relationship between the output. And if we talk about overfitting again the error becomes high because the variance goes very high, right? And [01:33:59] definitely it makes and more complex. So we don't want a algorithm to be or model to be underfitted or overfitted. We want a good balance fit. That is how do we judge that that the training error and [01:34:13] judge that that the training error and the testing error should be similar. Right? So today we will start with the first uh you know learning technique that is supervised learning technique and the regression model. Are we good to [01:34:27] So as now you are getting more familiar with my style you know first we will try to cover the concepts and then move on to the practical aspect because unless and until you happening how is how are we going to [01:34:41] interpret the model there is no point moving on to the code but definitely we'll move on to the code but let's start with the concept of regression start with the concept of regression so here we go with the regression so [01:34:55] so here we go with the regression so what is regression now if I ask you what is regression? Tell me. So when we're talking about regression learners, please be clear with this concept that always the target variable or the [01:35:11] always the target variable or the dependent variable are all numeric values or continuous value. Don't say the term data. Right? See everything is data. But if you're not very spec [01:35:26] then technically you go very wrong. Right? So you have to understand now data we're talking about structured data which cons in tabular form it consists which cons in tabular form it consists of input as well as output values right [01:35:41] whatever are the input or the output value x over here is referring to your value x over here is referring to your independent variables right but whenever we are choosing any model or algorithm we always and always take into [01:35:56] consideration the depend dependent variable. So it is very important that the dependent variable has to be numerical or continuous right. So we are numerical or continuous right. So we are trying to find out so whole so basically [01:36:11] what are we trying to find out? We are trying to find out the relationship between the independent variable based on the output. Clear? So let's try to break the term linear and regression. linear. We understand [01:36:28] and regression. linear. We understand from simple mathematics that anything of degree one any function which has value one is set to be linear. So this means that progressing from one stage to another in a single series of steps [01:36:45] sequential extending along a straight line or nearly a straight line. So we all understand the term linear right and as very rightly said by nanda [01:36:58] regression is a statistical terminology which is used to find out the relationship between one dependent variable. We always have one output. If the output is categorical, we will go in for classification algorithm. And if the [01:37:15] outputs are continuous or numerical, we will go in for linear or not linear but regression algorithms. So moving ahead does a linear regression [01:37:29] means a linear approach for so when I combine the two terms linear regression it means that the linear approach for modeling the relationship between the modeling the relationship between the dependent and independent quantitative [01:37:44] variables. It will plot a straight line because that's the simplest linear uh approach as a best fit along the data points to predict the target value. [01:37:56] Right? So what are the different types of linear regression available to us? So a simple linear regression is with one depend obviously the dependent variable [01:38:09] is always one that's the target or the output but if the independent variable is one it is simple linear regression and it is always represented by a straight line. So beta kn is the intercept and this refers to the slope. [01:38:26] intercept and this refers to the slope. So this refers to beta kn plus beta 1 x1. getting my point? And then we have multiple linear And then we have multiple linear regression that is over here when [01:38:41] independent variable then this straight line turns into a hyper plane. So line in higher dimension two dimension becomes a plane and in higher dimension [01:38:53] becomes a plane and in higher dimension it becomes a hyper plane. Getting my point? Are you all getting the concept mathematical equation along with the graphical view and of course with the [01:39:06] code try to I'm trying to explain you from all aspects all right and when I from all aspects all right and when I say polomial sure I can do that right over here we are talking about linear regression linear means something of [01:39:23] degree one so do you see all the variables have all the coefficients have variables have all the coefficients have the degree E1. Do you see this? [01:39:35] regression. So the term linear regression means we are trying to find out relationship between the independent uh variables with the dependent uh variables with the dependent variable. So a simple linear regression [01:39:48] variable. So a simple linear regression is beta KN plus beta 1 X1. Clear? is beta KN plus beta 1 X1. Clear? That is when we have one input uh variable and one output variable. Mathematically it is represented by a [01:40:03] straight line where beta kn represents the intercept of the line and sorry the beta kn represents the intercept and beta [clears throat] 1 represents the beta [clears throat] 1 represents the slope of the line. Clear? [01:40:18] Have multiple linear regression. Graphically a straight line now becomes Graphically a straight line now becomes a hyper plane that is a multiple linear regression. Of course you know generally practically we will have more than one [01:40:31] independent variables with output variables. And then we have polomial linear regression that is y is equal to beta kn beta 1 x1 beta 1 x² and q clear [01:40:46] beta kn beta 1 x1 beta 1 x² and q clear now. So now let's try to understand it from now the same thing from the machine learning point of view. So in simple terms it is a finding a ba best straight line [01:41:02] fitting the given data set. So in simple terms find a simple straight line which tries to find the relationship between the independent and the dependent [01:41:14] the independent and the dependent variable and best linear relationship. How do I decide the term best? Something with minimum error or maximum accuracy. with minimum error or maximum accuracy. Right? Something with minimum error or [01:41:28] maximum accuracy is taken in terms of simple terms. And if I talk about technical terms, it is a supervised machine learning algorithm that finds machine learning algorithm that finds the best fit relationship on the given [01:41:43] data set between independent and dependent variable. almost the same thing but the basic concept or the algorithm it works on it is OS the [01:41:56] ordinary le squared method also known as the sum of the square residuals clear so to explain these terms graphically [01:42:08] mathematically I want everybody to concentrate here on the graph or the slid share so this is an equation of a straight so this is an equation of a straight line. What is beta kn? Beta kn is the [01:42:22] intercept. What do I mean by intercept? What is the value of y when x is = 0? Are you all getting this point? What does the term intercept mean? That what [01:42:35] does the term intercept mean? That what is the value of y when x is equal to is the value of y when x is equal to zero? Clear? Is this point getting clear to everyone? And if I talk about beta 1, beta 1 [01:42:48] And if I talk about beta 1, beta 1 refers to the slope of line. Is this term getting clear to everybody? Since I have only one input variable or independent variable, it represents a straight line. [01:43:03] straight line. Clear? of the straight line, what do these blue data point uh blue dots represent? the [01:43:15] actual data points available right that the first data point has value one y is equal to three this is value four and value is equal to 6 all right now this [01:43:29] is where how we are going to calculate the error what is the error the actual the error what is the error the actual value this you can consider this y i as value this you can consider this y i as my actual value and the red dot on the [01:43:43] straight line is the predicted value if I want to fit a straight line on these I want to fit a straight line on these data points. Right? So this is my actual data points. Right? So this is my actual output and this is my predicted output. [01:43:59] Clear? The error is clear. How are we calculating the error? Now what algorithm are we technically using over here? ordinary le square method or the [01:44:13] here? ordinary le square method or the residual. Residual means the error sum of square error. So what does this mean? It is equal to the sum of the square of It is equal to the sum of the square of error which is actual minus predicted [01:44:28] value. Getting my point? So and why are we taking the square of the error? There is a story behind that the error? There is a story behind that also. Why? Because [01:44:41] is +2 for example and this error is minus2. So the total error is zero. So what I'm trying to say is if the error is plus over here and minus and if I add [01:44:53] them it will show me zero error. It does it is not zero error. So I will try to take the square of the value it will become four + 4 that is the error is become four + 4 that is the error is equal to 8. So always the errors are [01:45:07] taken as square of errors. And we are looking at minimum square of errors. That is why it is known as the ordinary le square method. We want the [01:45:19] sum of the squared error to be the least. This is known as the RSS. Getting my point learners? So this is the total RSS. by the total [01:45:33] for this particular data point. This is going to be my first error. For this data point, my second error for the third one, fourth one, fifth one. And if they are infinite points, then infinite data points are calculated. [01:45:48] data points are calculated. Clear to everybody? So why y Okay. Now Clear to everybody? So why y Okay. Now because this is my actual y1 and what is my predicted y1? This is my predicted value right base because we are using [01:46:04] nanda this particular function. This is my predicted value and predicted values are represented as ycap. So for the first one it will become beta minus beta first one it will become beta minus beta 1 x1 then y2 beta beta 1 x2 for each [01:46:18] 1 x1 then y2 beta beta 1 x2 for each data points. Now clear this is how based on this mathematical function I am going to predict my output for each data points. Clear? So the formal statement of simple linear [01:46:35] regression I'm talking about simple linear regression with one input and one linear regression with one input and one output. Y1 is the value of the response variable. Again, predict, target, predicted, you know, dependent, they all [01:46:51] mean the same thing. Do not get confused with the terminologies in the IAT trial. Beta not and beta 1 are the parameters. Xi is the value of the predictor variable in the IAT trial and epsylent I is the random error with mean E. epsylon [01:47:08] i is equal to zero and variance is equal to square. Clear? Okay. Now this random error this is very very important to understood that a [01:47:20] very important to understood that a random error we are talking about are we talking about the whole data set or a part of data set or a sample of data set part of data set or a sample of data set over here. This is my population and if [01:47:34] I take a part of it for training and as well as testing do I always consider the whole data set or a part of data set it's always sample and when I and is my [01:47:46] sample and we want the sample to be the true representative of the population true representative of the population but since it is a part of the data there will be some error also uh you know associated with it which is irreducible [01:48:01] that is represent ed by epsylon right so epsylent I is the random error with mean this is the random error which is always going to be introduced it's [01:48:13] it's not reducible like bias and variance or overfitting but it's part of the whole linear regression where the mean of that error is equal to zero and [01:48:25] mean of that error is equal to zero and the variance is equal to epsylon square the variance is equal to epsylon square now clear so The idea of whole uh linear now clear so The idea of whole uh linear regression or significance model is that [01:48:37] regression or significance model is that it is highly interpretable. it is highly interpretable. What is the idea behind this? That it is why the error is zero. We always want the error to be zero. In this case we [01:48:50] are considering that uh you know the average of the error not the square the the sum of the square or the average of the error is equal to zero not the Okay. So now the significance of linear [01:49:07] regression lies in the fact that we can easily interpret and understand the easily interpret and understand the marginal changes with with with input to marginal changes with with with input to the output. So linear regression is an [01:49:21] the output. So linear regression is an highly interpretable model. How is it interpretable? That if we increase the value of X1 by one unit keeping the other variables constant, then the total increase in the value of Y will be beta [01:49:37] 1 and the intercept term beta KN is the response when all the predictor terms response when all the predictor terms are set to zero and are not considered. are set to zero and are not considered. Right. So how much is my y value change [01:49:51] Right. So how much is my y value change will depend on how much is my beta 1 is it positive or negative beta_2 and beta n. So now what are we trying to find out in linear regression? The values of beta kn beta 1 [01:50:06] regression? The values of beta kn beta 1 and beta n right where the error is minimum. Clear? [01:50:20] lot of concepts lot of mathematical terms coming. So just hold on till we start moving on to the practical part and you know you will see them practically just try to grasp as much concepts as you can. So the different [01:50:33] assumptions which we follow in linear regression is linearity that is the relationship between the features and the target homoidastiticity. The error term has this is where you know the [01:50:48] linear regression has its assumption or constraint and the the term is known as homoidasticity that the error term is constant variance throughout the whole data. The error term will be constant. Multi-olinearity, [01:51:04] there is no multi-olinearity between the features. And if we talk about independence, observations are independent of each observations are independent of each other. getting my point? Normality. This [01:51:20] is where the uh point I was saying that the error residuals follow normal standard distribution where mean is equal to zero and the standard deviation is equal to one. Okay. So multicolinearity [01:51:36] means that we can have more than one input. Are they related with each other? input. Are they related with each other? No. every in in uh first input should be one should be independent with the third. The input should not be uh [01:51:52] dependent on each other. They should all be directly dependent on the output. First let's try to understand the concept. So how do we go about in the supervised uh learning process that we [01:52:07] supervised uh learning process that we have this full data set right and we will after doing all the pre-processing so today I told you to revise all your EDA that we will load the data set using PD dot read CSV command do head and tail [01:52:23] PD dot read CSV command do head and tail try to find out null values work on them uh do uh encoding if required and then split the data into training and testing. First step is getting clear and of course training can be se it's [01:52:38] generally 70 it's generally 7 to 2 to 80. If you want to keep it 50/50 also nobody is stopping you or 6040 that's completely up to you. Okay. And the function which is used in Python in the sklearn prep-processing library [01:52:54] in the sklearn prep-processing library is train test_plit. X is the first parameter which represents the features of the input data. So this is how we will create a list or array which will contain all the [01:53:07] list or array which will contain all the input. Y refers to the target or the label vector of the input data. Size of the test data test underscore size. And the test data test underscore size. And finally we have random underscore state. [01:53:21] What is this random state? It is the seed that initializes the pseudo random seed that initializes the pseudo random number generator. What does that mean? That I will give any fixed number. It could be 42. It could be three. It could [01:53:36] be zero. It could be 500. Right? I want to keep this number constant so that the shuffling or the randomization of the data happens in the same way. so that I can check the accuracy and do the comparisons. [01:53:52] Clear? Right? So again remember this point whenever data is given to you you have to identify it in terms or split it in to identify it in terms or split it in terms first in terms in number of inputs [01:54:07] and the output always separate the outputs from the input and then we will use this function and then we will get four outputs. train test_split [01:54:25] gives four output that is x train y train x test and y test clear training test_split is used to split the data right we are [01:54:37] preparing the data which will be used for training and testing separately so the parameters of this function are input output how much we want to give for testing. It could be 3/4 half completely up to you. No rules for that. [01:54:54] completely up to you. No rules for that. And random state gives me the seed that initializes or does the shuffling of the data in the same manner. And the output of this function is four output. I get input training and output training data [01:55:10] and input training and output testing data. Clear? data. Clear? And now if I talk about multiple linear regression right in threedimensional setting with [01:55:25] two predictors one response the le squared regression line becomes a plane. So this is a multiple linear regression model epsylon is that normalized error. [01:55:37] Okay which is irreducible. Nothing can be done much about it. And now we will be done much about it. And now we will practically deal with data which has practically deal with data which has three inputs. The budget of TV, [01:55:50] three inputs. The budget of TV, budget of radio, budget of newspaper and which one of them affects the sales most. So this is the first [01:56:02] uh you know algorithm that we are going to study. Okay. So this is simple. This is polomial and these are the errors. So let's start with the practical work and then get back to the different errors and analysis. [01:56:17] Got it? Have you got an idea what linear regression is all about? So I request you all learners to from the LMS2 open 3.2 [01:56:29] supervised learning regression linear regression five. Okay. So now do you understand the term independent independent variables straight line when we have in two dimensional when we have one input one output it's a straight [01:56:45] line blue dots are my data points and this is a straight line or the line of regression which I want to fit on the data points this graph is now very clear to everybody I'm starting with the uh pile now the [01:56:59] first 3.2 into regression. This graph is clear. And where do you think regression is used? It has a huge huge application in oil and gas uh industry. Various types of data collected in oil and gas [01:57:14] industry from surface subsurface to understand production sale processes. liquid? How much is the pressure that needs to be created? So linear nonlinear regression model forecast global oil production. Oh my god, the whole world [01:57:31] production. Oh my god, the whole world is fighting on oil and uh well we can is fighting on oil and uh well we can use our aggression models to forecast the global oil production. Right? So the whole world is on a big fight on oil [01:57:43] whole world is on a big fight on oil only. Right? regression helps to analyze the effectiveness of advertising campaigns. Predict sales based on marketing spend. segment customer based on demographic [01:57:58] data. Right? Then we have retail. Then we have linear regression is utilized in retail for demand forecasting, inventory management, pricing optimization and customer analytics. Health care. Linear [01:58:13] regression is applied in linear healthcare for predicting the patient outcomes, analyzing relationship between the different medical variables and the different medical variables and their diseases. getting my point? [01:58:27] And finally, we have the real estate that in the real estate industry, linear regression predicts the property prices based on the factors such as location, based on the factors such as location, size, mment, uh amenities and economic [01:58:41] indicators. Clear? So now what are the different types of regression? Regression can be classified into two categories. linear and [01:58:53] nonlinear. Linear regression finds a straight line relationship between the dependent variable and one or more independent variable. Nonlinear regression uh finds a relationship between the dependent variable and [01:59:09] between the dependent variable and independent variable using a curve or more complex shape. Okay. So getting back to linear regression it is so technically if somebody ask you what is linear [01:59:22] regression in in machine learning you will say it's a supervised learning will say it's a supervised learning algorithm which is used to please try to understand each and every term is used to predict a continuous target variable [01:59:35] please be clear it's not continuous data as you all were saying that's a wrong statement it helps us to predict continuous output variable by modeling its relationship with one or more independent variables through a linear [01:59:52] regression. It predicts a continuous dependent variable based on one or more independent variables. So it predicts the continuous dependent variable. Here it says target output dependent [02:00:06] completely anything you can say. It uses the le square criteria to estimate the coefficients of the regression equation. What do we mean by the le square criteria? That the value between the actual and the predicted value. The [02:00:23] actual and the predicted value. The square of the error of this error that is RSS should be the should be minimum that is the residual sum of squares. Getting my point and it can be applied if there is linear [02:00:39] relationship between the variables. So in case the dependent variable is continuous and independent variables can be continuous or decre discrete we are not bothered about our input variables. So for predicting the uh you know output [02:00:57] we are always considered about the output. So the relationship between a dependent variable Y and one more independent variable X is established using a best fit straight line which is also known as the regression line. [02:01:14] And there are two types of linear regression. One is known as the simple linear regression. Other one is known as the multiple linear regression. All right. So simple linear regression is clear to [02:01:29] everybody. We have one independent variable and one dependent variable. Do we understand now the terms beta KN and beta 1? This is the intercept and the [02:01:42] slope. Yes, everybody is understanding now what is simple linear regression? And as I told you a simple linear relationship is given by one of the input that is TV expenses over here as [02:02:00] my input and in the output we have that sales. That means in this particular data set we are trying to establish the [02:02:13] relationship that if what is what is going to be the budget of TV and how it is affecting my sales. So what do you conclude? Do you see the red points? These are the predicted values and the blue points are nothing but my actual [02:02:29] data points and I'm going to calculate or it automatically does the calculation or it automatically does the calculation of the error if once I you know apply the algorithm clear. So do you see a kind of positive [02:02:43] linear relationship? Can I conclude that this shows a little Can I conclude that this shows a little positive linear relationship? that is the value of TV [02:02:56] expenses increase the sales is also increasing. It shows a positive relationship between them. Getting my point and practically if you look at the real world problems they will not be simple [02:03:10] linear regression problem. They will going to be uh you know multiple linear regression problem. There will be more than one input from how we know positive and negative. See positive and negative button that's that's you know shows it [02:03:23] very clearly from the graph. If the graph is showing this thing that as my x is increasing y is also increasing it shows a positive y is also increasing it shows a positive relationship. See over here [02:03:43] decreasing. But if my graph is like this, what does it show? That x increases, y decreases. y decreases. And if x decreases, y increases. Getting [02:03:57] my point? So this is a negative relationship. Inverse relationship between input and output shows negative relationship. and [02:04:09] uh direct relationship between input and output shows positive relationship. So as I was telling you practically if I talk about uh you know real time data uh [02:04:21] you know we have more than one inputs and output. So multiple linear regression models the relationship between the two and more independent variables predictors features and the dependent [02:04:36] variable as a straight line. The equation for multiple linear regression equation for multiple linear regression is y is equal to beta kn plus beta 1 x1 is y is equal to beta kn plus beta 1 x1 beta 2 x2 and beta n xn. that x1, x2, xn [02:04:50] are the predictor variables and beta 1, beta 2, beta n are the coefficients of each predictor. So let's start practically working on this data set that this becomes now if I want to try [02:05:06] to find out the relationship between TV expenses, radio expenses and sales. So now my straight line turns into this hyper plane. Do you see the difference how the straight line has got converted into hyper plate? And of course there is [02:05:22] one constraint that we cannot view more than 2D or 3D graphs right we cannot than 2D or 3D graphs right we cannot view more than 2D or 3D graphs right view more than 2D or 3D graphs right that correlation covariance is same as [02:05:36] uh you know um regression but in regression it's a linear relationship but the positive and negative relationship is the concept same as correlation and covariance that correlation and covariance is founded by [02:05:51] Pearson's you know coefficient of correlation but here we are talking in terms of linear equation but if I talk about one input equation but if I talk about one input x1 and y how x1 is related to y then [02:06:04] x1 and y how x1 is related to y then that is r that is this + one minus one that's that exactly is the same concept the concept is not very different linear regression is using that particular concept all right so now let's start [02:06:19] practice practically working on the data set. So what I want everybody to do is on your uh you know go from data set folder open your Jupyter notebooks and [02:06:32] folder open your Jupyter notebooks and keep the TV marketing CSV data in the same folder where you have kept this 3.2 yes learners wherever you have kept this yes learners wherever you have kept this 3.2 to keep the TV marketing [02:06:45] 3.2 to keep the TV marketing dot CSV copy it from the data set folder dot CSV copy it from the data set folder and start uploading this data and start uploading this data clear [02:06:59] and then run this thing. So basically this data set consist of three input variables. This is my x1, x2, x2, x3, and this is my output. [02:07:14] All right. So once I have done my head, what does the head command do? What does the head command do learners? What does it prints? The top five rows by default. And these are my inputs, right? And what does the info function [02:07:31] do? Good, Rishi. Good Anush. What does the info do? It tells me there are no the info do? It tells me there are no null values in this data set. And of course unnamed is not required over here. And all other my inputs are also [02:07:46] in numerical value. And what is my output? My output is sales. And that's a numerical value. And that is why regression will be used. Clear? [02:08:01] regression will be used. Clear? And that is why regression will be used. And that is why regression will be used. Getting my point. [02:08:13] are going to do? Since this data does not require pre-processing, it's a clean data. So nothing required to clean or check on the null values etc. Right? So [02:08:25] we are considering the data to be clean. So now when I want to fit any model what is the first step that I will do? I will try to separate the input to the with try to separate the input to the with the output columns. Agreed? The first [02:08:40] step is to extract the features of the data set and output. What is this ILC in Python core? Index base filtering. So if I see very [02:08:54] Index base filtering. So if I see very closely that 1 2 3 are my input index. closely that 1 2 3 are my input index. This is of actually no relevance. This is of no relevance. These are my input values [02:09:20] and this is my output value. Clear? So in this case ilocc the first colon represents it. It talks about all the rows but only the columns which are there. It says one and two. So how many [02:09:34] inputs are we talking about? Only the first one. Only the first one is TV. And first one. Only the first one is TV. And what is the output? Only the sales. This point is getting clear. What are we taking over here? We are talking about [02:09:49] TV and sales. So we are trying to implement linear regression, right? We are trying to do simple linear regression. This X and Y is clear. I've only taken a part of it. First I've taken one input and then I'll take all [02:10:04] taken one input and then I'll take all the three inputs. This point is clear, the three inputs. This point is clear, right? And then we have the from skarn right? And then we have the from skarn model selection import train test_plit. [02:10:17] And now I am taking my input that is only the TV input. Y is my sales test only the TV input. Y is my sales test size. 30% of the data is used for size. 30% of the data is used for testing and random state is taken to be [02:10:32] 42. 42 is generally considered from a fiction that it's cons considered to be a you know a good number for the whole universe. Clear? [02:10:44] Okay. So now what are the different steps to now fit the model not much of coding that's why Python is a preferred language beautiful language that is from [02:10:56] sklearn.linear model import linear regression that is first we will import the library that is skarn is the main library and we are [02:11:09] importing linear regression from there. So when I imply linear regression it is lin_re and this is fitting that is the training and this is fitting that is the training input and output training data. [02:11:25] input and output training data. Getting my point fit a function is used Getting my point fit a function is used to fit the data or train the data. What do we mean by train the data? finding the parameters over here that is beta, [02:11:41] the parameters over here that is beta, beta 1, beta 2, beta 3 etc. Clear? Are you all there with me till this code? Anybody who still facing any [02:11:54] difficulty who's not been able to run this part of the code this part of the code till here are we good to go? First step is to import the library which contains the linear regression [02:12:07] which contains the linear regression model and then I will create an instance of this linear regression function right that is lin rig it is a instance what do [02:12:19] you understand by instance object of this particular function and then fit this particular function and then fit the training input and output data the training input and output data it is divided by using this train [02:12:34] it is divided by using this train test_split input training and testing output training and testing I explained you [02:12:46] train test split where we get the four outputs relationship between the sales and the TV so TV and sales so we generally we [02:13:02] TV so TV and sales so we generally we want it to be an input. regression model x label we don't want it to be that's I think so that's a [02:13:16] it to be that's I think so that's a little opposite of it over here this is TV change make changes in the code and TV change make changes in the code and output is sales [02:13:35] constraining factors and model. What do we mean by that? Delineate. What do we we mean by that? Delineate. What do we mean by that? represent? What does this graph represent? This plot shows a positive [02:13:49] linear relationship between sales and TV. And where the blue regression line indicates the model's prediction, the green data points are generally close to the line suggesting the model fits the data reasonably well through those some [02:14:05] data reasonably well through those some variability exist. Right? So does it show a positive relationship basant? Right? Do we see that there is a positive relationship between TV and [02:14:19] sales? No, it's not the train data. It is the test data. Why test data? Because I'm plotting input as training and predicting the output. The xaxis is the input. The x is input, [02:14:37] right? And the output is prediction of the training. Right? So, we are trying the training. Right? So, we are trying to plot over here. Plot over here is the to plot over here. Plot over here is the straight line. [02:14:53] Right? But why am I saying the test data? Because actual data points are my data? Because actual data points are my input and output test data. Now clear, we are trying to plot uh do a scatter plot over here. Scatter plot is between [02:15:08] input and output test data. The actual data points are by testing data points. Okay? And the plot the straight line is input [02:15:20] training and predicting the output. That is why it say it says the linear is why it say it says the linear regression model for test data set. Now clear. So uh what what do we conclude from this [02:15:35] graph? We can conclude that the plot shows a positive linear relationship between sales and TV where the blue regression line indicates the model's predictions. The green data points are generally close to the line suggesting [02:15:49] the model fits the data reasonably well though some variability exists. Okay. So this shows a positive relationship. Now the important point is [02:16:02] how do I understand is this model underfitting or overfitting. So when we are developing machine model what is the most important point important concept that needs to be taken into account? The concept of overfitting [02:16:18] and underfitting. So when developing machine learning models achieving the right balance between the complexity and simplicity is crucial. What are the points of overfitting? Let's quickly go through that. Overfitting occurs when [02:16:33] the model learns the noise and details in the training data too well to the extent that it is negatively impacts the performance on unseen data. [02:16:45] impacts the performance on unseen data. Right? That means when the variance is high sign high accuracy on the training data but poor accuracy on the test data. [02:16:58] So how do I know that my water has overfitted? That I get a good score for overfitted? That I get a good score for training data but a bad score for testing data. And why is it? What is the reason? [02:17:13] Because the model is too complex. It has got too many parameters. Underfitting happens when the model is too simple to capture the underlying pattern. Poor accuracy on both training and testing data. Underfitting is more easily to [02:17:28] detect but overfitting is a constraining problem. So what are we looking? How do we solve this problem? We want a bias we solve this problem? We want a bias and variance tradeoff. Bias error due to [02:17:41] overly simplistic assumptions in the learning algorithm. And high bias always causes underfitting. Variance is error due to excessive complexity in the learning algorithm. High variance causes overfitting. Clear? So low bias, high [02:17:58] overfitting. Clear? So low bias, high variance again it is failing uh to be a good fit. This refers to overfitting. High bias low variance it leads to underfitting and optimal tradeoff finding a balance between the model [02:18:11] performance on both training and test data. So this is what we are actually data. So this is what we are actually trying to achieve in the output. Right? So just to understand things you have to understand the machine learning model [02:18:26] from the mathematical intuitive behind it. If it is possible to understand it higher complexity higher dimensional plane it becomes difficult to access uh to understand it visually but mathematical intuition the concept and [02:18:42] mathematical intuition the concept and how do I interpret the results are you understanding and of course how we are implementing those functions in Python that is where you know your concentration should be mathematical [02:18:57] concept or or what is the concept of that particular algorithm. How are we implementing in Python and how am I interpreting the results? That how am I interpreting the results? That should be the option. Clear? [02:19:11] Is this point getting clear to everybody? And how do I check the error? The checking of the error is going to be through the training and testing error. So if you just look at this, I I'll tell [02:19:27] you what these errors are. So if you just look at the training error is more than the testing error. So is it underfitting or overfitting? Definitely [02:19:39] underfitting or overfitting? Definitely underfitting. Consider increasing model complexity. That means linear regression or a straight line fit is not a good fit or a straight line fit is not a good fit on this particular datas. Agreed? [02:19:54] So now getting back to the concepts of errors, how do I evaluate my regression errors, how do I evaluate my regression model? [02:20:09] Okay, so one of the simplest error is known as the mean square error. This is known as the mean square error. This is nothing but your RSS, the residual sum [02:20:21] of square error. Please try to understand this is nothing. So errors are the do you think one error is easy to define linear regression? No. Right? There would be several errors. So now let's try to understand different errors [02:20:37] let's try to understand different errors available in the linear regression model. Yeah. So if I talk about the mean square [02:20:49] error, this is nothing but my RSS or the residual actual minus predicted value the whole square. So error square of the error and if I take the average of it, error and if I take the average of it, it becomes mean square error. Agreed? [02:21:04] it becomes mean square error. Agreed? Are you understanding this error? And when I talk about root mean square error that is when I take the under under root square root of this particular MSE. But what is the need of [02:21:21] taking the root mean? Can anybody tell me we what is the what is the need of taking the root of this mean square error? Anybody who can explain this is exactly the same uh you know [02:21:35] question which relates to standard deviation and variance. Anybody who understands the relationship between standard deviation and variance [02:21:53] after square error will increase or square root will get proper value. No nanda that's wrong answer. Yes ma'am. rooting variance is standard deviation like but why do we take that under root that's the question [02:22:12] variance what is why what is the significance no it doesn't mean that significance no it doesn't mean that squaring will uh get a proper value no square root will yeah because ultimately we have to [02:22:28] compare with other values such as such as average. So average is a simple as average. So average is a simple value. Simple values are cannot be compared with a squared value. Mean square value cannot be used for [02:22:41] comparison with other values. So to make it at you know can I it's something like can I compare a centime squared the unit of area that is equal to cm squared with [02:22:53] cm can I compare a cm squared or a meter squared a unit of area with normal meter? No that goes mathematically wrong. So just by taking the square root of MSE now it becomes [02:23:11] at at the same level and then the comparison can be done. Yes. And if we talk about mean absolute error that is actual minus the predicted value and actual minus the predicted value and this is I divided by N I get the mean [02:23:25] this is I divided by N I get the mean absolute error. Clear? Now these are all absolute error. Clear? Now these are all you know um you can say absolute terms of error and if I want to do comparison between different models now what does [02:23:39] it what difference does it make if I'm saying actual minus the predicted the whole square and then I say predicted minus actual the whole square is there any difference between that [02:23:58] mathematically whatever way they write you know the answer will be same. Yeah. Clear? Now now let's start let's understand another factor of comparison that is R². Please try to understand R². [02:24:12] Okay. Okay. Now to understand R square it is a relationship between S STTO. What is this total sum of square? This these are [02:24:24] my actual points and total sum of square is the difference or the total variation between the actual points and the average. Generally we want to always find out how far my points are from the average. [02:24:41] Right? How my points are far from the average. Right? So SSTO is getting clear to everybody. everybody. So over here this is my actual [02:24:56] data point and here I'm doing my comparison. Is the total sum of square clear to everybody? What was what is SSC? Is SSC same as the RSS? Yes. Here we are trying to find out the [02:25:13] difference between actual value and its difference with the average yi minus difference with the average yi minus ycap v² again y square because I don't want any no positive and negative values cancelling each other clear [02:25:29] and if I talk about SSC or the error sum of square this is like RSS actual minus the predicted value this is my actual this is the predicted value. Do you see this? So the uncertainty of the data around Y [02:25:45] observations lie around the regression line. If SSC all observations fall on regression line, larger the SSC, the greater is variation of the Y greater is variation of the Y observation around the regression line. [02:25:59] Then we have regression sum of the squares. Right? This is actual uh sorry the average predicted value with the difference with the average. So the [02:26:11] difference between the predicted or the fitted value on the regression line and the mean of the fitted value that is equal to SSR. The measure of variability [02:26:23] of Y associated with the regression line. larger the SSR in the relation to SSTTO greater the effect. Right? So if we look at now the whole [02:26:36] Right? So if we look at now the whole picture that my total sum of the square picture that my total sum of the square is equal to SSC plus SSR that is the difference the sum of square values with the actual value minus the [02:26:50] average the whole square is equal to actual minus the predicted the whole actual minus the predicted the whole square and predicted minus the average. Agreed learners? But what is the use of all these? So now there is this another [02:27:05] factor another uh unit which helps in comparison of value that is R² that is known as coefficient of determination [02:27:20] determination which is also known as the R square statistics proportion of variance explained the value varies between zero and one and independent of between zero and one and independent of the scale. scale y right so r² is equal [02:27:35] the scale. scale y right so r² is equal to sstto minus s upon sstto ssr upon sstto so what is the advantage in short let's not get into the mathematics that let's not get into the mathematics that the r² value will always lie between 0 [02:27:51] to 1 so the values which are more close to one are better fit as compared to to one are better fit as compared to zero yeah mean square error and RMSSE relationship is clear but they are not relative errors. I cannot do comparison [02:28:08] between the different models that this model is better than the other. They are model is better than the other. They are good to tell the errors are less or more comparison between training but R square is a relative error. Even absolute error [02:28:22] is not a relative error. What do I mean by that? What do I mean by that? Let let me explain you with an example. For example, you scored 70 marks and I have scored 30 marks. Which who scored better? You would say 70 marks. Ma'am, [02:28:38] you you scored better than me. But I scored 30 out of 30. You scored 70 out of 200. Getting my point? So there how will I find the relative? If I start taking percentage of this that 30 out of 30 is [02:28:54] 100% and 70 out of 200 is somewhere around 40 50%. then I can do comparisons. Got my point? That is why we are now heading towards ratio between [02:29:07] we are now heading towards ratio between the SSTO the total sum of squares with RSS and SSR taking into all account all the variability of the data points and that's why I now reach on this R squared [02:29:22] term now better is it better sum of square error ranges between 0 to 1 which measures the amount of variability that is left unexplained after performing the regression. An SST minus S SSE measures [02:29:37] the amount of variability that is explained or removed after performing the regression. R²AR and the proportion of variability in Y can be explained of variability in Y can be explained using X. So R² is trying to take into [02:29:52] account all the errors and variability of the data point. So when R² is equal to 1, SSC is zero. That means it's a perfect fit. It's an ideal situation. [02:30:05] Will it ever happen? Hardly. I don't think so. You know it's going to be but something which is close to 1.99 is a good and when I get this horizontal line that you all were getting that means r² is equal to zero error is equal to total [02:30:21] that means there is no linear relationship between x and y. Now clear relationship between x and y. Now clear right but R² suffers from one one [02:30:35] right but R² suffers from one one limitation right since R² is also known as the if it is greater than one it can never be greater than one that's the assignment that's not possible mathemat can percentage can ever come beyond 100% [02:30:50] the formula is such no because we're taking a ratio out of 100 right Simon So mathematically it can never be uh true. Got it? [02:31:02] Okay. So R square also known as the coefficient of determination measures the proportion of the variance variation in your dependent variable X and in your dependent variable X and explained by your independent variable [02:31:17] X for the linear regression model. But the problem that R² also suffers is that it will always remain or increase as we are adding more number of independent are adding more number of independent variable. So this problem is solved by [02:31:31] variable. So this problem is solved by adjusted R squared. What is adjusted R That it measures the proportion of the variation explained only those variation explained only those independent variable that really help. [02:31:45] So what do we do? We divide the error by degree of freedom. Do you understand the degree of freedom. Do you understand the term degree of freedom? Learners [02:31:59] Degree of freedom is the number of independent the number of independent variables. Right? Suppose if there are 10 10 if suppose if there are 10 input features then the degree of freedom is n [02:32:15] -1 that is 9 clear how many input features are we depending on abinatri that is referring to my [02:32:29] degree of freedom it's something like now let's understand it over here now let's understand it over here now if If I want to predict the price of the car, what are the factors that are determining the price of the car? Safety [02:32:43] is a very very important factor. How many airbags does it have? Does it have many airbags does it have? Does it have a parking sensor or not? The branding. What is the most important factor? The mile, the fuelage, consumption, the fuel [02:32:55] type in today's world, assistant, infotainment system, is it a luxury car infotainment system, is it a luxury car or not? the color, the kind of uh you know the design, the seats, the the the the [02:33:09] wheel, the alloy. There could be several factors. Agreed learners. But do you think the the kind of wheels or the wheel alloys the kind of wheels or the wheel alloys are as important as the engine of the [02:33:23] car or the model of the car to determine the price? the price? No. Right. So that is what is taken care by adjusted R square and when we divide it by the [02:33:38] number of degree of freedom getting my point see normally what will happen is if I keep on adding the factors in R square that I've added the number of [02:33:50] alloys the number of uh you know the seat covers the color of the car the the R square will keep on improving it will give me an illusion Oh, my model is giving so much well answers but that's not true. Maybe you know adding the kind [02:34:05] not true. Maybe you know adding the kind of wheel alloy is not that concrete uh factor to determine the price of the car. See the these numbers of degree of freedom as I told you this is some [02:34:19] example that they are taken. Degree of freedom depends on the number of input features. So that is why adjusted R square is a So that is why adjusted R square is a better better uh you know way of [02:34:33] better better uh you know way of determining the results. determining the results. So if I get back to the file over here and now look at how I calculate the results. So what am I trying to do from [02:34:50] sklearn metrics import mean square error R square error and then I do the prediction. Prediction is done by the predict function for the training output [02:35:03] predict function for the training output as well as for the testing output. And then the different metrics can be simply calculated in Python by using the mean square error for training data for testing data and R square. And then I [02:35:19] testing data and R square. And then I can do the comparisons. So if you do you see the training error is 0.573 and testing error is more than that that's more closer. So they all refer that the model is underfitting. [02:35:37] refer that the model is underfitting. Now clear how are we trying to fit the model? How are we trying to uh [02:35:52] calculate the error. So if we talk about linear regression let's let's look at it again that this is my actual output right and if I talk [02:36:05] about a single variable as my input and there are two coefficients beta KN and there are two coefficients beta KN and beta 1 this is the epsylon error is this first equation getting clear to everybody [02:36:24] is getting clear to everybody Everybody then we have linear regression multiple then we have linear regression multiple variables right we have x1 x2 x3 right absilent over here refers to the random [02:36:38] error and if we talk about a model evaluation are we clear with these metrics now the mean absolute error the mean square error and the root mean square error right so comparisons between the [02:36:55] different models is not possible. Therefore, we have R square error which talks about the ratio and even better than R square is adjusted R² because it is not affected by the number of inputs [02:37:11] is not affected by the number of inputs in the data. Clear? So basically we in the data. Clear? So basically we divided by degree of freedom mean model. divided by degree of freedom mean model. Clear? And we have also understood [02:37:24] Clear? And we have also understood different psych learn objects. Please try to understand fit function. Fit function helps us to train the input [02:37:36] parameters or train the parameters for the data. Transform is transforming the the data. Transform is transforming the input to output. Fit transform is mixing the function of fit as well as transform. And it is the predict [02:37:51] function which is used to predict the output of the training data or of the testing data. To understand it better, types of skarn To understand it better, types of skarn objects are based on transformers. [02:38:06] Transformers are nothing which transform the data set. Have you learned feature the data set. Have you learned feature engineering learners? Do you understand feature engineering where we do feature scaling, [02:38:19] standardization and normalization, encoding of the data that we need to encoding of the data that we need to transform the data before we actually fit into the model. Let me try explaining you over here. [02:38:35] So before the data is actually fed into the model, it needs to be transformed. Why it needs to be transformed? Because if they are missing values that will not fit into the model. If they are categorical features then we have to [02:38:51] categorical features then we have to involve uh encoding. Feature scaling is a feature which brings the data into one particular range and outlier detection. That is why data science is important because these concepts are covered in [02:39:08] detail. How do you deal with data? How do you handle missing values? How do you deal with categorical data? Scaling of the data is all part of the feature engineering process. So in feature engineering, we are not [02:39:24] So in feature engineering, we are not just building models but we are playing match makers for data as well as algorithms. Getting my point everybody? Okay. So [02:39:36] getting back for transforming the data set right. So fit learns what does a fit function do? It is an estimator. Please try to understand. Fit is an estimator [02:39:48] which estimates the model parameter based on the training data and hyperparameters. And finally we have the predictors that is it predicts the data makes the data set as input and does the prediction [02:40:03] score method to measure the quality of the predictions. Getting my point? So here we have the data pre pre-processing training and inference [02:40:15] pre-processing training and inference model clear and then we have fit and fit model clear and then we have fit and fit transform right what do we do that when we are trying to train the data what I'm trying to explain let me explain it with [02:40:28] trying to explain let me explain it with this slide that again please try to this slide that again please try to understand let me rewind things in supervised learning we have label data first First of all, what is label data? [02:40:41] That we have the input as well as the output. Now, in the story, have you understood two types of output in supervised learning? That we will have the actual output of the label data and one we will predict the data using the X [02:41:01] test. This point is getting clear. what this point which I'm trying to explain regarding the output is getting clear right so this is output actual as well right so this is output actual as well as predicted now we also understand that [02:41:16] we divide the training and the testing data right so ultimately we are trying to when we are doing training it is all on the training data by the fit parameter so the model all the training will be [02:41:31] done using the fit function in simple terms, right? And the testing and output will be on the predicted. But if I want to transform the data, so I [02:41:44] will fit it and then transform it. And can I do the function of fit and can I do the function of fit and transform together? Yes, by using fit transform together? Yes, by using fit transform function. Getting my point. [02:41:58] transform function. Getting my point. Can I do the function together? Yes. For Can I do the function together? Yes. For only training data I can use fit transform and for testing data I will only use the transform function. Now [02:42:13] only use the transform function. Now clear code. So now are we able to also understand [02:42:25] how when the model is underfitting and overfitting that point is also clear. And if we talk about nonlinear regression, polomial regression, uh polomial regression is a subset of [02:42:38] linear regression that includes polomial terms. The relationship between the independent variable x and dependent variable y is modeled as an nth degree. So here we are not trying to do it as a straight line but polomial regression [02:42:55] which is a subset of linear regression which includes polomial terms. The relationship between an independent variable X and dependent variable Y is variable X and dependent variable Y is modeled as nth degree polomial. [02:43:14] called a special case of multiple linear regression? Some polomial terms are added to multiple linear regression equation to convert it into polomial regression. So it is a linear model with some modifications made to increase its [02:43:29] accuracy. The data set used in polomial regression for training is nonlinear. So what do we what do we observe that in [02:43:41] this particular graph this is a straight linear regression and if I change it to a curve it becomes polomial linear equation. Do we see that exponential curve over here? Okay. [02:43:55] Okay. So I have this uh you know simple data So I have this uh you know simple data right in this simple data this is the project statement. Okay. Wait. [02:44:14] Yeah. So a certain spare part manufactured company once a month in lots which vary in sizes. data on lot and size numbers and man of hours that that means the input is the lot size of the people and how many man hours are [02:44:28] the people and how many man hours are required right it's numerical right so if I start analyzing the data actually you know it's done through an excel it's you know it's done through an excel it's simple that I do x - xar yi y - y bar [02:44:42] and I calculate this what is this known as rss residual sum of squared actual minus minus the output the whole square x i minus x i the whole square do [02:44:54] you see this and why am I taking the square value because I simply take the difference between y minus ycap the error is zero which is actually not which is actually not therefore we take the square values [02:45:08] not therefore we take the square values and I get the answer as 13 six okay so the sample size the number of rows in this data set is 10 degree of freedom over here is 9. Mean is equal to xar [02:45:25] sigma xi upon n 50. Variance over here is this much. Standard deviation is in is coming into picture. We have calculated the value. Similarly we can [02:45:38] calculate it for the man hours yi. Okay. And if I try to see the relationship, it's coming out to be straight line. Right. Now coming on to the concept of how do we calculate the best fit line right? [02:45:58] So this is the le square concept where we are trying to find out the RSS. So over here what am I trying to do again? Y I minus [02:46:10] what am I trying to do again? Y I minus the ycap the whole square not this one. Okay, here it is doing hidden trial method. But ultimately the le square estimators can be found out by trial and error [02:46:26] method. But that is not the case. We try to calculate it by the normal equation. This is the equation to for calculating of beta KN and beta 1. So these are [02:46:38] already predefined formulas which run at the back end when I run the linear regression function. So this example shows that whatever values you are getting by solving mathematically or through linear [02:46:53] mathematically or through linear regression model I get the same values. Okay. So what does polomial regression on data set work? Again we are using the on data set work? Again we are using the TV marketing CSV features from the skarn [02:47:08] prep-processing we are including the polomial features again first we are polomial features again first we are splitting the data so now I am creating an object of polomial features right with degree equal to two and now I am [02:47:24] using the fit transform function not the fit function itself but the transform function function because I'm transforming the input values as well as fitting but I'm only and only transforming the output. Do I want to [02:47:39] transforming the output. Do I want to train the testing data? No. Why I don't I want to train the testing data because so if I give you what is going to be if in your test if I tell you the questions is it actually a test? No. That calls [02:47:55] for paper leakage. So if I tell you what questions are going to come for an exam that is known as paper leakage and the same concept is known as data leakage [02:48:07] over here. Getting my point learners? So are you understanding the difference between fit transform and transform? The fit transform will only and only work for training data. Transformation [02:48:22] can happen will happen for input as well as output will happen for training as as output will happen for training as well as uh testing but fit will only happen for training data. When I create an object of linear regression [02:48:38] I get this output. So basically the polomial features transformer is configured to generate polomial features up to four degree. Then we are only transforming the input to generate a new feature set that includes polomial [02:48:53] features and interaction. Then we transform the test data corrected from transform the test data corrected from fit transform. It is used on x test to apply the same transformation and then train the model. The linear regression [02:49:07] model is trained using the transform training data. Clear? And finally we are going to predict the output [02:49:25] see the scatter plot is testing data and plot is x range polomial predict the output. Now clear you are getting all the code [02:49:37] how the graph is getting created. So what what is the observation and this is exactly you know why Jupiter notebooks are hit in the market because we are able to see the graphs over here directly which VS code lacks right we [02:49:53] can see the code as well as I can write my observations also so as you can see the regression line is able to fit majority of the data points you can majority of the data points you can infer from the above implementation that [02:50:08] nonlinear inputs require by a nonlinear model such as polomial model clear. So I mean regression linear regression also has capability of dealing with [02:50:21] nonlinear data up to degree 4 but not very high nonlinear data right then we have different variations of the algorithm and if we talk about performance metrics for analysis why different metrics [02:50:38] because it calculates the average of the squares of error which is differences between the actual and the predicted. Then we have the root mean square error. Then we have the root mean square error. Oh, [02:50:55] I hope you all are clear. Mean absolute error, R square error, zero value indicates model explains none of the variance in the dependent variable. The independent variables have no explanatory power for the changes in Y [02:51:10] and one represents a perfect fit. The model explains all of the variance in the dependent variable. The change in y are perfectly captured by the changes in the x. Clear? [02:51:25] So now we have to start understanding another technique important technique. another technique important technique. Okay. So now let's understand what is cross validation. And before we understand cross validation, we need to [02:51:41] understand cross validation, we need to understand few more concepts. terms that you need to understand. One is parameters and other one are hyper [02:51:56] is parameters and other one are hyper parameters. between parameter? A lot of learners are you know get confused because they both think they are the same thing. No, again I'm telling you I try to give you [02:52:12] exactly you know the cris concepts differences between them. So when I talk about parameters they are values learned by the model during the training. So [02:52:24] during the training what the model learns is referred to as parameters and hyperparameters are defined by the user to control the learning process. [02:52:38] Purpose over here is it directly impacts the model predictions. So model automatically learns the parameters you know and that is directly impacted in [02:52:51] the output. Whereas hyperparameters are controlled by us. For example, how much is going to be the training data and the test data is a hyperparameter [02:53:04] and the test data is a hyperparameter test size or train_plit. Right? They the factors or parameters which are controlled by us during the training process or model fitting process that are known as [02:53:20] hyperparameters. Whereas when the model learns the parameters on its own using the fit function are parameters. Clear? They are estimated during model training. They are generally uh you know estimated [02:53:35] before the training begins. learn from the data using optimization. the data using optimization. Whereas the hyperparameters are set [02:53:52] methods like grid search, random search or basian optimization. Influence on the model affects the output directly affects the speed and the quality of the learning and this is dependent on the data set. These are independent of the [02:54:08] data set. So the model parameters are nothing but your coefficients beta, beta 1, beta_2 or they can also be represented by weights in general. So [02:54:20] generally coefficients become weights. That's the general term used in deep learning. Whereas the test underscore size, the number of iterations and there are several other factors that will be controlled by us even random state are [02:54:37] all hyperparameters. Is the difference getting clear? getting clear? Right? So now what is the need of cross validation? Do we cross validate results? Do we want do we like to cross [02:54:54] validate views, reviews of doctors, lawyers? What does that mean Adinetri? And why do we want to do that? What does it mean? And why do we want to do that? [02:55:06] So that you know we are more sure of the answer, right? If one of the doctor is say saying that you know you need to be uh you know maybe you know your disease has this and maybe you need to get operated and even the other doctor says [02:55:19] that means you're more correct that yes or if the other doctor says no no no you it doesn't need to get operated you know you can cure it through these medicines you can cure it through these medicines right so cross validation or also known [02:55:33] as rotation estimation or out of sample testing refers to the process of testing refers to the process of rotating or splitting the data into rotating or splitting the data into different subsets. So it helps in uh you [02:55:47] know in the process of rotating splitting the data into different sub splitting the data into different sub sets. Okay. So one part we are very very clear that when we take the data set we divide it into training and test and [02:56:01] definitely we don't want to mix any of the training and testing data sets to avoid any kind of data leakage. That point is also clear that we want to avoid any kind of data leakage. That [02:56:15] points also makes it clear. Now to make our results better, more confident, more accurate, rather than using only one part of the set, if I use [02:56:28] only one part or one sample of the set, if I use multiple samples, I can get better results. Right? [02:56:40] So this is my mini training data set. So it's always the training data set which gets divided into further training and validation split right. So what is the idea that we [02:56:55] initially the data is getting split into training set and test set train and tune tune your models using cross validation. So we will try to find out okay we want to make certain changes this is better not better only through the training [02:57:09] set. The test set is not involved. Test set will we do not touch this until the very end because ultimately that is going to give us whether the model is going to give us whether the model is robust or not. Is this point clear now? [02:57:26] So how do we go about cross validation? So model evaluation is this when we fit the model and then we predict the test set. That point is clear that we have one data set that is divided into training and test set. Right? This point [02:57:43] this is normally that we do we fit it and then we predict the output and do the comparison. But if I want to do model selection between different comparisons what I will do now the training data set is [02:57:59] going to get uh you know divided into train as well as validation data set. What is the use of this model selection? It helps us to explore the grid. Now there are different hyperparameters. [02:58:16] How do you how do I know that this is the best hyperparameter for me? I will try to do it through the validation set. Fit on the train evaluate on the validation. Pick the best hyper parameter. [02:58:30] Clear? Why are we doing this? because we want to validate our training and results better. It helps us in finding out the different hyperparameters, right? So as we move along more uh you know [02:58:46] algorithms, you will understand it more. It's not part of regression but part of supervised learning also. So cross validation when data set is too small validation when data set is too small for apply hold out strategy then cross [02:59:00] validation can be used for evaluation and model selection. So what do we do? Suppose this is my 100 rows in data set. Okay. And now I divide [02:59:12] it into 2020 20 groups. So I have five groups 1 2 3 4 5. So the one of the part is going to be used for validation other is going to be used for training. Then [02:59:26] the next group is use going to be used for validation next. So I will iterate for validation next. So I will iterate it for five iteration. Got it? Of course it is increasing the complexity but it is giving me a much more validated [02:59:42] result. Agreed? So one of the uh you know u uh So one of the uh you know u uh techniques or variance of kfold is leave techniques or variance of kfold is leave this is known as leave one out [02:59:58] cross validation right so we have 1 2 3 and n over here right so if we take one part of the sample and the rest n are taken for [03:00:13] training then the next row is taken and the rest for training. So how many times the loop will be executed end times. Is it a good idea? No. Taking each sample for test sorry for validation and others for training is not a good idea because [03:00:29] it will increase the number of computation. But if I have kfold that is I divide it into number of groups my number of iterations decrease and I get [03:00:41] better results. So how do we go about the performance and the output metrics right? So what will happen it will so the cross validation will automatically create an array in the first iteration I will get [03:00:57] the first metric when it takes the second kfold I get the second performance metric when I take the third iteration I get the third performance metrics fourth the fourth one and last one I get the fourth fifth performance [03:01:11] one I get the fourth fifth performance metrics clear so to answer that point everybody please understand the test data is not touch it is the training data and the different folds which I use and final evaluation is done on the test [03:01:26] data whether I am getting the correct output so the concept is same we are testing and that will give me the results whether it's overfitting underfitting etc clear yeah but when we talk about [03:01:42] categorical data what does categorical data means that the data is in category data means that the data is in category in terms of males and female. [03:01:55] cross validated will be taken up for final testing. Final testing will be on the test data only. Once the parameters have been found out okay the number of have been found out okay the number of folds okay the number of uh you know u k [03:02:09] value for this thing is this. So now we will test on this value. Are we getting the minimum results or not? It will help us to calculate the different us to calculate the different hyperparameters. Got it? Okay. So when [03:02:22] we talk about categorical data dividing it into male and female. So this is my round one. So stratified kfold is being used for [03:02:34] categorical class which helps us to keep the ratio of the different categories same. Okay. Okay, this we will do it when we do uh classification. Okay, so what is the [03:02:48] difference between kfold and stratified kfold? K-fold is random. Stratified kfold helps us to maintain proportions may vary across folds maintains class uh distribution across folds. Imbalance data set not ideal preferred use cases [03:03:04] balance and imbalance. And how do I check the metric that my cross validation is good or wrong? First my uh you know parameters for cross validations are the estimators input output scoring and CV equal to 5. [03:03:20] output scoring and CV equal to 5. Estimator is equal to the model object. X is an array of the feature values. Y is an array of the target value. CV the number of folds. Scoring the metric and [03:03:34] it returns score an array of scores at each split. So you know the beauty of this particular function is that it returns at you know score of each uh [03:03:46] iteration. So you are able to see which one is better but generally we take the average of that clear. So now let's get back to the file and do it practically. [03:04:02] So this file is quite long. We will be using this file in the next session. Of course, we will not be able to complete all the concepts today. So, we are on cross validation. After cross validation, then we have regularization. [03:04:18] After regularization, then we have uh hyperparameter tuning that is model optimization and then the pipeline. So [03:04:30] it's it's a long journey that definitely we are going to continue okay in this file. Yeah. So cross validation technique everybody is there validation technique everybody is there with me? 3.7.2. [03:04:45] cross validation technique? Tell me learners. So cross validation is a machine learning technique that evaluates the model performance on evaluates the model performance on unseen data by dividing the data into [03:04:59] multiple folds. In each iteration, one fold is used as a validation set and the remaining as training. That point is clear that the process is repeated. So we have to give the number of iterations also. [03:05:13] Yes is repeated so that each fold serves as a validation set once and the results from all iterations are averaged to provide a robust estimate of the model [03:05:25] performance. Some of the common cross validation techniques [03:05:41] the number of equally sized folds that if it is unequal it will make it equal by adding randomly some of the you know randomly the data set only for the last [03:05:53] one right and there is no harm also in doing it will not affect it much so the model is trained on k minus one folds and tested on the remaining fold that point is clear so if I have five folds then one of the fold will become the [03:06:09] valid for validation test and four will be tested. This process is repeated K times. So is this K also a hyperparameter? [03:06:21] Yes, we decide they are going to be fivefold, sixfold, 7fold, eight folds, not the system. They are not parameters to the algorithm. It is decided by us. [03:06:33] So it's a hyperparameter. And with each fold exactly once as the test data set the results are averaged to produce a single performance estimate. We generally average it uh them. What is the advantage of kfold [03:06:49] cross validation? It provides more accurate estimate of the model performance. Yes, that this is the answer. K is the number of folds. Yes, AJ. And consationally intensive for large data set. Yes, you have to [03:07:03] reiterate the training process. But it gives definitely better results. Clear? Then we have stratified kfold. Similar to kfold but ensures that each [03:07:16] fold has the same proportion of different classes as the original data set. This is especially use useful for imbalanced data set. More reliable performance estimates for imbalanced data set [03:07:31] data set and still computationally expensive. Third is the hold out method. Simple and fast. One uh we we divide the data into training and test and the training data is first further divided into training [03:07:46] and validation data set. So the model is trained on training set and evaluated on the test set. It is simple and fast. The evaluation may be noisy and variability in the training. The normal one that we do we divide it into training and test [03:08:01] do we divide it into training and test is known as the hold out method and leave one out cross validation is again a special kind of cross validation technique used in kfold. A special case of kfold cross validation where k is [03:08:16] equal number to the number of data points in the data set. Each observation is used once as a test set. The model is trained on remaining data points. Clear? [03:08:29] So this maximizes the amount of training data used. Cons: Extremely computationally expensive especially for the large data set. Clear? Four variants of cross validation techniques. [03:08:44] Kfold stratified hold out and leave one out cross validation. Okay. Now can we begin with the [03:08:57] code? So I hope everybody is familiar with the So I hope everybody is familiar with the pandas and the mattplot lib library. We are also familiar with the skarn model selection. These are the different uh [03:09:11] selection. These are the different uh techniques and metrics linear regression over here and this metrics to compare mean square error absolute R square etc. Clear? So here we are talking about housing. [03:09:27] CSV. Yeah, take it as housing. So what are the pre-processing steps? What are the steps involved in EDA? Tell me how do you perform EDA? You'll do head tail to view the data, check null values, do info, [03:09:42] info, right? So let's look at info. So if you look at info, it has around 20,640 rows and 8 columns. Do you see [03:09:54] this? All are integer values. Do you see this? me which is the what is the output over here? Why it is a regression problem? What are we trying to do in this data set? Can you analyze it through the [03:10:10] number of column based on the latitude, longitude, based on the latitude, longitude, housing, age, total rooms, bedrooms, they all are inputs. I am trying to predict the house value. What will be [03:10:24] the value or the price of the house? House price prediction. Now clear all these are my inputs and one output. House values are numerical continuous value. Therefore, [03:10:40] value. Therefore, this is known as regression. Clear? Okay. Now, let's let's see the observation that price lies uh between 1.1 million to 2.6 million. Houses are generally 18 to 37 years of old. Housing [03:10:56] generally 18 to 37 years of old. Housing data. Let's check the null values. Yeah. Now, let's check the null values. So total number of bedrooms we have 27 So total number of bedrooms we have 27 values. How do we deal with null values? [03:11:11] Either we will replace it by mean, median or zero. Since null values make median or zero. Since null values make up only 1% of the total data, rows and column features with missing values will be removed. So what are we trying to do? [03:11:24] We are going to drop now because 207 is hardly 1% of 20,000 rows. So we can drop hardly 1% of 20,000 rows. So we can drop those data and now my data set is clean [03:11:36] right and this is the categorical data which we are not using we don't have that and now can you tell me what does this mean x and y now I have dropped the median house value access one because that's my [03:11:52] output and now this becomes my output this becomes my input Right? So I say print X and then [03:12:09] print Y. Do you see? Yeah. So now do you see this? So X basically has longitude, latitude, house, median, total bedrooms, [03:12:23] etc. And Y is only that. So let's remove Y. So I keep making changes in the code. Are you understanding? So these are my now the input and my output is median house value and [03:12:39] and my output is median house value and then I use the train test_plit. then I use the train test_plit. Clear? Okay. Now let's perform kfold cross validation. It implements kfold. Number of splits is 10. Divides the data [03:12:52] set into 10 10 folds automatically. random state is 42 shuffle is also equal random state is 42 shuffle is also equal to so I have created an object of this to so I have created an object of this kfold now I initialize the model there [03:13:06] kfold now I initialize the model there with me score based on the model was linear regression training data and the target [03:13:20] variable scoring I'm using negative mean absolute absolute error as the performance metric. So this is where the absolute error comes into the picture. CV cross validation is KF. Number of jobs is equal to minus1. What does this [03:13:37] mean? It utilizes all the available processes for parallel computation. Do we understand the concept of threading or the number of core processors learners? So this parameter takes care of that. [03:13:53] So this parameter takes care of that. Okay, this parameter takes care of that Okay, this parameter takes care of that particular point. Clear? So from statistics import mean I'll take the mean of the k4 cross validation kfs [03:14:08] fours are not defined. So let me run this. Let me you know show you that since the number of folds is 10. So the output of kfold crow is an array object. See do you see they all are mean negative [03:14:23] you see they all are mean negative absolute errors negative does not mean that the error is less it's just the sign okay so we will try to take the average of these right there will be 10 outputs [03:14:37] since we have given 10 phones 1 [clears throat] 2 3 4 5 6 7 8 9 10 clear it's an array [clears throat] and then I will try to [03:14:49] take the absolute absolute average of this clear. So we have trained the model and evaluate on the test set. So are we now understanding how are we going about cross validation score learners [03:15:06] and now finally after cross validation now we can use the test data. So this is how I have fitted the data predicted it through the test and my MSE test MSE is [03:15:18] through the test and my MSE test MSE is here. Similarly, my R square is here, right? So, my test MSE and R square scores are here, right? This the look at the error. It's so huge indicating that and on average the squared prediction [03:15:33] errors are large. This can be interpreted in context of units of dependent variable which are likely in order since the numbers are very high. That is why the errors is coming out to be large. [03:15:47] This can be mitigated by scaling the features. Feature scaling is critical in machine learning to ensure that all features contribute equally. features contribute equally. So this is how and then we demonstrate [03:16:00] leave one out also loves loves and we get the absolute mean score. So the mean absolute error is high to improve the model complex more complex [03:16:17] model can be considered which will be discussed in further lessons. So the concept of cross validation is clear. Let let's do a quick knowledge check on that. Okay a quick knowledge check. First question [03:16:31] First question first question. Which of the following which of the following cross validation versions may not be suitable for very large data set with hundreds of samples? We just now studied that [03:16:45] practically you know practical has a major uh impact. We all have just now seen the impact and it is leave one out cross validation. Great. [03:16:58] Which of the following is a disadvantage of kfold uh cross validation method? Training algorithm has to return from scratch. Do we understand this? every [03:17:11] scratch. Do we understand this? every time it has to restart and do it again. Next question. Suppose you have picked the parameter for model using 10-fold cross [03:17:23] validation. Which of the following is the best way to pick a final model to the best way to pick a final model to use and estimate its error? Train a new model on the full data set using the parameter you found. Use the average [03:17:37] parameter you found. Use the average cross validation error as its error estimate. So you can use average cross validation error as its error. Everybody got this? Why C is not correct? Because we will train the model on the full data [03:17:51] we will train the model on the full data set. Okay. So what is the idea? The best way to pick why is the answer B correct? The best way to pick a final model is to The best way to pick a final model is to train a new machine learning model [03:18:04] on the full data set using the parameter learned to use the average cross validation error as its error estimate. So please be clear with this particular So please be clear with this particular point that is why I have added this MCQ. [03:18:18] We can compare different models using cross validation. Cross validation is mainly used for comparison of different models. For each model you may get average generalization error on the K validation sets. Then you will be able [03:18:33] to choose the model with the lowest average. Clear? And cross validation is also used for And cross validation is also used for model checking not model building [03:18:45] because it allows to repeatedly train and test on a single set of data set. Let us suppose we have a linear regression model and a neural network. To select the best one among these we can use k-fold cross validation. So to [03:18:58] compare between the models also we can use cross validation to select a better use cross validation to select a better performing model. Clear? performing model. Clear? So when we are now looking at the [03:19:11] regression outputs right or the regression analysis some of the outputs are like this. So the output over here that we receive is [03:19:23] in these terms of coefficients that this is my beta kn this is my beta 1 beta_2 and this is my beta 3. So this is my north south east and constant term and [03:19:38] these are the values right. So what is it showing? This is the constant value positive. East has positive relationship but south has more positive relationship with the heat flux. Heat flux is the output and north has a negative [03:19:53] relationship with the output. Are you now understanding it better? How do we get the output and how are we relating it with the uh you know the equation of linear regression. So this is nothing but like [03:20:08] beta kn plus beta 1 x1 the value of east plus beta_2 x2 minus. So a lot of you had question [03:20:23] that how do we understand negative it will automatically get this. So this is will automatically get this. So this is multiple linear regression. Yes learners this is multiple linear regression and there was lot of confusion regarding [03:20:37] ma'am what does multi-olinearity mean? Multi means referring to multiple independent variables multiple inputs with multiple regression. Call means to [03:20:50] uh join or together referencing to the linear movement or correlation as I told you tries to find out correlation in terms of minus1 + one occurring within the linear equation and suffix means the idea. So if we look at over here from [03:21:07] statistics we understand that p value is a very very important term a very very important term right that if I have the value lesser than 000.5 then this is accepted this is accepted [03:21:23] this has a strong evidence statistically significant but if this is not lesser significant but if this is not lesser than 0.05 05 that means east is 2.12 is than 0.05 05 that means east is 2.12 is not very statistically uh proven or [03:21:37] not very statistically uh proven or confident that the value is this clear confident that the value is this clear and when I talk about VIF how do I calculate multi-olinearity it is the variance inflation factor [03:21:51] which is coming out to be 1.21 21. What does that mean? This is the correlation map, a heat map that you understand. So multi-olinearity that you understand. So multi-olinearity is the phenomenon of high correlation [03:22:05] between the predictor variables can create instability and bias in regression model. To identify and address multi-olinearity, we use the address multi-olinearity, we use the variance inflation factor. So the VIF is [03:22:19] variance inflation factor. So the VIF is equal to 1 - 1 upon R² and if the value is equal to 1 that means all the input values are independent and if it lies between 1 to 5 it suggests moderate correlation over here we are getting the [03:22:35] value between 1 to 5 so there is moderate correlation or we can say moderate correlation or we can say independent also and if it is greater than five then it indicates high correlation that means [03:22:47] indicates high correlation that means then linear regression cannot be fit uh then linear regression cannot be fit uh can be fitted on that input values. So can be fitted on that input values. So where e vif should [03:23:03] dummy variable or a nomial um variable and multiolinearity reduces the statistical significance of the independent variables. VIF is used to [03:23:16] detect these variables. A large variance inflation factor on an independent [03:23:28] linear relationship to other variables that should be considered or adjusted that should be considered or adjusted for structure because multi-olinearity is one of the assumptions that we uh you know uh you [03:23:41] assumptions that we uh you know uh you know assume when we are uh trying to build the linear regress. regression model. So that means if the VIF value is greater than five then linear regression should not be used. Clear? [03:23:58] should not be used. Clear? I hope these points are clear. Now first concept that we will learn data leakage in machine learning. So now we will understand the concepts of pipeline. Today pipeline is a technique [03:24:13] for automating different processes. What what is the different processes that we do whenever the data is loaded? We want to do transformations such as encoding, scaling, right? So we'll try creating a pipeline and [03:24:29] what is our main aim that we want to avoid data leakage of course. Why? Because if there is data leakage, the data gets lost. A lot of information is also getting lost. So now if we look at the data from [03:24:47] the supervised learning perspective, it is divided into two parts. One is known as the training data and the other one is known as the test data. Agreed? [03:25:01] is known as the test data. Agreed? Okay. So a scenario when the ML model Okay. So a scenario when the ML model already has information of a test data already has information of a test data test data in the training data. Do you [03:25:13] think that test data should be present in the training data? If I tell you okay in your exam these questions are going to come that going to be very beneficial for you for scoring marks but do you actually learn out of it? [03:25:30] actually learn out of it? No, that's not a good uh way of learning right that you know it can give you good results but you are not going to become you're not going to be a robust model. That means if any other question is [03:25:43] asked from that particular topic you will absolutely fail right but this information would be available at the time of prediction called data leakage. So what is the disadvantage and how can we avoid it [03:25:58] that it causes high performance while training set but performs poorly in the training set but performs poorly in the deployment or the production. [03:26:15] say when there is an overlap of training and test data then it causes data and test data then it causes data leakage right and basically data leakage happens due to two reasons. First we [03:26:30] understand train and test contamination or the target leakage. What do we mean by that? Target leakage occurs when the model is trained on the training data [03:26:42] that contains target or the feature information. So we don't want that and that should not be available at the time of prediction. So we have to be very very careful when we are doing this. The other one is the train test [03:26:59] other one is the train test contamination into the training data and the data prep-processing steps for transformation [03:27:11] for example scaling encoding are applied before the splitting the data set. So contamination is when the test data leaks into training that and the pre-processing step so that uh the scaling encoding should not be applied [03:27:30] before the splitting of the data set to avoid data leakage. Got it? Basically if the there is mixing of training data in the test test data set [03:27:45] or rather test data in the training data set then this causes data leakage and we have to avoid it in every case that is we should not apply any transformation [03:27:59] that is scaling or encoding before the splitting of the data set. Now clear. splitting of the data set. Now clear. Now another very important concept of regularization. There are two types of regularization [03:28:15] There are two types of regularization available under regression that [clears throat] is ridge and lasso. Preventing overarning. Yes, normalizing the data. Yes, we can say that. So [03:28:28] regularizing thing in in a normal way, right? We want to prevent it from overfitting overarning. Right? So regularization is a technique in machine [03:28:40] regularization is a technique in machine learning which prevents overfitting of the model. And how do I know that the model has overfitted? How do I know that the model has overfitted? [03:28:54] Test error much higher than the training. The training error. And why is it that model is trained so complex than requests high variance it covers every possible train outcome making it yeah so the complexity of the model increase it [03:29:07] tries to train on each and every data point and fails on the final test data due to noise due to learning of noise in the training data so to avoid it so as [03:29:20] I've been telling you that in machine learning do we suffer more from learning do we suffer more from underfitting or overfitting overfitting right so [03:29:33] we need methods so that our model does not overfit and we need certain control parameters to achieve that. So there are two you know algorithms under this that [03:29:45] two you know algorithms under this that is ridge and lasso. So lasso is known as L1 regularization technique. Can anybody tell me what does this term mean? Anybody? It is summation of the actual value or [03:30:03] the true output minus the predicted output the whole square. What is this term known as? It's an error. I want the typical name of this error. It is the typical name of this error. It is the residual sum of square error. But rather [03:30:19] residual sum of square error. But rather than only taking this error actual minus the predicted value the whole squared I take into account the another hyperparameter lambda which I will use to control or [03:30:36] regulate the training process. So do we have a regulator? Do we have a regulator to control the speed of the fr fan? Similarly lambda we will use the Similarly lambda we will use the regulator and similarly summation of [03:30:50] beta. Can anybody tell me what is this beta over here? beta over here? What does the term beta mean? It is the coefficients. It could be a single, it could be many multiple variables. So [03:31:03] this is known as the regression coefficients. is only in twodimensional case like beta plus beta 1 x1. I agree both of you are [03:31:18] clear uh correct on those perspective but in general betas are known as regression coefficients or the weights parameter. [03:31:32] So what is the advantage that we are getting through this L1 regularization getting through this L1 regularization that in the error term in the RSS term that in the error term in the RSS term now we have added the penalty lambda [03:31:45] which is controlled by us and which controls the value of these regression coefficients. Getting my point? [03:32:01] regularization. So this is now also you know known as my error or it is also known as the cost function. Try to understand in deep learning the error with which we are calculating [03:32:15] error with which we are calculating becomes my cost function. In this case it is represented as L over here. It represents it as my loss function. there is slight difference but actually in machine learning they all mean the same [03:32:31] thing. Okay. this W over here? Is it the same as beta? [03:32:50] regression coefficient. Is the lambda same? Yes. The lambda in Python is known as alpha and then we have sum of the square of the weights and what is this [03:33:03] actual value predicted value. Now are you understanding whether it's w or beta they same mean the same and this is the lambda parameter with the square of the coefficients. Okay. So what are the advantages and disadvantages and now [03:33:19] advantages and disadvantages and now then we will practically jump onto the file and start learning from there. So L1 regularization performs feature selection. What do we mean by features [03:33:34] independent variables or the inputs? Yes, the inputs. So now we are trying to control. So now what is the use of overfitting? We understand that [03:33:46] overfitting generally leads to high variance. Agreed learners high variance. Agreed learners and complex models. Agreed? Do we and complex models. Agreed? Do we understand this concept everybody? So to [03:34:01] reduce the complexity now I will try to control my coefficients regression control my coefficients regression coefficients beta KN B1 and beta n or in [03:34:13] general I can also call them as weights now getting my point do not get confused with the terminologies it somewhere it would be written as weights somewhere as beta somewhere as independent variable somewhere as features so you should be [03:34:29] able to understand the concept of it. Clear? So understand the concept of it. Clear? So the beauty of L1 regularization is that [03:34:41] it performs feature selection by shrinking the less important features weights to zero. So if I have a multivariable multivariable data set from beta 1 to beta to beta 7, [03:34:55] it will shrink few of the features. That means some of the features or the weights will become zero. That means they have no significance relationship they have no significance relationship with the output. For example, if we want [03:35:10] to yeah just try to listen and absorb as much as you can that suppose you know we want to predict the price of the car. Okay. So there are several factors from fuel to design to color of the car to the alloy of the wheels to the security [03:35:25] systems to the infotainment system to the sunroof. They can be several uh features but maybe the alloy of the car or the wheel of the car might not be or the wheel of the car might not be important. So it can be reduced to zero [03:35:39] that will help us to reduce the complexity and that will try to reduce the overfitting of the model. So very very overfitting of the model. So very very important concept L1 is used to reduce [03:35:54] or shrink the less important features weights to zero. So it helps us in feature selection. Right? And it can also be used for high dimensional data [03:36:07] set with many number of columns such as 20, 30, 50, 100 with many irrelevant features. And the disadvantages it is not effective for data set with many important features. Now you might say [03:36:23] ma'am how can we remove a feature? My every feature is important for the every feature is important for the output. Then we use L2 regularization. Okay. Where the number of the where the value [03:36:37] of the coefficients of the weights will not become zero. It provides a smooth solution and improves the generalization performance of the model. So ma'am ridge [03:36:50] is going to doing the opposite instead of zero it is reducing the weight. of zero it is reducing the weight. Ridge is trying or uh you know it will [03:37:02] it is trying to see the first one is making it zero and the L2 will try to making it zero and the L2 will try to reduce the value. L1 is lasso and the ridge one is making the value of the coefficient small not huge or big but [03:37:19] small. So therefore the biggest advantage is that it can handle data advantage is that it can handle data sets with many important features. sets with many important features. Okay. So now moving ahead to great ain. [03:37:35] Everybody has the file. Everybody's ready in the Jupiter notebooks. Okay. ready in the Jupiter notebooks. Okay. So, regularization in regression in linear regression. Regularization encompasses a set of techniques employed [03:37:51] encompasses a set of techniques employed to address the issue of overfitting. So, what is regularization? Regularization is the method techniques to achieve the is the method techniques to achieve the objective by introducing a penalty term. [03:38:05] Please try to understand again it's a very very important question from the point of interview introducing penalty term lambda to model's objective term lambda to model's objective function to prevent it from [03:38:19] overfitting. This objective function typically measured by mean square error is minimized during the training process. [03:38:31] So generally the error of the mean square error or the RSS to be precise is minimized during the training process. What is the advantage of that penalty term or the lambda? The penalty term discourages the model from attaining [03:38:47] excessive complexity by penalizing the size of the model coefficients. So model coefficients, regression coefficients beta, beta 1, beta_2 or the W1, W2, WN [03:39:01] thereby mitigating the overfitting process. Got it learners? All right. So now if we talk about the regularization term alpha, it can be [03:39:13] regularization term alpha, it can be known as alpha as well as lambda. Okay, they all mean the same thing. It's written as lambda or sometimes as alpha. So do not get confused. Okay. So there is a little terminology mishap [03:39:28] happening. So everyone uses their own technique but try to grasp the concept. technique but try to grasp the concept. Okay. [03:39:40] scales the penalty term. It controls the strength of regularization. Higher the alpha, it imposes stronger penalty on the coefficients. That means they tend to become zero leading to greater regularization. [03:39:55] This tends to produce a simpler model that may underfeit the training data but that may underfeit the training data but often generalizes better to unseen data. If we talk about lower alpha, it imposes a weaker penalty leading to a model that [03:40:11] is less restricted by regularization and more complex potentially capturing more more complex potentially capturing more details in the data but at the risk of overfitting. So what are we looking at? We are definitely looking at a value [03:40:24] which is not very high and very low. So again minimum mid value to find out. Got it? How do we find it? What are the best methods that we will understand [03:40:37] today? That is known as hyperparameter tuning. That's part of this today's tuning. That's part of this today's session also. Okay. Now, what are the benefits of regularization? [03:40:50] It enhance the generalizability of the model by mitigating the overfitting model by mitigating the overfitting factor. Regularization fosters model that can perform well on unseen data. reduce model complexity. It promotes [03:41:05] reduce model complexity. It promotes interpretability and potentially reduces computational cost associated with training complex models. And the two common regularization techniques are L1 and L2. L1 is lasso. [03:41:23] Lasso is a full form of least absolute shrinkage and selection operator. Right? shrinkage and selection operator. Right? So absolute uh term is there as the penalty. So Simon is there any standard to label it as high, alpha or low? Yeah, [03:41:39] to label it as high, alpha or low? Yeah, we'll understand. negative values and high values are in thousands and lakhs. Let's do it [03:41:51] practically to understand that point better. Okay. So is the full form of lasso clear to everybody? So which is the penalty term? Of course, it is the alpha or the l uh you know the lambda. But what are we trying to add? The [03:42:06] But what are we trying to add? The absolute value of the regression coefficient. So the least absolute shrinkage and selection operator regression relies upon the linear regression model but additionally [03:42:20] regression model but additionally performs a so-called L1 regularization which is a process of introducing additional information in order to prevent overfitting. As a consequence we can fit a model containing all possible [03:42:35] can fit a model containing all possible predictors. What are predictors? feel ma'am is ma'am just reads the data but I'm purposely reading it to make you [03:42:47] understand line by line. So what is containing all possible predictors? Predicted value is an output. Predictors are input. So we can fit a model [03:42:59] containing all possible predictors and use lasso to perform a variable selection by using a technique that regularizes the coefficient. So what is that lambda parameter doing? It is having control on the value of the of [03:43:14] having control on the value of the of the the lambda or the alpha has it's the coefficients the coefficients the regression coefficients that's up to you how you want to represent it as betas or [03:43:31] weights right so it performs variable selection or feature selection it forces some of the coefficient estimates to be exactly equal to zero with the help of large tuning parameter. So more the value of [03:43:47] the lambda some of the features coefficient value will turn out to be zero. And that is why this is a technique which is also used in dimensionality reduction. What is dimensionality reduction? Reducing the [03:44:02] number of features in the data set so that it reduces the complexity. So L1 that it reduces the complexity. So L1 again plays a major major role in that. It reduces it helps to reduce learning of more complex data and overfitting. It [03:44:17] decreases the variance of the model without increase in the bias. All right. So in minimization objective does not include RSS like the OS regression but include RSS like the OS regression but also the absolute value term. So this [03:44:32] RSS now is clear. This is the residual sum of square and this is how we can expand it and write it. This this point is also clear. [03:44:44] What is yi? This is the actual output. And what is this output? This is the predicted output. Are you all understanding it mathematically, conceptually? But in lasso, where does the difference [03:44:59] come? The RSS is the same but we have added the penalty term alpha or lambda along with the absolute value of the summation of the coefficients. Now clear [03:45:12] if the alpha is equal to zero then there is no regularization that will happen because the error term will be same as the RSS. If it is equal to infinity all [03:45:25] the coefficients will become zero. And if the alpha is greater than zero lesser than infinity coefficients are between zero that of le square linear [03:45:37] zero that of le square linear regression. Got it? Is the theory part clear? Now let's move on to the practical part. Start with a new data set. So let me share it with you. hitters CV ca dot csv okay here we go [03:45:55] please download this data set and be ready I hope different libraries are also clear the numpy pandas then linear model lasso the metrics mean square [03:46:07] error r square all these are there okay so basically now let's understand the description of the data set so I I have [03:46:19] description of the data set so I I have provided a link over here. So uh this is for like we have like now IPL matches. So here we are trying to predict the So here we are trying to predict the salary of the player based on atbat [03:46:32] salary of the player based on atbat number of times he batted in 1986 hits the number of hits the number of home runs the number of runs the number of runs the number of runs the number of runs batted walks number of times a bat [03:46:45] runs batted walks number of times a bat during his career. So lots and lots of during his career. So lots and lots of um parameters to decide the salary of um parameters to decide the salary of the player that is 1987 annual salary on [03:46:57] opening day in thousands of dollars. Got it? And then we have another important column that is new league a factor with a and n indicating the players league at the beginning of 1997. [03:47:13] So what is the first step? We've loaded the libraries. We load the data set and the view of the data set with df do head. Okay. What are the other functions that we would perform info. So this data set [03:47:28] we would perform info. So this data set is quite huge in the sense that it has is quite huge in the sense that it has the number of rows are only 3 30 uh 322 but the number of columns are many. So this is an object uh data type. Atbat [03:47:45] So this is an object uh data type. Atbat everything is integer but we have league everything is integer but we have league division as categorical data and then we have salary over here as the output and new league. So what are the [03:47:58] different steps that need to be performed? Can we directly apply the model onto this data set? So these are my columns. So I have used df.drop Drop [03:48:10] unnamed equal to zero. Encoding if any. Do you think encoding is required? Coding is encoding is definitely required for league division for all the columns which are of object data type. [03:48:26] Please remember this point learners encoding is definitely required for object data type. Without that you will not load it into the model. duplicate [03:48:38] values. If they are then we need to check that. Okay, we are removing unnamed column because that is not required. We are specifying the access. What does in place equal to true mean? Permanent removal of that [03:48:54] true mean? Permanent removal of that column from the data set. So now if I see my first column gets removed. So the first processing that I have done. Now let me check the null value. So are there any null values over here? Only [03:49:10] the null value is in the output in the salary field. Right? Which is my output field. Agreed learners? Do I need to separate my x and y also? Yes. So the number of missing value in [03:49:26] Yes. So the number of missing value in salary is 59. So 59 out of 30 322 observation with null values correspond to columns salary. Since we will use the lasso algorithm from the scikit, we need to encode our categorical also. Okay. [03:49:45] Now, how do we deal with categorical data? What is the use of valueore counts function? It gives the category along with the frequency along with the count. [03:49:59] Right? That's incomplete. It gives the category along with the count. So in league there is A and N with these categories. In division there is W and E western and eastern and again in new league we have A and N [03:50:15] and again in new league we have A and N the American le or the national le. So if I separate this into a data frame. Do you see this you see this data set data frame rather? [03:50:29] This code is getting clear to everybody. Are we here till here? Okay. So, what are the different ways of encoding data? One is one hot encoding. [03:50:42] Other one one hot encoding. Other one label encoding. What is the difference between the two? One encoding adds column and then gives the binary output. If that column value is there then it has one. Else all values are zero. Yes. [03:50:59] and label and assign integers to each of the category. Yes. And label encoding is the category. Yes. And label encoding is used when we have a ordinal data. One hot encoding is used when we have nominal data. That means there is no [03:51:15] order. So over here we are using which function learners? function learners? Yes, it is now one hot encoding that now it will have instead of three six columns with zero and one value [03:51:31] right so do you see this league A league N division E do you see league A league N division E do you see now so wherever we have value of Lee [03:51:43] get dummies is one hot encoding and label encoder that is the function so is the output of dummies do head is also clear that now if these are the six the three categorical value get converted into six and wherever the [03:51:58] value was there this is one or this is zero this is zero and this is one clear okay so what are we doing we are separated the output x numerical value [03:52:11] we are dropping the output the league the division and the new league because they all are categorical data and everything gets converted into float type. So the numerical columns are now clear. [03:52:24] The input X numerical and the output Y salary. Good chakra. Good. Now do we salary. Good chakra. Good. Now do we need to concatenate these columns with the original data? So we'll just take one of them since they it's it has [03:52:39] binary category. Please try to understand. Since all of them have binary category, therefore it can be used for league N, division_W and new league N. Getting my point? [03:52:55] N. Getting my point? So now you see all my inputs have become integer values, numerical values and that is what is required before you feed that is what is required before you feed any data into the model. [03:53:09] Now once my data is ready now we go in for testing and splitting of my input for testing and splitting of my input and output. Test size is 25% rest is and output. Test size is 25% rest is training and I get my four outputs. [03:53:23] Right? So lasso you know performs best when all the numerical features are centered around zero and have variance in the same order. Homoidastiticity needs to be maintained and if a feature [03:53:38] has a variance that uh that is orders of magnitude larger than others, it might dominate the objective function and make the estimator unable to learn from others. So it should not be that you know one of the columns have very high [03:53:54] values and the other one low. This means it is important to standardize our features. We do this by subtracting the mean from our observations and then dividing by standardization. What is this concept known as [03:54:10] a feature engineering? Okay, I've given you one hint. This is the concept you one hint. This is the concept involved in feature scaling and this involved in feature scaling and this refers to standardization that is [03:54:23] scaling the factors uh with mean equal to zero and standard deviation equal to to zero and standard deviation equal to 1. Clear? To avoid data leakage, standardization of numerical features should always no be performed after data [03:54:42] should always no be performed after data splitting only for the training data. Please try to understand to avoid data leakage, we will only and only use it after the splitting of the data. We saw that. So what is data leakage? You [03:54:58] understood what is data leakage that it occurs when the information from the occurs when the information from the outside the training data set is used to create the model. This can happen if the data would not be available at the time [03:55:11] of prediction is included in the training process. So data leakage can lead to optimistic performance estimates and models that fail to generalize well [03:55:23] and models that fail to generalize well to new as well as unseen data. Clear? So which is the Python function? Which is the Python function to perform standardization standard scaler which creates an instance of the object fit. [03:55:41] creates an instance of the object fit. Fit trains the X train numerical data that is training means to find out the parameter to compute the mean and the parameter to compute the mean and the standard deviation for each feature in X [03:55:55] standard deviation for each feature in X train list numerical. X train is your for each uh data set and the list numerical is a list of column names corresponding to the numerical features. Got it? [03:56:10] Got it? What is the use of transform? training data which transforms each feature in the training data set to have mean zero and standard deviation one. Right? And then we only use the [03:56:27] Right? And then we only use the transform function on the test data. Now clear okay [03:56:39] function to perform standardization. Standardization is one of the feature scaling technique where we take mean equal to zero and standard deviation equal to 1. Anoj that point is getting clear. [03:56:53] Fit is always performed on the training data that is numerical data. Till here also this point is getting clear [03:57:06] numerical data and transform we are doing it on the train data that is how we are going to avoid data leakage now clear [03:57:18] leakage now clear got it but when we talk about uh you know uh testing data are we going to apply fit function on it No, we are only going to transform it. So it applies the same standardization parameters mean and [03:57:35] standard deviation computed from the training data to test data and this ensures that the test data is scaled in the same data as training now clear. [03:57:49] Okay. Right. So practically how do we go about it? So we have imported the standard scalar from the skarn library and then we perform this function. So and then we perform this function. So the explanation is given above [03:58:04] right. So what are we trying to do? We are fitting the standard scala to numerical features so that there is no biasness in the data set. Now uh you know you'll see positive negative value because the sum of all these values [03:58:19] should be zero. The mean of all them should be zero. clear to everybody? [03:58:31] So the training data has 241 rows and 19 columns. Clear? Similarly, test one will only have 81 rows because of the 25% data and [03:58:43] 19 columns. Clear? Now, there were missing values in the output data. How do we go about it? [03:58:55] How do we go about it? Let's see. Yeah. So, what are we doing over here? Yes. So, what are we doing over here? Yes. What are we doing? We are taking the training model as well training data as well as the testing data and filling it [03:59:10] with the median value. Wherever there is null value, it will fill it up with median. Not mode over here. Dr. Henry we see it practically that it's getting filled with median. What is the difference between mean and [03:59:25] median? Besides that which one is affected by the outliers? Where do we use median when we do not Where do we use median when we do not want to get affected by the outliers? So [03:59:39] now let's move ahead and understand the lasso model. So whenever we have to create how do we begin with the hyperparameters? [03:59:53] how do we begin with hyperparameters? We will randomly assign any value and We will randomly assign any value and then start training on it. And to find hyperparameter tuning. So we apply lasso regression on [04:00:09] the training set with regularization parameter that is alpha equal to 1. So we begin with a very simple value alpha equal to 1. This value is commonly used [04:00:21] equal to 1. This value is commonly used as default and provides a good balance between maintaining model complexity and reducing overfitting. Okay. So by default whenever we have to do [04:00:35] thing you can randomly assign any value to the hyperparameter alpha is equal to one maximum iteration is equal to 10,000 and now we are fitting the training data and this is now my new intercept coefficients the regression coefficients [04:00:52] beta KN and the other coefficients because since they are how many parameters can you tell me how many parameters are can you tell me how many parameters are there in this particular question. [04:01:05] How many are there? 19. All the numerical values, right? So, we would have 19 coefficients. And do you do you see after applying lasso some of them see after applying lasso some of them have become zero? This one, this one, [04:01:21] have become zero? This one, this one, this one and this one. Getting my point? So if I see the output these are my coefficients and some of my coefficients [04:01:33] have become zero. Do you see? Then you might say ma'am what is negative0? It just puts the sign. There is nothing no term in mathematics as0. Clear? So what do we conclude? This is the [04:01:47] intercept term of your lasso regression model. It represents the uh expected mean value of the independent variable when all the independent variable when all the independent variables are set to zero. [04:02:02] In practical terms, it is the baseline prediction when no other information from the variables is provided. So the lasso coefficients represent the relationship between the variable positive coefficients. Now what do we [04:02:17] interpret out of it? Let's see that a positive coefficient indicates that as independent variable increases, the dependent variable also increases. That's the positive relationship. A negative coefficient indicates that an [04:02:33] independent variable increases, the dependent variable decreases. The magnitude of the coefficient shows the strength of the impact. Now clear? [04:02:45] So the magnitude of the coefficient shows the strength of the impact. A larger absolute value indicates a stronger effect. [04:03:00] ability to perform feature selection. So ultimately the values have become zero. ultimately the values have become zero. So now my 19 features have reduced two. So I have reduced these two. Then I have also reduced [04:03:22] features get reduced. So 19 gets reduced to 16. Right? When we have taken lambda to 16. Right? When we have taken lambda equal to 1. All right. Right. But in this model however it seems that none of the coefficients are [04:03:38] exactly zero suggesting that all included variables have some impact on the models through some impacts are those some impacts are [04:03:50] very very small. So basically four features yeah okay four feature. So now how do I evaluate it? I'll calculate my mean square error for the [04:04:02] training data and then the for the testing data. This code is also getting clear that helps me to decide whether it's overfitting or underfitting. So now you tell me is it an overfitting model or an underfitting model [04:04:22] than the training. So is it underfitting overfitting? So is it underfitting overfitting? it's overfitted [04:04:36] the value is very very less. So to better understand the role of alpha the regularization parameter. So what is alpha known as the parameter. So what is alpha known as the regularization [04:04:54] parameter? Plot the lasso coefficients as a function of alpha. Maximum iterate as a function of alpha. Maximum iterate are the maximum number of iterations. [04:05:07] function. What does the np numpy lindspace function do? that the now the values of alpha will range between 1 to five and with 10 equidistance value. So [04:05:19] now what are we trying to create link space function do we understand of space function do we understand of numpy? So now if I say now I will start understanding the range of the values range of the values from 0.01 01 right [04:05:36] range of the values from 0.01 01 right to 500 and the value is still 100 right to 500 and the value is still 100 right so over here these are the values or the array created of alpha the maximum iterations are 10,000 and now I will try [04:05:53] iterations are 10,000 and now I will try to run or fit on each of these alphas to run or fit on each of these alphas and then look at this beautiful graph Simon specially for So the value of the alpha goes from 0.01 [04:06:07] So the value of the alpha goes from 0.01 to 100,000 and we see as the value of the alpha is increasing some of the coefficients actually become zero. You coefficients actually become zero. You see this? [04:06:25] So that is why L1 regularization is also known as feature selection. Agreed? Okay. So, have you understood this graph now? So, we create now the best value of [04:06:40] alpha. How do we find out the value of alpha? That now we will use cross validation to find out the best value for alpha. Lasso regression comes with [04:06:52] builtin cross validation. So, beauty of lasso is that you can combine lasso with lasso is that you can combine lasso with cross validation. Alpha again ranges between 0 to 1,000 CV is equal to 10 maximum [04:07:07] to 1,000 CV is equal to 10 maximum iteration 10,000 and num n jobs is equal to minus1 can you anybody tell me what is the use of this parameter [04:07:23] I'll do that in the meanwhile tell me what is the use of n jobs learners training and the testing process and we found out that it is an overfitting [04:07:38] model the test error was too high than the others. So now we have to find out the others. So now we have to find out the uh you know the best value of alpha. Okay. Right. So use the maximum course and [04:07:52] which is the you know ma best alpha value which has come out over here six comes out to be the best uh regularization. So get the best alpha [04:08:04] regularization. So get the best alpha regularization strength selected by the cross validation clear. So it was not the value of alpha was not one but it the value of alpha was not one but it will come out to be six. And now I will [04:08:18] create and fit the lasso model taking this value. Now I will fit it and now check on my training and testing error. So have they improved? Well the test So have they improved? Well the test error is still higher than the training [04:08:33] error is still higher than the training right tuned parameters. So zero coefficients such as runs, RBI, [04:08:45] CAT and all these new leagues have several coefficients shrunk to zero. This indicates the lasso has deemed these features less important irrelevant for the data set and the nonzero coefficient give us the detail [04:09:00] relationship. Clear? Moving on to ridge regression. Ridge regression is another technique for handling uh overfitting of the data that's known as L2 regularization. Here [04:09:15] none of the features will become zero rather they will values will become rather they will values will become less. Right? So where it is used why where when to use that and when to use the other one get the best alpha [04:09:28] the other one get the best alpha selection validation. Now what is this added below library? Okay. Okay. So when is ridge to be used? It is useful for handling multi-olinear data where two predictors independent [04:09:42] variables are highly correlated to each other. AJ you are absolutely correct AJ Khana you are absolutely correct but the [04:09:57] second reason where we use ridge regression is that it is useful for handling multi-olinear data where two predictors independent variables are predictors independent variables are highly correlated to each other. Now [04:10:11] highly correlated to each other. Now getting my point. Collinearity refers to a situation where two or more predictor variables in multiple regression model are highly correlated meaning they have a linear relationship and this makes it [04:10:27] difficult to determine the individual effect of each predictor on the target variable. Right? So that's why the penalty term is added the regularization terms are added. Getting my point. [04:10:50] ridge CV function which can do so on the same data we will first apply the alpha fit function and get the intercept value. Yes, reach ridge needed to use if [04:11:04] features. So I've not been running the code. So I will get error. So sorry for that. I'm not running the whole code going back and running it. So now let's move and then evaluate the model again. Testing error is more than the training. [04:11:19] So overfitting is there. And do we see the coefficients? Now none of them become zero. Now we can use cross validation again to find out the best [04:11:31] value. These are the values. I have used rich cross validation and in this case rich cross validation and in this case the value comes out to be 204. Right? Then we train on it and these are my new coefficients. [04:11:47] my new coefficients. These are my new coefficients. Right? All values have reduced. If you compare with the original all values have reduced now. So the negative error relationship is with errors division W [04:12:04] they show the negative relationship. So the coefficients with highest value suggest cumulative career statistics total runs the total hits are most important and showing a positive relationship with salary. division_W the [04:12:21] only feature which is notably high negative coefficient which indicates that being in western division is associated with negative effect on the target variable. So do not be in the uh western division ears assist and new [04:12:37] league and these features are smaller. So are you understanding this is how you're practically supposed to run the code and write down the observations code and write down the observations over here [04:12:51] this case the values will not become zero but near to zero. zero but near to zero. Clear? [04:13:04] regularization? How it helps in preventing overfitting? Whereas L1 helps in feature selection and ridge helps in uh reducing uh the [04:13:16] and ridge helps in uh reducing uh the values if all are important. And of course this analysis helps in understanding the relative importance and influence of different aspects of a baseball player statistics on the [04:13:29] baseball player statistics on the predicted target. Clear? predicted target. Clear? Now moving on to the next concepts right model optimization. Now what does the term optimization mean? What is the [04:13:48] the term optimization mean? What is the meaning of model optimization? Optimizing resources. We want minimum number of uh code less number of time to giving the best. Yes. how the model performs efficiently. Absolutely. So if [04:14:06] we talk about the journey of data science or analysis, see every thing on if you if we are doing a normal Excel maximum, minimum drawing graph that's [04:14:18] also analysis and even when we are doing machine learning that's also analysis. But where is the difference? Are you able to get this um graph that I'm showing where is the difference? Can you tell me if we [04:14:35] already have a data set? So this is the side for that we are going to describe side for that we are going to describe you know about the uh historical data you know about the uh historical data that we have that is simple finding out [04:14:49] hindsight in the data. When we try to find out why that happened that becomes find out why that happened that becomes diagnostic analysis that is inside. diagnostic analysis that is inside. And if we move further, [04:15:03] what is going to be predicted sorry or what will be the output that becomes predictive analytics and similarly why that uh you know we [04:15:17] are predicting that output whether that is going to be good or bad that is known is going to be good or bad that is known as prescriptive analytics. So this is where we start from inside and now we want to move to the [04:15:29] and now we want to move to the foresight. How does it happen? When we foresight. How does it happen? When we start optimization of the model. Are you understanding this point? Very very important graph from one analysis [04:15:43] to another to explain you in more lame language. Suppose you have cold and cuff, right? And you go to the doctor. Doctor sees okay you know your nose is running. you have cold and cuff. He says you know um [04:15:58] medicine and you know do a lot of steaming and take this medicine you'll be fine. But that was his you know analysis which he done to just by seeing you. So then you come back you're not [04:16:11] fine then you go back after a week to him you say no I have fever also I have other thing then he thinks that the fever he's not getting well so maybe he has some kind of an infection then he prescribes you certain blood test so [04:16:24] based on the blood test he would give you certain medicine but you're still not fine your further analysis is needed maybe you have a chest infection now or something and maybe you want to go in for an MRI so that's how the level of [04:16:38] analysis increases is all right. So basically what do we want to do and what do we want to optimize in this whole machine learning process that if we have this input data that is the training data including the target output. Of [04:16:54] course if it is supervised learning we would also have the target output. This is my model right and output prediction calculated [04:17:06] right and output prediction calculated by model. So when we try to compare the predicted output with the actual output, therefore we are able to calculate the error and the loss function. Are you all [04:17:19] there with me? Then we are able to calculate the error and the loss function. And using this error and all loss And using this error and all loss function we can use optimization method [04:17:33] function we can use optimization method to sorry further reduce the values. to sorry further reduce the values. How we can further reduce its value. How we can further reduce its value. Getting my point right. So that's how we [04:17:48] are trying to improve on the error part of it. And then another way to optimize the model is through hyperparameter tuning versus model training. What is [04:18:02] hyperparameter tuning? That to find the best hyperparameters which give us the most efficient results. And this is a very very important slide because it gives the whole crux of the [04:18:20] data uh of all the points that we have studied till now. Please look here. This is my data set. Feature engineering is performed before. What are the different aspects of feature engineering? And why is feature [04:18:35] engineering required? Tell me what are the different aspects of feature engineering and why is it feature engineering and why is it required to prepare the data to fed into [04:18:49] the model not tune we will say to prepare the data prep-process the data to fed into the model and what are the different type what is this step known as splitting but what what are we splitting and why are we splitting it is [04:19:06] splitting of the data set. Why are we splitting it into three parts? Normally it is training and testing. This is this step is known as cross validation that [04:19:18] it divides the training data into training and validation and a separate training and validation and a separate test data. Second step is splitting of the data. Yes, but we generally keep using it for cross validation by taking [04:19:32] multiple samples. Got it? by taking multiple samples. Got it? by taking multiple samples. Now this is an iterative process. This keeps on getting repeated. So the training data we build models, train the [04:19:48] training data we build models, train the results, right? And finally once the training is done then on the validation uh thing we do the training results do hyperparameter tuning and keep on repeating this iterative [04:20:04] process. Finally the best model which is selected Finally the best model which is selected is then used for testing the output and compare the results. Do not get confused. See every concept [04:20:19] Do not get confused. See every concept that we study has its own role and concept right? So where will the regularization fit over here? If I compare the result my testing result is much higher than the training result [04:20:35] much higher than the training result then it is an overfitting model. Right? We always start with the basic mega then we will use the model over here as ridge [04:20:47] we will use the model over here as ridge or lasso and then do the comparison. Now getting my point everybody all right and the two types of hyper tuning methods are grid search and random search. Grid search is the one in [04:21:04] random search. Grid search is the one in which we it's an exhaustive re search technique where we have different hyperparameters one such as a b and c it's not necessary that you know we have understood that regularization as one [04:21:17] hyperparameter we have other models like decision tree which have many decision tree which have many hyperparameters hyperparameter to xyz hyperparameters hyperparameter to xyz value and then we try to take all the [04:21:29] combinations and then see which one works. works best for a model. Disadvantage of this method is it's an exhaustive method uh computationally and [04:21:41] timewise it is expensive. Clear and random search that randomly we will select any hyperparameters select any hyperparameters and do training on it. [04:21:56] Clear now and we have seen so cross validation is a very very important technique. Why? Because it helps to split the data into Because it helps to split the data into training, validation and test set. [04:22:11] Please make this point clear. Cross validation technique will almost be used in supervised learning because it will divide the data set into three parts [04:22:23] that is training, validation and testing set. Useful if you want to have a metric set. Useful if you want to have a metric on how well your model is performing. [04:22:35] Clear? So once the grid CV search returns with tuned parameters, build the model using this set with the tuned parameters and test the new model with the test set. Clear? And this is another uh you [04:22:50] set. Clear? And this is another uh you know um slide which I always uh show it to my learners even in my feature engineering class. Now are you able to interpret this particular slide everybody? [04:23:04] Yes. Right. So we again have a data set, retrieve the data set, perform pre-processing, wrangling, feature extraction, feature engineering, train the model, then model evaluation, [04:23:17] hyperparameter tuning and reiterate the process. But this slide gives a much better picture of each step. Got it learners? So getting back to the file. So this file is pretty long. You know [04:23:30] last week also we have done and I think so this session also we would be using. So model optimization means hyperparameter tuning is the process of finding the best settings for the parameters in the machine learning [04:23:46] model. Hyperparameters are settings that are not learned during the training but are set before the training process begins. Hyperparameter tuning involves [04:23:58] trying different combinations of the hyperparameters and evaluating the model's performance using validation techniques. So the two techniques that techniques. So the two techniques that we can use is grid search. [04:24:17] grid search? It systematically works through the multiple combinations of It performs an exhaustive search on a specified parameter grid. So how do we go about it? Please try to understand [04:24:31] the different steps. The first step is we define a parameter grid that is defined with the help of dictionaries because they can be more than one hyperparameters in the model. Combination evaluation. [04:24:46] The algorithm evaluates all possible combinations of these hyperparameters. Model training. For each combination, the model is trained, evaluated using cross validation. And then we have the optimal parameters combining yielding [04:25:03] the best performance. Highest accuracy is chosen as the optimal set. [04:25:19] defines the parameter distribution. Specify the distributions and ranges of Random sampling. Random sample combinations of hyperparameters from these distributions. Model training for each sample combination. And optimal [04:25:35] each sample combination. And optimal parameters. we are running the code. I might get a lot of error. It will [04:25:51] lot of error. It will let me run the x tree. validation. Here it we are creating a cross validation object with repeated [04:26:05] kffold function. Number of splits is equal to 10. Number of repeats is three. equal to 10. Number of repeats is three. Random state is equal to 1. Are you now there with me? Everybody other code else code is running then we create a grid or [04:26:20] a dictionary and then try to find out the value. So still it's giving error. Is it running for everybody else? Oh well I've tried to do a shortcut but well I've tried to do a shortcut but shortcuts never work in life. Value y [04:26:35] input consist of nan. Okay. Yeah. Are you getting it? I'm not running the code. So this point is clear. [04:26:48] it's coming out to be 0.9. Is it the same with everybody? So then we create and fit the ridge regression model to avoid training data [04:27:00] with the optimal alpha. So here we have the ridge function results into best parameters alpha. fit them through the training uh input as well as output and [04:27:12] find the mean square value that comes out to be 341 out to be 341 and R square comes out to be 0.38 which suggests a moderate fit the model captures some of the variability in the [04:27:27] data but not a large portion and now if I see the coefficient this is how the coefficient values are now let's move on to the next concept last concept by to the next concept last concept by creating pipelines. So skarn pipelines [04:27:42] are nothing but an automation of the model fitting and data transformation steps for training and data set. So what are we doing at the moment when we upload the data we separately use fit and transform do encoding do scala [04:27:57] standardization. Now we will try to implement it with the help of a single pipeline. So what do we do? If we have test data, test labels, we would perform [04:28:09] feature scaling standard scaler initially. Then if feature selection or extraction, dimensionality reduction is required, that can also be done and then the model data can be fed to the different models. [04:28:25] So how do we go about it? We are very clear in supervised learning. We have a training set and a test set with class label, right? So in the first step we will try to fit the training data that is missing in the test set because we do [04:28:41] is missing in the test set because we do want to avoid data leakage right the fit will ne fit function will never ever work for the test data [04:28:56] right and therefore we have the fit transform function over here transform function over here this is the pipeline right so over here this is the pipeline right so over here you know We um do the transformation [04:29:08] such as for for scaling, dimensionality reduction and then do the predictive model and over here it is fit and transform for the test training data but [04:29:21] for the test data it is only the transform function. Clear? And the transform function. Clear? And the predict will only be for the test data. Now let's see how do we practically implement it. So why sklearn pipelines? [04:29:36] Pipelines provide an organized approach to managing your data prep-processing, modeling the code. They combine the pre-processing and the modeling steps pre-processing and the modeling steps into a single streamlined process. [04:29:51] Cleaner code pipelines eliminate the need to manually manage uh training and validation data at each reprocessing step reducing the clutter and the [04:30:03] complexity. Fewer bugs by uh bundling steps together. So bugs are few. It gives a much more cleaner uh code and easier to productionize. that is they [04:30:15] simplify the transition from a prototype to a model to a scalable deployable to a model to a scalable deployable solution. Clear? So we have skarn is a [04:30:27] beautiful library. We just need to uh you know create a class or we do import the pipeline function and we will create an object of pipeline class whatever steps that we want to perform. Do we want to use any kind of memory or [04:30:43] want to use any kind of memory or verbose clear? So a pipeline is a sequence of data transformers that can include a final predictor also final output also. It lets you apply reprocessing steps to your data in order [04:30:59] reprocessing steps to your data in order and u optionally end with a predictor for modeling. Each intermediate step in the pipeline must have fit and transform the pipeline must have fit and transform methods while final step only needs the [04:31:13] fit function. Memory stands that you can use cache memory for these transformations to make the processing fast using the memory argument. The pipeline main goal is to combine [04:31:31] validated together and have their parameters adjusted. You can set parameters for any step by using it name followed by a underscore. You can set [04:31:43] any of the parameters by default followed by a underscore name and the parameter name and you can replace any steps estimator with another estimator steps estimator with another estimator or remove transformer by setting it or [04:31:58] pass through or none. So now over here they are using a different data set that is the one with ocean proximity that is housing one I think. So I gave you for [04:32:10] testing also. So it is little different from original because one of the parameter is categorical ocean pro uh you know object. So how do I check that? What are the different uh uh you know categories [04:32:27] of this particular function? How do I check that? So we have this categorical data. So what are the different encoding techniques? either we use one hot encoding that is better because there is no order in this so we'll not use label [04:32:42] encoding we'll use get dummies function to perform it value counts function is a function in python which gives the categories along with their frequency categories along with their frequency count very very important function [04:32:57] what is happening in this particular code of line number 46 come on learners why is getting the output value and then we are separating the X. Yes, we are [04:33:10] defining input and out. What is happening in the next step? Line number 48. What is happening in the next step? Splitting of the data. Very very Splitting of the data. Very very important. And then we see the info and [04:33:24] sum over here. Right? So where are the null values? Where are the null values? null values? Where are the null values? Total bedrooms has Total bedrooms has 62 null values. Data is huge. 14,000 [04:33:40] 448. Right? So definitely cleaning is required. Removing of the null values encoding is required. Right? So rather than now performing each of these steps [04:33:54] separately can I use the concept of skarn pipelines. First I will feed in my for training data perform feature scaling feature extraction and ML [04:34:06] scaling feature extraction and ML algorithms and then work on the test data getting my point learners. Okay. So what are the steps that we want to perform before building the model? Feature [04:34:19] engineering steps. So feature engineering part is used as skarn. Even the other parts can be combined as pipeline. First missing value treatment imputation 162 missing values in the total bedrooms [04:34:34] uh numeric data column. Then we have the dummy variable creation for categorical data. and finally standardization of the numerical value. So how do we go about [04:34:46] it? These are my different libraries that I will impute import and most that I will impute import and most importantly is the pipeline library? [04:35:09] column transformer that importing column transformer class to apply different pre-processing steps to different subsets of the feature. So column transformer has the capability that encoding is applied only on categorical [04:35:26] data not on numerical data. Getting my point? So it will only be applied on categorical. So let me explain this. [04:35:43] transformation happens. But if these are my features which is a combination obviously of numerical, categorical and others. So column transformer has the beauty that for numerical columns only scaling will happen. for categorical one [04:36:01] hot encoding and maybe others can simply pass through. It's not necessary that some operation needs to be performed on it and finally we can get a transform [04:36:13] data in very one go and it can be combined together. Got it? Column transformer concept is clear to everybody. [04:36:29] now when you read about column transformer, you will understand that it allows you to apply different pre-processing steps to different subsets of the features in your data set. This is particularly useful when [04:36:42] you have a mix of numerical and categorical data that require different types of pre-processing. Column transformer ensures that each column or group of columns gets appropriate transformation before combining the [04:36:57] results for further processing or modeling. [04:37:10] names of the numeric and object type variables differently. So how are we going to do? And now we have splitted the training and the testing data. And we are exclusively selecting the data types include object [04:37:25] exclude object and result. Are you understanding this code learners? understanding this code learners? So housing_cat will only have one since this data set has only one categorical data that is ocean uh proximity and the [04:37:42] data that is ocean uh proximity and the rest are all numerical in nature. Clear. Now the next step says to set up skarn pipelines for numeric variables we need to perform missing value imputation and [04:37:59] then standardization. So how do we go about it? Now we will create a pipeline about it? Now we will create a pipeline object using this pipeline function num pipeline. of a numerical data. We will import uh you know do the missing value [04:38:14] import uh you know do the missing value imputation and standard scaling. Is this point getting clear? Simon, are you there? Are you understanding this code there? Are you understanding this code now of creating pipelines? [04:38:34] What is imputation and what are we trying to achieve over here? Imputation is the concept of missing value and here we are filling the missing value with we are filling the missing value with the median value. Good. Good. Mega very [04:38:47] clearly it says that right and standardization is that we are trying to create the mean of the data zero and standard deviation equal to 1. [04:38:59] Then unified data processing with pipelines and column transformer. Now we need to create another pipeline. No. So we've created a numeric pipeline. Now we create a column transformer. The numeric pipeline and the other one categorical. [04:39:16] For categorical data we only need to perform only and only one function that is known as the one hot encoder function. Clear? So this ensures both numerical and categorical data are pre-processed [04:39:32] categorical data are pre-processed approximately within single unified work. Now got it. The beauty of column transformer that we have combined the numerical as well as categorical values. [04:39:47] numerical as well as categorical values. Okay. Then we do pre-processing fit transform x train. We're going to fit and transform. And now check the uh you know data set for training. Clear. Now do you see [04:40:04] Clear. Now do you see it has already done this? So right. So now we fit and transform. And now check the training data. This is my now check the training data. This is my training data. And now if I check my the [04:40:18] scaling has already happened. Negative values are there as well as the null values are there as well as the null values have been removed. Got it? Pipelines. Are you understanding? So what is our observation? Good to see [04:40:32] all missing values are treated. Numeric values are standardized. Option proximity ocean proximity is converted. Now the other way of creating model is again you can create a final pipeline for [04:40:47] pre-processing and model ridge over here we are creating ridge regression model straight away and then moving ahead. So have you printed that also the final have you printed that also the final pipeline also? Yeah. So final pipeline [04:41:01] gives you this that after applying the basic thing we are applying the ridge model right and then we define the grid of hyperparameters to search you can set parameters for any uh step. So over here this is my ridge [04:41:17] alpha the hyperparameter which will range from 0.1 to 2.1 and then I use my grid search CV along with cross validation estimator is my final [04:41:30] pipeline these are my grids that I have passed uh onto it scoring I'm using negative mean absolute error number of jobs is equal to minus1 and we fit and [04:41:42] jobs is equal to minus1 and we fit and get the output clear. So the mean absolute error is 49,000 and the alpha value is coming out to be 0.1. Then I check my test result. The test [04:41:56] result is slightly higher. So it's almost a good fit of the value now clear almost a good fit of the value now clear how you know and even in classification model we would be using a lot of pipelines to get the result. getting my [04:42:12] point? So finally today after skarn we come to So finally today after skarn we come to the end of regression analysis which is you know part of supervised learning predicting numerical value. Regression [04:42:27] analysis is an essential method for estimating examining and predicting variable relationships. In this lesson we type as we had discussed initially types of regression then the different error matrices the cross validation [04:42:44] technique the regularization technique the pipelines. Are we all now clear with this [04:42:57] particular file 3.2 finally we are ending with regression finally we are ending with regression today. [04:43:09] of regularization in uh machine learning model? Why do we want to do regularization? Yes, if we want to prevent overfitting. Which of the following is a characteristic of L2 regular [04:43:23] regularization? Squaring of the value. [04:43:41] learners. It is both B and because it adds the squared magnitude its ridge of coefficient and it reduces overfitting by reducing the model complexity. Next question. What effect does increasing the alpha parameter in lasso [04:43:57] increasing the alpha parameter in lasso regression have on the model lasso? Do we want to increase model complexity or decrease? Obviously when alpha increases the model complexity. What is the primary difference between L1 and L2 [04:44:13] regularization in terms of the effect on the model coefficients? Remember L2 means squaring of the value. Remember L2 means squaring of the value. L1 is absolute. So it's simple. L1 [04:44:26] L1 is absolute. So it's simple. L1 encourages sparity. What is sparity? Having value zero. So it makes most of the features equal to zero. Well, while L2 regularization does not uh encourages sparity. Now [04:44:42] clear. How does regularization help in improving the generalization ability of a machine learning model? You want to reduce the variance or the bias? Overfitting. Overfitting gets reduced by [04:45:00] variance. What is the risk with tuning hyperparameters using test data set? What will happen if we tune hyperparameters? What is the risk with tuning hyperparameters using test data set? The [04:45:16] model will overfit the test set and data leakage can happen. Which of the optimizers? They try to avoid overfitting because [04:45:29] that's the work of regularization. If searching among large number of hyperparameter, you should try values in a grid rather than random value so that you can carry out the search more systematically and [04:45:44] out the search more systematically and not rely on chance. True or false? We should carry out more systematic We should carry out more systematic searches or random searches [04:45:56] when searching a large number of hyperparameters. grid rather than random values. Why? Because then it would become an [04:46:09] exhaustive search, right? We should always try random values. I hope this knowledge check was helpful for everybody. Start with another part of supervised learning that is classification. Now tell me how will you [04:46:25] define classification? It says supervised learning used when the output is output is categorical. Right? Yeah. So this is supervised learning [04:46:40] technique when the output is categorical. So these are the independent input variable and the model classifies them as animals and fruits. So are you aware about different classification techniques learners? [04:46:57] Do you understand different classification techniques? Do we understand the concept of binary classification? What what can you interpret from this particular diagram that the uh you know [04:47:13] the model will classify the input or not input sorry the outputs into two categories. For example, if a mail is there, it will either put it into the spam folder or the inbox. So when there are two categories as output, that is [04:47:31] are two categories as output, that is known as binary classification. Getting my point learners? Can we have more than two classes as output? Can we have more than two classes as output? Yes, we can have you know when we want [04:47:46] to label different vegetables such as capsicum, carrots, tomato, radish, turnip, it could be endless number of classes and that is known as multiclass [04:47:59] classification. Are you all getting this point learners? That is known as multiclass classification learners. [04:48:12] classification learners. What is this? Suppose I have a image and the in the image there can be multiple labels. This image consists of a truck, boat, dog as well as a plane. So if we have more than one, it's not multiclass [04:48:29] have more than one, it's not multiclass but a same uh you know uh image or a same input. For example, a movie. A movie can be uh romantic, thriller, horror also, right? or comedy, thriller, horror. So the same item can have [04:48:45] multiple labels that is known as multilel classification. What do you observe over here? If the classes are not balanced, right? The the [04:48:57] images of trucks are around 60% plane only 25% boards 15 then there is a different techniques to deal with imbalance class data set. All right. So [04:49:10] this basic concept of different classes is clear. Now let's straight away dive is clear. Now let's straight away dive into how we will analyze the classification uh output right or what are the [04:49:24] different metrics to analyze the classification output. So what do you think? How will I check whether my model is classifying the males and the spam males correctly? The matrix which is used or the major [04:49:40] role that is being played is done by the confusion matrix. confusion matrix. Please try to understand. For example, for example, I'm taking a very simple example. There is a uh you know bag of [04:49:54] balls of red and green balls. Now if I tell you to classify them, it is good that in the red ball bag you put all the red balls and in the green ball bag you put all the green balls. Right? So that is true positive and true negative. So [04:50:09] the green balls which belong to the green bag are in the green bag and the red balls which belong to the red bag are in the red bag. But is there a possibility that a green bag a red ball [04:50:22] can be in a green bag or a green ball can be in a red bag? Is that possibility? Yes. machine can create those errors and that kind of errors are known as false negative and false positive. Right? When the machine does [04:50:40] not predict the correct output. [clears throat] So what is our aim as a model? We always want to increase the true negative and true positive to create the best results. Right? And we always want to decrease [04:50:55] the errors that is the false negative as well as the false positive. Getting my well as the false positive. Getting my point right? So a confusion matrix [04:51:07] please remember again it is very very important uh point when we always talk about classification that a confusion matrix is a n byn matrix used for [04:51:19] evaluating the performance and that is why statistics is said to be the strong foundation for machine learning concepts. Bang on mega appreciate this. Yeah. So confusion matrix is an n byn [04:51:36] matrix used for evaluating the performance of a classification model where n is the number of target classes. So it will always be a square matrix. A good model as we have understood will have high true positive and true [04:51:50] negative rates and whereas a low or a bad model will have uh you know high uh uh low true positive and true negative length. So what do we do when we have imbalanced data set? That [04:52:06] means the ratio of the classes is not same or equal. It's always better to use confusion metrics as your evaluation criteria for your machine learning model. So this is how you know you will get the [04:52:22] So this is how you know you will get the output in terms of uh your Python code. It will create this kind of a matrix uh and display for you. And I hope the terms true positive and true negative, false positive, false negative are [04:52:36] clear. True positive is predicted positive which are actually positive, false positive, predicted positive but are actually negative. Getting my point? And this is the exactly the same example I always use in [04:52:52] my hypothesis testing class also I use the same example. interesting example. We have the images of cat and dog right. So we have about [04:53:09] of cat and dog right. So we have about 20 images of cat and dog right. So for cat they were only seven images and dog they are around 13 images. Right? So I you know ran this particular images on [04:53:25] the model. So six images which were actually CAD were predicted positive by the model. So that becomes my true positive. 11 images which were not CAD [04:53:37] were actually predicted that they are not CAD. They come under the true negative category. So very good. But the dis the the errors were that the image the image was of the cat but it says that you are a dog. So one error over [04:53:54] here of false negative and similarly the image was of the dog but it says that you are a cat. So this is false positive. So the false negative errors are known as type two errors also and the false positive are known as type one [04:54:11] the false positive are known as type one error or also known as alpha and beta. If you remember from hypothesis testing same concepts same concepts absolutely same concepts are being used. [04:54:25] Are you understanding is this? So there are different metrics that can be evaluated from this confusion from this confusion metrics accuracy precision recall F1 score FPR and FNR all can be [04:54:43] calculated. So if I talk about accuracy, accuracy simply measures how often the classifier makes the correct prediction. It's the ratio between the number of correct predictions and the total number of predictions. So what does accuracy [04:54:59] mean? I'll take my true positive, true negative divided by the total number of output. That gives me accuracy. But what is the illusion over here? The illusion is the illusion over here? The illusion over here is that it is not valid for [04:55:14] imbalanced data set. Anu says why type one and type two. Why type one and type two? It is there a possibility that the classes might not be uh classified correctly. So type two and type one actually arise because I am [04:55:31] considering one of the statement as positive. Positive means that the image is of the cat and negative statement means that you are not a cat. Right? So false positive error means [04:55:46] Right? So false positive error means that actually it is not a that actually it is not a image of a cat but still says you are a cat. That is why false positive that it is not a cat and false negative because [04:56:00] it is not a dog. the same game we play between null but we'll be uh you know resolving it into different metrics. So one of them is accuracy the other one is precision. So the precision is defined as the ratio of [04:56:16] the total number of correctly classified positive classes divided by the total positive classes divided by the total number of one predicted. So precision is useful when please try to understand precision is a [04:56:31] useful metric in cases where the false positive is a higher concern than false positive is a higher concern than false negative. What do we mean by that? That false positive is a higher concern in precision. So what is more detering? [04:56:48] whether you know the male is not a spam but the model predicted as spam. Is that more important than a male which was spam and model predicted as not spam? [04:57:07] male is considered as spam then that is a serious issue. Therefore the precision needs to be high. So what is precision? It is true positive upon the true positive plus false positive. Getting my point? [04:57:23] All right. So the other metric that we can use is recall. Recall is classified can use is recall. Recall is classified as the ratio of the total number of correctly classified positive classes divide by the total number of positive [04:57:40] classes or out of all the positive classes how much we predicted correctly. classes how much we predicted correctly. So recall should be high and idly one. And again when where is the recall function mostly used? It is useful [04:57:55] metric in cases where false negative trumps false positive. Getting my point? So what would be the case where false negative is important [04:58:08] than false positive? For example, medical cases where it does not matter matter if we raise a false alarm but actual positive cases should alarm but actual positive cases should not go undetected. [04:58:23] Got my point? So this is the formula for recall the true positive divided by the true positive plus false negative. [04:58:35] Clear ma'am do we need to calculate all these formulas? No, you just need to understand the mathematical intuition. The Python programming will do it all for you. [04:58:47] Right? So some might say ma'am u sometimes you know the false negative error as well as the false positive error both are important right then how do we go about then the metric which is used to balance [04:59:03] it out is known as the f_sub_1 score. So the f1 score is a number between zero the f1 score is a number between zero and one and it is the harmonic mean of precision and recall. So the definition already has the formula. Okay. Harmonic [04:59:18] mean. We use harmonic mean because it is not sensitive to the extreme values. So F1 score is valid when we have imbalanced data set because it maintains [04:59:31] imbalanced data set because it maintains a balance between the pre your classifier. If your precision is low, the F_sub_1 is low and the recall [04:59:45] is low again and F_sub_1 score is low. So what do I mean by that? This is the harmonic mean and this is the formula. So F_sub_1 score is the harmonic mean and it should be ideally equal to one. So there will be cases [05:00:02] where there is no clear distinction whether to use precision or recall then whether to use precision or recall then we can use the F1 score. Got it? [05:00:15] when true positive and true negatives are more important. Accuracy is better metric for balanced data. False positive is much more important. Use precision. Whenever false negative is much more important than use recall. F1 score is [05:00:31] used when the false negative and false positives are important. Now clear this gives you more clarity when to use this gives you more clarity when to use what and where. [05:00:48] And how do we create this confusion matrix in Python? By simply importing skarn matrix import confusion matrix and running the confusion_matrix function. So this is my y test y predicted and this is how I get the [05:01:05] predicted and this is how I get the output. [05:01:18] right but the best part about it that simply by using skarnmetrics import simply by using skarnmetrics import classification report we can get the all the metrics in one particular output this is what we actually use [05:01:37] not see Simon, it's a very very relative question, right Simon? Obviously 100% is perfect. But depending on the model, the kind of problem that you are facing, if you're not getting accuracy more than that, then this is the best. But if you [05:01:53] that, then this is the best. But if you can get accuracy up to 92, 95 or 97% can get accuracy up to 92, 95 or 97% then of course 85% is not acceptable. whether to accept or to reject the project [05:02:09] or you want to make further improvements or optimizations onto it. All right. So this is a very very beautiful uh you know code that you know [05:02:22] for classification that we will have this classification report in which we get all the output pre precision recall f1 score in one tabular form. All right and this is what we are going to actually use for uh you know comparing [05:02:38] actually use for uh you know comparing the results getting my point. So again a quick recap accuracy, precision, recall, F1 score, specificity and different metrics can be calculated. But as I told you we do not use different metrics. We [05:02:54] will straight away uh straight away import classification report and get the output. Now another very very important factor Now another very very important factor how graphically we can judge the output. [05:03:10] how graphically we can judge the output. So in classification problem we have AU So in classification problem we have AU and ROC curve. What does AU stand for? AU stands for area under the curve and ROC stands for receiver operating [05:03:26] ROC stands for receiver operating characteristics. learn file. This is my material. I'll be sharing it. even the regression [05:03:40] material, classification material. Right? So another very important factor to understand classification is the graph A and ROC curve. AU stands for [05:03:55] graph A and ROC curve. AU stands for area under the curve and ROC stands for area under the curve and ROC stands for receiver operating characteristics. So AU ROC curve is a performance measurement for the classification [05:04:09] problems at various threshold settings. Now what do we mean by various threshold Now what do we mean by various threshold settings that we will try to change the you know the that whether the classes have been uh classified correctly. You [05:04:24] know the threshold value always ranges between 0 to one. So it's like a meter rating. Okay. Now this is a point or this is the boundary. Yes, these two classes have been classified correctly or the other two right. RO is a [05:04:39] probability curve. Please try to understand. ROC is nothing but a understand. ROC is nothing but a probability curve and AU represents the probability curve and AU represents the degree or the measure of separability [05:04:52] between the classes. So here we are trying to measure how well the two trying to measure how well the two classes have been classified classes have been classified correctly based on the threshold value. [05:05:07] It tells how much the model is capable of distinguish between distinguishing between the classes. Higher the area under the curve, the better the better [05:05:19] under the curve, the better the better the model is at predicting the model is at predicting zero classes as zero and one classes as zero classes as zero and one classes as one. Now getting my point. [05:05:31] What do we mean by a? It represents the degree or the measure of separability. better the model is at distinguishing between patients with diseases or no [05:05:47] diseases. So how does it work? Basically, so this graph is drawn between the false positive rate on the x-axis and true positive rate on the x-axis and true positive rate on the right hand axis. Right? And the this [05:06:03] graph is and this is area under the curve and this is the ROC the receiving operative sorry. So this is receiver operating characteristics and these are the [05:06:17] different formulas that we have just now seen. So sensitivity and specificity are inversely proportional to each other whereas TPR and FPR are proportional. [05:06:30] Okay. So sensitivity and specificity. Okay. Recall the name. Recall. Have you understood recall? Recall is nothing but the ratio between true positive and the [05:06:43] true positive and false negative. Okay. The recall is also known as the true positive rate and the other name is sensitivity. Now clear [05:07:01] true negativity uh divided by true negative plus false uh divided by true negative plus false positive. So ROC curve is the plot between the true positive rate and the false positive rate. That point is clear [05:07:15] across the all possible threshold. So what is this threshold? The threshold what is this threshold? The threshold value is always between 0 to one and in between you know uh it tells how well the two classes have been classified. [05:07:29] Right? and AOC is very clear that it is the area uh under the curve. So the more the area under the curve which is equal to one the better is the classification. [05:07:42] Now clear Anush. So ROC AOC curve is the area under the curve. It sums up how well a model can produce a relative scores to discriminate between the positive and the negative instances across all classification threshold and [05:07:58] the values will always lie between 0 to 1. So a is desirable for the following that why do we want to understand the area? Why why do we want to use this [05:08:12] area under the curve? Let me go through it and still you have errors. I'll explain them. First let's get the crux of it. I understand a lot of points are there. First let's get the crux of it. Area under the curve is scale invariant. [05:08:26] It measures how well predictions are ranked rather than the absolute value. Right? So it's scale invariant. Does it tells how well the predictions were ranked. It's not uh absolute value. It helps in comparison. Secondly, it is [05:08:42] classification threshold invariant measures the quality of the model's predictions irrespective of what classification threshold is chosen. So this is what I mean by my threshold generally it should be like 50 50% [05:09:00] generally 5050% that my true negatives have been classified as true negative green balls in green bag red balls in red bag and the total area under the [05:09:13] curve is one. So this is one of the ideal situations. This is one of the ideal situation that there is no overlap area under the curve is one and the model has ideal measure of separability. [05:09:28] Now moving on to the next point. If we see it practically if and if my If we see it practically if and if my threshold is 0.5 there would be some you know uh you know errors that you know the red ball going into the green bag [05:09:42] and the green ball going into the red bag but these should be less. And if my area under the curve is close to one around 9.8 data 7 then also it is said [05:09:57] to be a good classification problem. So AU.7 means that there are 70% chances that the model will be able to discriminate between negative and [05:10:09] positive classes. Now getting my point is everybody now are you understanding the meaning of threshold? Now if I decrease the threshold suppose if I make it to 0.2 what will [05:10:26] happen? My false negative error will become less but my false positive error will become more. So the threshold we have to try to always create a balance [05:10:38] have to try to always create a balance to get the minimum errors. to get the minimum errors. Okay. where the model was not able to differentiate between negative and [05:10:54] positive classes. One of the worst cases that area under the curve is 0.5. So we don't want the area under the curve to be equal to 0.5. So this is the worst [05:11:07] situation where area under the curve is approximately 0.5. The model has no discrimination capacity to distinguish between the positive [05:11:19] classes and the negative class. Now better and what could be the even worse situation that the true negatives have been classified as true positive and true negatives uh positives as negative and the area under the curve is zero. So [05:11:37] when a c is approximately equal to zero, the model is actually reciprocating the classes. It means the model is predicting a negative class as positive class. And do we want this situation? Never. We don't want this situation area [05:11:53] under the curve to be zero. Area under the curve to be close to 0.5. But close to one is a good option. That means the classes are able to [05:12:06] distinguish each other clearly. What is the role of threshold? Classification threshold in machine learning is a boundary. [05:12:19] At what boundary? As I told you, we will cut the two classes or cutff point used to assign a specific predicted class for each object. Any machine learning algorithm for classification gives output in probability format. That point [05:12:36] is also clear. So in the morning also I was teaching So in the morning also I was teaching the LLM. So LLM, chat, GPT is uh you come under the category of classification problems and they all [05:12:51] give the output in terms of probability. Okay. So a little bit of GI and LLM models. So, LLM, chart, GPT, your co-pilot, all the different models, they [05:13:03] are all come under the category of classification problem and generate the classification problem and generate the word based on probability. Okay, so that's the tip of the day today. So in order to assign a class to an instance [05:13:17] for binary classification we compare the probability value to the threshold and if the value is greater than that then the probabil it belongs to class one if it is less than that it belongs to zero. So the story of [05:13:32] threshold is like this that what is the cutff point that we can achieve to get the best distinguished classes that if you increase the threshold you move left on the curve and if you decrease you move right to the curve. [05:13:49] Okay let's get back to the file. So now we are on lesson number four. You can go ahead with 4.1. So there are a lot of algorithms that we [05:14:01] would be covering in this uh classification logistic regression, navebased classifier, KN&N, decision tree, support vector machines. I so today uh we will be doing touching on to the classification algorithms. Let's [05:14:17] see how many are we able to do. Let's logistic near live bias and KN and definitely we are doing three. Let's see if we if we to begin with decision tree if we if we to begin with decision tree also. Let's see. [05:14:33] of classification? How will you define it? Tell me what is the definition of classification? Yes learners. [05:14:47] questions so that you have the you know clarity of the concepts and then you are clarity of the concepts and then you are also well prepared for the interviews. Supervised machine learning algorithm using categorical data always say it it [05:15:02] is part of which learning technique and where the model is trained to predict where the model is trained to predict the class label the output is important when I say label the output of the given input data. So it is important that the [05:15:19] output is categorical. Getting my point? Please try to cover all the points in the definition. If you're going to just put three four words, it will never give you a complete answer. Getting my point? Okay. So, [05:15:35] answer. Getting my point? Okay. So, classification is a supervised machine learning technique used to predict the category of class of n observations category of class of n observations based on the training data. So, I think [05:15:49] based on the training data. So, I think so this definition is quite correct. So classification algorithms categorize data into categories. So do you think that uh in classification there is a role of this [05:16:04] regression? Do you think that can we use a straight Do you think that can we use a straight line to uh separate two classes? Let's see since linear regression is a simple function. Let's see can it be [05:16:19] used. But before that you know classification examples of classification in healthcare whether you are diabetic or not or your medical condition is there all come under the health care category finance banks [05:16:35] financial institutions our classification algorithm whether you are eligible for loan or not marketing classification aids in customer segmentation and target marketing. It helps business identify potential [05:16:49] customers and tailor marketing strategies by categorizing customers based on their behavior. Retail in retail classification algorithms are crucial for managing inventory forecasting demand and manufacturing. [05:17:05] Classification is essential for quality control, fall detection. So there is no domain where classification problems cannot be used. I hope you all are getting this point in retail, marketing, manufacturing, HR, everywhere. [05:17:22] manufacturing, HR, everywhere. Right learners. classification that we have understood? Can you tell me? [05:17:35] This is what we are going to try to solve. Rashan, can we use regression in classification problem? This is just we are just about to solve it. We're just about to solve it. Just just be patient. [05:17:49] Okay. Okay. So, what are the different types classification? So, when the model is able to classify only two classes for example, whether the male is fraud [05:18:04] for example, whether the male is fraud or not fraud or spam or not spam comes under binary classification already predefined. Yeah. to classify into predefined multiple classes whether it is a tomato or whether it is a vegetable [05:18:18] is a tomato or whether it is a vegetable it is a fruit or it is an animal or is it is a fruit or it is an animal or is it a domestic animal or pet so many classes can be used one to many no chakra pani that's not the correct way [05:18:32] to answer the model is capable of predicting more than one classes of the input that's that's the way to answer It multilel multilel classification [05:18:47] multilel classification be careful while answering this that each data point can be assigned multiple labels simultaneously rather than just one. Yes, simultaneously like in an image it can it can have an image of a [05:19:01] image it can it can have an image of a dog, cat, rat, a chair etc. Okay. And the last part is imbalanced classification. What is imbalanced classification? When the classes are there, they can be [05:19:16] two classes or more than two classes. But they are not equally balanced. Right? one class like 70% of males are there and you know 30% of females when the classes are do not have the same frequency count or we can say when one [05:19:32] frequency count or we can say when one class has higher significantly more than um other observations right when one becomes a majority class other one becomes a majority class other one becomes a minority class okay [05:19:45] so binary classification under binary classification some popular algo algorithms are used. So please try to understand binary classification can be achieved by logistic regression, knives bias, KN&N decision trees and support [05:19:59] vector machine. We are supposed to you know do all of these. So while these methods excel in bind capable of handling multiclass except [05:20:11] logistic regression or other are capable of handling multiclassification task also this versatility allows them to be used in wider range of applications such as recognizing multiple categories of [05:20:25] objects in images or predicting several types of behavior. So now we'll start types of behavior. So now we'll start with logistic regression. Before that [05:20:37] let me also let's do a recap. Now what are the different metrics involved in classification? How do I judge that my [05:20:49] classification problem is correct or not correct? Confusion matrix. Yes. So what are the different components of a confusion matrix? accuracy, precision, recall. These are the different metrics. But [05:21:04] components of confusion matrix I said components true positive, true negative, false positive, false negative. Absolutely correct. And based on these components, [05:21:16] we calculate different metrics which can be used to analyze the classification model. So if this is my data and I divide the data because classification [05:21:28] also comes under supervised learning technique. So if this is my data and I divide it into training and test set and the training set over here we have right from the training we develop the model. Test set is used for [05:21:45] model evaluation and we understand for classification problem the different classification problem the different metrics used are accuracy precision and recall right so it's a quick revision again so this is my n to n matrix if I [05:22:00] have two classes what is what are what are true positive and true negative so when the output is also true the predicted value is also true when the output is false when the predicted value is also false but what What are the two [05:22:14] types of errors? What are the two types of error that we What are the two types of error that we analyze in confusion matrix? Type one error is also known as alpha or the or the false positive [05:22:31] alpha or the or the false positive error. Right? The type one is also known error. Right? The type one is also known as the false positive or the alpha error. And type two error is also known as the beta errors. [05:22:47] Clear? Then which metric is to be used when we have accuracy. When out of the prediction model has been made what the percentage is. If it is an imbalance class is accuracy a good metric to be [05:23:02] class is accuracy a good metric to be used right. If we have more of positives right true positive right out of all the yes how many of them were correct then we use the precision recall and sensitivity are the same when we want to [05:23:16] answer the question how good the model was predicting at real yes events recall and specificity so there is this always this confusion about recall so [05:23:28] recall is associated with specificity also and sensitivity also also and sensitivity also The question can accuracy so F1 score so The question can accuracy so F1 score so F1 score is a more matured score which [05:23:42] can be used uh which is not biased towards the precision or recall and u it is used when deal dealing with imbalanced data set meaning that they are more of one class label than they are of the other. It corresponds to the [05:23:58] are of the other. It corresponds to the harmonic mean of precision and recall. Right? And what is the graph that we understood which helps us to identify whether the two classes have been uh classified correctly or not. If you [05:24:13] classified correctly or not. If you remember that graph learners remember that graph learners the AU it's it's AU and ROC curve. ROC stands for receiver operating characteristics. A stands for area under [05:24:29] the curve, right? So what was that graph telling me? This graph generates probability value instead of binary 0 and one. It should be used when your and one. It should be used when your data is set roughly balanced. The ROC is [05:24:45] imbalanced data set leads to incorrect interpretation. So ROC curves provide good overview of trade-off between the true positive and the false positive rate for binary classifier using different probability threshold. So the [05:24:59] value if the value is below 0.5 it's a poor classifier 0.5 random classifier but if it is 7 or greater than one we can set that it is a good classifier. [05:25:14] can set that it is a good classifier. Now clear. So this is how we understand the ROC and the AOC curve is between the false positive rate and the true false positive rate and the true positive rate. Right? And if you look at [05:25:27] the more closer look at the AU and the ROC curve, this is my X, this is my Yaxis. These this is my optimal threshold. Practically we are going to draw this graph also. and the selected threshold. The ROC curves gives a quick [05:25:44] visual understanding of the classifier's accuracy. The closer to the right angle curve, the more accurate the model is. So more closer it is to the right. Classification threshold that turns the up upper left corner of the curve [05:26:00] minimizing the difference between the true positive rate false positive rate true positive rate false positive rate is the optimal threshold. Clear? Okay. So now let's go ahead and [05:26:13] understand the first algorithm for classification logistic regression that is transform linear regression which gives output linear regression which gives output between zero and one. So there was this [05:26:26] between zero and one. So there was this uh query of you all can straight line be used for classification problem right that was the query. So when I say you know binary classification, how do you think that binary [05:26:41] classification would appear? Yeah. So I'm asking you how do you think that binary classification points would actually look [05:26:53] actually look in the graph? Tell me. Absolutely correct. Right. they can be on the either side of the straight line. Right? So if I say it's a typical good binary classification data, it's a [05:27:10] binary classification data, it's a typical binary classification data that either the output will be zero or the output this is this is an ideal situation agree the [05:27:24] will be zero or the y will be one. Do you all agree? good fit on these data points or is it covering is it consistent or is it [05:27:40] underfitting? No, this is not this is not a good fit. And moreover, the values are also getting lesser than zero. Some of the values and some of the values are greater than one that we don't want [05:27:54] because in binary classification the output will either be zero or be one in output will either be zero or be one in idle case. Agreed? So now I need to idle case. Agreed? So now I need to modify this straight line into a [05:28:09] modify this straight line into a sigmoidal curve or an S shaped curve. Is it helpful for me? First of all, the values will always be within the range between 0 to 1. [05:28:22] Is this point getting clear to everybody? And suppose if I want to find everybody? And suppose if I want to find out that my threshold value over here is out that my threshold value over here is 0.5. So if my value is 02 it it is very [05:28:36] clear that it belongs to this class and if the value is 7 it belongs to this class. Uh only the confusion would arise at 0.5. So it is up to me the 0.5 I want [05:28:48] to keep it on this class or that class. Clear? Yes, linear regression that question that regression are incapable of solving [05:29:01] that regression are incapable of solving binary classification problem. Why? Because the straight line gives under fit to all data points. The you know the values can exceed greater than one or zero. Right? So are all these points [05:29:16] zero. Right? So are all these points getting clear to everybody? mathematically? So we understand this sigmoidal function [05:29:29] So we understand this sigmoidal function equation is written like this right? P equation is written like this right? P is the probability 1 upon 1 + e to the is the probability 1 upon 1 + e to the power minus y where y is nothing but the [05:29:43] why can you tell me it's like the linear polomial not even polomial multiple linear regression equation. Agreed learners? So this is your linear regression. This is your sigmoidal function. [05:30:05] So now if I start taking log on both the sides. So now on the left hand side this is the equation that log to the power p is nothing but the probability. We want the answer in terms of probability that the values will lie between 0 to 1 that [05:30:21] the values will lie between 0 to 1 that is p upon 1 minus p beta kn beta 1 x1 beta_2 x2. So again ma'am do we are we need to solving the coefficients over here. Yes. Logistic regression is a modification of linear regression. But [05:30:37] here also we are trying to find out these regression coefficients. Got my these regression coefficients. Got my point? Is this point getting clear to everybody? [05:30:59] So taking this point further log px upon minus px is known as the log it function. Now getting my point where does this term logistic term comes from? does this term logistic term comes from? because of this equation and simple px 1 [05:31:12] because of this equation and simple px 1 minus px is known as odds of p that's the technical term. So now you are clear that the straight line or linear that the straight line or linear regression is not capable of solving the [05:31:26] logistic regression problem. So the straight line can gets converted into the sigmoidal scurve because the predicted y lies within the range zero predicted y lies within the range zero and one. I hope this is clear to [05:31:41] and one. I hope this is clear to everybody. success 1 minus P is the failure of the probability right? Same we are using absolutely the concept remains the same. Yes. [05:32:05] learning sorry what is logistic regression? Please try to understand and try to answer all the points. This is what the interviewer will catch. How what the interviewer will catch. How many concepts you have grasped. Okay. So [05:32:17] logistic regression is a supervised uh machine learning technique primarily used for binary classification. In this method we apply the sigmoidal function to the linear combination of [05:32:33] independent variables and predictors. Now clear always say that it is a part of supervised learning classification problem used for binary classification. And here we are applying the sigmoidal function to predict uh you using the [05:32:49] linear combination of independent uh variables predictors of features. What are what does this mean? What does that mean? Independent variable predictors of features. This refers to the input part of it. [05:33:05] Right? Absolutely correct. Absolutely correct. Great. between 0 and one. And this probability represents the likelihood of a data [05:33:20] point belonging to a specific class positive or negative outcome. positive or negative outcome. Clear? [05:33:34] regression it uses the sigmoidal function. The core is the sigmoidal function. The core is the sigmoidal function sigma zs. So over here logistic regression is also known as sigmoidal function. This is how it will [05:33:48] always range between 0 to 1 1 + e ^ minus zed where zed is nothing but the minus zed where zed is nothing but the linear combination of input [05:34:03] multiple variable linear regression and the output is in terms of probability ranging between 0 to 1. Are these points 1 2 3 getting clear to everybody? So now you're getting the crux of machine learning. You have to first [05:34:18] understand the mathematical intuition of that algorithm or model. Then how do we implement it in Python? Understand it from three perspective as I always tell [05:34:30] in my data science class also. First try to understand the concept the mathematical intuition behind it. Second, how do we implement it in Python? And third, how are we going to interpret the output? All these things [05:34:45] interpret the output? All these things need to be taken into account. Got it? [05:35:02] What is cost function? The error function. How are we calculating the error function in logistic regression? That is calculated with the help of binary cross entropy. Please try to understand the cost function used in [05:35:19] logistic regression is binary cross entropy which measures the discre dis differences between the predicted probability and the actual class label [05:35:33] where m is the number of the training samples y is the true label ycap is the predicted probability. Clear? [05:35:55] learners do we need to optimize our models also? See the concept same concepts will be applied. So what is optimization? Why do we need to optimize optimization? Why do we need to optimize the model? [05:36:17] optimize? What is optimization? Extracting are the best, right? Getting the best. So the goal is to find weights W and bias B that minimize the cost W and bias B that minimize the cost function [05:36:36] gradient descent. So we'll try doing the gradient descent algorithm as we move ahead with further algorithm mostly in onsemble learning I've covered. Yeah. So basically we want the best output from the model. So now we are [05:36:53] going to use the breast cancer data set. Do I need to share this with you all? [05:37:05] So this breast cancer Wisconsin diagnostic data set is widely used data set in the field of machine learning particularly in classification problems related to medical diagnosis. This data sets consist of breast cancer cases [05:37:21] derived from a group of patients who underwent surgery and had their breast mass tissue sample whether they are cancerous or not. And basically there are two types of cancer. One is malignant and the other one is benign. [05:37:35] Malignant malignant I don't know how I really pronounce it correct or not and really pronounce it correct or not and the other one is benign. Right? the other one is benign. Right? And this data set contains 569 instances [05:37:48] each representing an individual sample of the breast tissue. Number of attributes the they are 30 numeric attributes computed from the digitized attributes computed from the digitized images of the tissue sample. Thements [05:38:03] of the cell nuclei present in the images. So based on that we are there. So what are the different attributes or inputs in this data set? Radius, mean of the distance from the center to point. So [05:38:17] it's not an image just they have taken uh different parameters which will help us to uh decide whether it's cancerous or not. Textured standard deviation of or not. Textured standard deviation of the grayscale parameter area smoothness [05:38:31] compactness concavity concave point symmetry fractal fractal dimension. And the target variable diagnosis indicates the cancer type diagnosis indicates the cancer type which can be malignant or benign. So how [05:38:45] many what is the output diagnosis and what are the two outputs of the uh diagnosis whether the cancer is malignant or benign. Now clear is the [05:38:58] objective. Please always try to give time on the data set without that you will not be able to understand the analysis what you are trying to do and what you are trying to achieve. So we will use this data set to explore and [05:39:15] compare various binary classification algorithm examining how their performance varies depending on the type implementation of each algorithm accompanied by the mathematical explanations. [05:39:31] Yes learners are we good to go? Why do we do regularization? First let's understand that is if regularization is required then only we will do anush why required then only we will do anush why do we do regularization tell me [05:39:46] yes if observed overfitting if the model will overfit then we will okay anush yes now clear so now let's start implementing so what are the steps uh in [05:39:59] python come on tell me quickly learners what are the steps in python that we what are the steps in python that we need to perform tell me quickly Okay, importing of the libraries. I hope the basic libraries are clear to everybody. [05:40:12] Then these are for the pre-processing model scaler pipelines. Have we understood the concept of pipelines? Can we implement it in classification? Yes. What is the concept of pipelines? EDA. What does EDA mean over here? No, [05:40:27] we don't. Yeah, we start with EDA. But first is import of libraries. I'm starting from very basic assignment. So what is the use of pipelines? Tell me learners. Yesterday we did skarn pipelines [05:40:52] deployment after package. What do you mean by that? automating workflow for building ML models. Good AJ, that's a better way. [05:41:08] It's automating because we would do all the pre-processing in one workflow. That's the correct way of answering it. Okay. An skarn linear model logistic regression and accuracy. My skarn version is 1.5.1. [05:41:25] version is 1.5.1. What about you learners? cancer data set everybody with ID, diagnosis, radius mean, texture mean? [05:41:40] diagnosis, radius mean, texture mean? Oh, so Subra yours is higher than mine. Oh, so Subra yours is higher than mine. Right. So what do we observe in the info Right. So what do we observe in the info of this data set? Tell me. [05:41:57] data set? Tell me where is my output? This is my output. And of course it has to be categorical as malignant or benign. And all others [05:42:10] as malignant or benign. And all others are numerical values. [05:42:22] So do we need to convert the output categorical value into numerical? categorical value into numerical? Do we need to encode the output? No. AJ, Do we need to encode the output? No. AJ, we don't need to encode the output. [05:42:41] Yes, we definitely need to because computer will not understand benign or malignant. We all the categorical value need to get converted into numerical need to get converted into numerical value. Okay. [05:43:12] This is the radius mean, texture mean, the per par meter mean. Do you see this? The count, the average value, standard deviation, minimum and maximum. So you have to uh you know observe whether the values are lying within the range or are [05:43:27] values are lying within the range or are they too deviated right so this is if I do df.escribe describe T then I get the values for all other [05:43:40] then I get the values for all other parameters. [05:44:02] complete null values. So can we drop off this column? [05:44:18] 1. It is for the column in place equal to true. We understand that we that means it will make the changes original in the data set. So now we will remove [05:44:30] this column because it is of no relevance. Got it. [05:44:58] So when I do a df dot shape it has got this particular data set has got 569 this particular data set has got 569 um rows and 32 columns. Agreed? So this particular data set has 569 rows and 32 column. And how and what is [05:45:16] the best way to check categorical data by using unique function which gives me the different categories. N unique gives me the number of categories but I always me the number of categories but I always prefer value counts function that means [05:45:32] it gives me the frequency along with the category. So they are 357 benign category. So they are 357 benign cancerous patients and maling 212 maling [05:45:45] cancerous patients and maling 212 maling patients. Got it nas? of it I get this type. Everybody is [05:45:58] of it I get this type. Everybody is getting this point. Learners, [05:46:15] world, always try to you know create graphs. Write down the observation what graphs. Write down the observation what the graph is doing. Right? Then again which which ID has been dropped? Which column has been [05:46:30] dropped? The ID column has been dropped. And this is my original data frame. Got And this is my original data frame. Got it? Now [05:46:47] understand the relationship between the features. The heat map visualizes the correlation between different in the data set. So what does the heat map do [05:46:59] or what is correlation? Yesterday we discussed about this. We've discussed this point. What are the values of correlation? What does the correlation do? What is the range of values of [05:47:13] What is the range of values of correlation? uh relationship which will always range between minus1 to + one. Did I share that with you or not? So what are the values of correlation? [05:47:35] Negative correlation and no correlation. Yes. Good. Good. So you all remember other learners? Great. So heat map is the map which gives me So heat map is the map which gives me relationship between the various [05:47:49] dimensions. So what do we observe that there is high correlation value between radius mean parameter mean and area mean. Whereas when we talk about other [05:48:01] observations, feature groups, features related to worst, largest value of these features for being each image, mean [snorts] and error, standard error [snorts] and error, standard error calculations are grouped over here. [05:48:14] Now tell me what is happening in the next stage. What is happening over here? That from skarn pre-processing import label encoder. Right? Label encoder is one type categorical encoding. Please [05:48:32] try to understand the correct word. This is categorically encoding and we are is categorically encoding and we are using label encoder right and only one [05:48:44] one column is categorical that is diagnosis and here I am trying to fit and transform only the diagnosis column. So now my benign and malignant values [05:48:58] So now my benign and malignant values change into zero and one clear. Is this point getting clear? So now when I see my DF dotted I see now not M and [05:49:12] B. Earlier please look at the output the column of uh it is in terms of M and B. column of uh it is in terms of M and B. Now I can see it in terms of [05:49:24] zero. Can we define what is one and what is zero? How else? Yeah. So when one is malignant and zero is benign. Yes, we malignant and zero is benign. Yes, we can. Now what is happening over here? [05:49:37] Yeah. So basically we are trying to separate the input and output and then separate the input and output and then split the training and the testing data. Right. And after the splitting only then we move to skarn pipeline uh pipeline to [05:49:55] streamline the pre-processing and the training process. Why do we do this? Why do we do this after splitting the training and the testing data that the features in the training and the test sets are standardized to have a mean of [05:50:09] sets are standardized to have a mean of zero and standard deviation of one. A logistic regression model is trained using standardized training data. to using standardized training data. to avoid data leakage. We don't want our [05:50:21] avoid data leakage. We don't want our testing data to be part of the training data otherwise there would no point uh be of the prediction. Got it? Right. So do we understand this pipeline function? What are the two things that [05:50:37] we are doing in this pipeline function? Learners, we are creating a pipeline. First step we are doing. What is standard_cala do? Come on learners tell me it's not that difficult. We've all done the EDA [05:50:52] part of it. What is standard scaler do? Feature scaling. Yes. Standardization. Feature scaling. Yes. Standardization. And then we are applying model selection logistic regression which will be iterated 10,000 time and randomly the [05:51:06] iterated 10,000 time and randomly the data 42 is getting selected. So till here are we clear? We have splitted the data right into training and testing. Now after training and testing split and even the categorical encoding has been [05:51:20] done I am doing standard scalar function. What is standard scalar function. What is standard scalar function that all my data points will have will be standardized with mean equal to standard deviation equal to 1. [05:51:37] Right? They will all be having that value mean equal to zero and standard deviation equal to 1. And next step we will apply logistic regression on the [05:51:50] will apply logistic regression on the data. Now better Simon now clear. And finally we go in for the pipeline dofit. What does pipeline dofit do? It will try to find out the parameters or the training data happens on input as well [05:52:07] training data happens on input as well as output training data till here. So now when I do fit it will automatically first do standardiz standardization of all the input points and then fit the model and then we will predict the [05:52:23] model and then we will predict the output using training and testing data. Clear? But since it is a binary classification [05:52:35] problem, we will do not use simple predict function. We use predict probability a. What does that mean? Probable outputs are needed that will [05:52:47] Probable outputs are needed that will range between 0 to 1. So this is my range between 0 to 1. So this is my output. Obtain the predicted output output. Obtain the predicted output after training the model. [05:53:07] my actual label. This is my predicted label and this is the probability associated with it. Are you understanding the three aspects Are you understanding the three aspects of it? So when I do the probable A, it [05:53:21] of it? So when I do the probable A, it gives me the values in binary. Yeah, the values will be 0 and one. The values will be 0 and one. But will it be zero or one will depend on the probability value. Got it AJ? This is what I was [05:53:37] saying when I say about the LLM models, right? The probability decides what is going to be the output and then we can calculate the accuracy [05:53:49] score and the training score. So in this case both of them are coming out to be quite same. So we can say it's overall coming out to be a good fit. But but which are the metrics which are actually used for classification problem? Which [05:54:05] are the metrics which are actually used for classification problem? the confusion matrix because it's a little imbalanced not highly imbalanced because if two classes were absolutely equal then it would have [05:54:21] been balanced data 50/50 but it is somewhere around 6040 I would say okay p value counts or this graph is telling me it's imbalanced right so the training [05:54:34] and the testing accuracy is there but ultimately who is the judge the confusion matrix. The components of confusion matrix are true positive, true negative, false positive, false negative. I hope these points are very [05:54:48] much clear to you. So the provided confusion matrix provides a summary of the effectiveness of COVID 19 tests that identify [05:55:00] individuals with with the virus and those without it. So we have understood this concepts that the K is correctly classified as positive and predicted also as positive. These terms are very clear. We've discussed it in detail [05:55:14] yesterday and I had done a recap. So now significance of confusion metrics is in terms of these metrics such as accuracy, precision, recall which is sensitivity, precision, recall which is sensitivity, true positive rate, specificity and the [05:55:29] true positive rate, specificity and the F1 score. Right learners and do we understand the formula that's also we've gone in detail. [05:55:42] Then we also understand the ROC curve or the AU curve which is used as a graphical representation of a classification model's performance [05:55:54] across different classification thresholds. It plots the true positive rate against the false positive rate at various threshold settings. So true positive rate or recall at these settings. Clear? So what is our main aim [05:56:12] settings. Clear? So what is our main aim in uh you know plotting the AU or the ROC curve that the default threshold for many classification algorithm is 0.5 meaning that if the predicted probability is greater than.5 the [05:56:28] probability is greater than.5 the instance is classified as positive. However this threshold might not always be optimal especially in the case of skewed classes. Do we understand the term skewed? What do we mean by skewed [05:56:42] class distribution learners? If the data is not normally balance, it is not known to be skewed data. Okay. So the methods of finding optimal threshold, there are [05:56:54] two ways. One is maximizing your Jordan's J statistics. In this we will uh find out the value of J by subtracting the true positive rate with the false positive rate. The optimal threshold is where the statistics is [05:57:10] maximized. This method balances the TPR and the FPR aiming to maximize the true positive while minimizing the false positive. Closest point to 0 comma 1. Another method is to point on ROC curve that is [05:57:26] closest to the top left corner representing ideal classifier. Clear? an area under the curve also we are clear. Now getting back to the Python code. So again confusion matrix [05:57:44] code. So again confusion matrix classification report ROC curve all these are part of which library? They are part of which library learners They are part of which library learners learn metrics be very clear skore [05:57:59] uh sorry skarn.metrics even the RSS MSE they were all under these curves. here. So the confusion matrix is uh is [05:58:12] created simply by calling the confusion matrix by passing the uh actual value matrix by passing the uh actual value predicted value output and based on that predicted value output and based on that I get this display [05:58:26] what do I observe 70 are my true positive 41 are my true negative and if I talk about type one and type two error it's quite less so logistic regression it's quite less so logistic regression is done well it gives me two and one [05:58:40] is done well it gives me two and one errors clear and then the best way to go about is creating a logistic regression report. So over here the precision is 97 report. So over here the precision is 97 98 for 0 and 1 recall F1 score is 98 and [05:58:55] 96 and if I look at the overall accuracy overall accuracy is also coming out to overall accuracy is also coming out to be 97%. So overall this particular logistic regression has come out to be a good classifier for this particular data [05:59:10] set. Clear? So over here they've given in detail class 0 negative class precision is.97 97% of instances predicted as is.97 97% of instances predicted as class 0 are actually zero. 99% of actual [05:59:25] class 0 instances are correctly predicted as class 0. And here we move predicted as class 0. And here we move ahead. [05:59:37] Right? All of the detail is given over here. Okay. And finally, how do I do the judgment that I need to do plotting of the ROC AU curve which shows the trade-off between the positive rate and the false positive rate at various [05:59:53] threshold settings. The AU value indicates the model's ability to discriminate between the positive and the negative classes. So parameters used in plotting are these many parameters that first of all ROC curve is the [06:00:11] function which gives me the different values of FPR, DPR and threshold right and based on that I'm going to calculate my Yordan's J and find out my optimal [06:00:23] threshold value. So what do I observe that optimal threshold value is coming that optimal threshold value is coming out to be 0.4848 48 48 and ROC or the area under the curve as one. So is it a good classifier? [06:00:35] Area under the curve is nearing it's almost one. It is one not even almost almost one. It is one not even almost one. So it is a very good classifier [06:00:47] one. So it is a very good classifier getting my point learners. So the ROC curve in the image reaches the top left corner TPR equal to 1, FPR equal to zero which indicates a perfect classification performance. The model [06:01:02] perfectly distinguishes between positive and negative classes at various threshold settings. And the ROC curve for logistic And the ROC curve for logistic regression model has AU1. This model has [06:01:16] perfect discriminatory power and AU of one means model correctly classifies all positive and negative instances without error and the optimal threshold of [06:01:28] 0.4867 defines the decision boundary for the classifier. Probabilities above this value indicate stronger belief that an instance belongs to the positive class whereas the probability below this value [06:01:42] indicate stronger belief that it belongs to the negative class. Clear? You know this practical example has made all the points clear. [06:01:56] But over here are we doing any kind of sigmoidal function anything that's already part of the implementation. So again I'm repeating this point. First try to understand the mathematical intuition behind the algorithm or the [06:02:10] model. Secondly how we will implement in Python and thirdly how do we interpret the results? Is it a good model or not? Clear? AJ says ma'am one doubt despite I [06:02:22] seen that this model is perfect in matrix itself. What is the main purpose matrix itself. What is the main purpose of saying ROC curve also see ROC curve tells me whether the two classes have been classifi are they well separated [06:02:34] from each other or not that is the idea is there more that is the idea is there more overlapping or not so just to reconfirm things it's not it's not by just doing one test you know if you are suffering [06:02:48] from some problem the doctor wants to reconfirm from different test maybe the blood test MRI report and X-rays or other things something like so it's something which is more visually appealing tells us has more impact yes [06:03:03] appealing tells us has more impact yes this is this is a good model accur the confusion matrix also tells us the different classification report also different classification report also tells us and AOC curves also tells us so [06:03:15] they there shouldn't be any doubt left you know when we are doing our analysis all the analysis should give the same result right [06:03:27] so now do we understand this graph let's see so now this is how is also this is also one of the reason the threshold is used that if I take it as 0.5 any value coming above it as I told you would belong to this class you know so [06:03:43] distinctly it is there is no overlapping it distinctly classifies all the points and any value which is lesser than 0.5 belongs to this class right so can we go in for a quick knowledge check about logistic regression and classification. [06:04:00] So question number one, which of the following metrics are used to evaluate following metrics are used to evaluate classification models? Confusion metrics ROC and F1 score is calculated using confusion. [06:04:14] What is a classifier? So when we talk about classifier, the output has to be categorical. So which one gives the output as category? It's both A and B. No. Output has a single discrete value or [06:04:28] output as a single discrete value. So both A and B. Input can be anything continuous or discrete. Yes, the answer is C. Next, false negative. What does false negative represent? Which option is [06:04:43] correct? It's A. Predictive ne negatives that are actually positives. Yes. In binary logistic regression definitely it is C. The dependent variable consist of [06:04:56] is C. The dependent variable consist of category. Next question. Why is linear category. Next question. Why is linear regression model output a poor predictor of probability? Why? Because D is the correct answer. It can give you know it [06:05:09] correct answer. It can give you know it cannot the range can be beyond zero and one. Next. The output in logistic regression problem is yes equivalent to one or true what is the pro what is its possible [06:05:24] value? So of course it is B it would depend on the threshold value. Okay. depend on the threshold value. Okay. So now we have knives base classifier. [06:05:36] Okay. Now moving on to the next algorithm. The next next algorithm is algorithm. The next next algorithm is completely based on probability. So what is probability learners? Chances of occurring of an event. And what is the [06:05:49] formula? Mathematical formula of probability. The number of successful outcomes upon the total number of outcomes. Right? What is the probability of getting ahead? When I when I throw a coin, it's 1 upon 2.5. [06:06:06] When I throw a coin, what is the probability of getting head? What is the probability of getting one? when I throw a dice that is 1 by 6. But there is another concept which is associated which is known as the conditional [06:06:21] probability. Do we understand what is the probability of event A given the probability of event A given probability of event B? Do we understand conditional probability? [06:06:35] This particular algorithm is based on that basian theorem. That's why data science is important. So learners who have done data science should should be clear about the basian theorem. This [06:06:51] conditional probability or the posterior probability is equal to the likelihood function that's opposite. What is the probability of event B given A into the [06:07:05] probability of event B given A into the prior probability A divided by the marginal probability? Do you remember this point or not? this point or not? No. [06:07:26] is directly proportional to the likelihood function. probability. Okay. So the nave base classifier this [06:07:40] is the formula that if I have these random figures or you know shapes the classifier is capable of classifying them into different shapes or groups [06:07:55] them into different shapes or groups right and the basic formula lies in that the posterior probability is directly proportional to the likelihood one function into the class prior probability divided by the predictor [06:08:09] probability divided by the predictor prior probability. Right? So how is this uh you know posterior probability actually calculated? It is nothing but the product of the likelihood function. Do you say see this X1 my