paudelsamir/365DaysOfData

A year long journey with ai from data, exploring adjacent techs

Jupyter Notebook

50

960 commits

updated Dec 26, 2025

See the code

README

Last Updated Repo Size

365DaysOfData cover

[!note]
I'll share progress and demos on linkedin and twitter.
I won’t post daily or raw learns, updates will be for specific topics, concise, and focused on what i actually built or explored. plan is 4–5 posts per week.
This journey is about AI from scratch with data, not my entire learning history. I’ll keep building in public while also learning other adjacent techs beyond AI.

Projects Completed

ProjectsDescriptionDeployment
Football Players Market Value PredictionA 10-day end-to-end machine learning capstone project involving data scraping, cleaning, feature engineering, model training, and deployment. Achieved 94% accuracy using gradient boosting algorithms.Live Demo 👆🏽
Movie Recommender SystemAn end-to-end content-based movie recommender system leveraging a dataset of 5000 movies from Kaggle. Built with cosine similarity and TF-IDF vectorization.Live Demo 👆🏽
Cat vs Dog ClassifierA deep learning model leveraging VGG16 architecture, trained on an RTX 3050 Ti for 30 epochs, achieving 95% accuracy using the Kaggle Dogs vs Cats dataset.Live Demo 👆🏽
Guess The Footballer By EyesAn interactive game where users compete against AI to recognize 25 famous footballers by their eyes alone. Built with ResNet18 achieving ~70% accuracy. Features scoring system and streak tracking.Demo 👆🏽
Seq2Seq ChatbotA sequence-to-sequence chatbot trained on Cornell Movie-Dialogs Corpus using encoder-decoder architecture with Luong attention mechanism. Built from scratch in PyTorch.Live Demo 👆🏽
GPT from ScratchComplete implementation of GPT transformer architecture from scratch following Karpathy's tutorial. Includes bigram model, self-attention, multi-head attention, and complete transformer blocks.Notebook 📓
Image CaptioningAn end-to-end image captioning project using the Flickr8k dataset. Explored the "Show, Attend & Tell" paper, built vocabulary, extracted features with ResNet-18, and trained a transformer decoder. Achieved a BLEU-4 score of 0.18 and deployed a Streamlit demo app.Live Demo 👆🏽
cineRank - A movie ranker appCommunity-driven movie leaderboard app with trending picks, sentiment reviews, and personal watchlists. Built using IMDb reviews, advanced text cleaning, EDA, vectorization (BoW, TF-IDF, GloVe, BERT), and BERT fine-tuning for sentiment classification. Features leaderboard, watchlists, and real-time updates.Live Demo 👆🏽
Choose Your Own AdventureInspired by interactive fiction like AI Dungeon, this app lets you become the protagonist in a personalized adventure story. Enter any theme—haunted mansions, space exploration, and more—and AI generates a unique branching narrative with multiple paths and endings. Features include an interactive visual map, concise story nodes (~40 words), meaningful choices, and a clean black-and-white interface for all devices. Explore different decision paths and control your own dynamic storytelling experience.Project Demo 👆🏽
Projects-Based-GenAIHands-on GenAI projects including text generation, multimodal models, and advanced LLM fine-tuning. Explore practical implementations of state-of-the-art generative AI techniques.Project folder
Project-Based-AgenticAIApplied agentic AI projects focusing on autonomous agents, multi-agent systems, and real-world agentic workflows using LangGraph and LangChain.Project folder

Resources

Progress

DaysDateTopicsResources
Day12024‑12‑14Basics of Linear Algebra3blue1brown
Day22024-12-15Decomposition, Derivation, Integration, and Gradient Descent3blue1brown
Day32024-12-16Supervised Learning, Regression and classificationMachine Learning Specialization
Day42024-12-17Unsupervised Learning: Clustering and dimensionality reductionMachine Learning Specialization
Day52024-12-18Univariate linear RegressionMachine Learning Specialization
Day62024-12-19Cost FunctionsMachine Learning Specialization
Day72024-12-20Gradient DescentCampusX, Machine Learning Specialization
Day82024-12-21Effect of learning Rate, Cost function and Data on GDCampusX, Machine Learning Specialization
Day92024-12-22Linear Regression with multiple features, VectorizationMachine Learning Specialization
Day102024-12-23Feature Scaling, Visualization of Multiple Regression and Polynomial RegressionMachine Learning Specialization
Day112024-12-24Feature Engineering, Polynomial RegressionMachine Learning Specialization
Day122024-12-25Scikit-Learn revision, Linear Regression using Scikit LearnMachine Learning Specialization
Day132024-12-26LR lab, ClassificationMachine Learning Specialization
Day142024-12-27Logistic Regression, Sigmoid FunctionMachine Learning Specialization , CampusX
Day152024-12-28Decision Boundary, Cost FunctionMachine Learning Specialization , CampusX
Day162024-12-29Gradient Descent for logical regressionMachine Learning Specialization , CampusX
Day172024-12-30Underfitting, Overfitting, Regularization Polynomial Features, HyperparametersMachine Learning Specialization
Day182024-12-31Neurons, Neural Netowrk, Forward PropagationMachine Learning Specialization
Day192025-01-01Forward Propagation, Tensorflow implementationsMachine Learning Specialization
Day202025-01-02Building and comparing models (Binary Classification)Machine Learning Specialization
Day212025-01-03Vectorization, Model training using TensoflowMachine Learning Specialization
Day222025-01-04Activation Functions, Softmax IntutionMachine Learning Specialization
Day232025-01-05Implementing SoftmaxMachine Learning Specialization
Day242025-01-06Backpropagaton, What and how??Machine Learning Specialization
Day252025-01-07Backpropagation - Why? Advices for applying machine LearningMachine Learning Specialization
Day262025-01-08Model selection, training test, cross validation, Bias and Variance, Learning curvesMachine Learning Specialization
Day272025-01-09Machine Learning Development Process, ML workflowMachine Learning Specialization
Day282025-01-10Implementing ML model: Error Analysis and Transfer LearningNotebook: Implementation, Machine Learning Specialization
Day292025-01-11Error Metrices, Encoding of Categorical Data, TransoformersMachine Learning Specialization , CampusX
Day302025-01-12Scikit-Learn Pipelines & Ridge Regression (L2 Regularization)Documentation: Scikit-Learn , CampusX
Day312025-01-13Lasso Regression (L1 Regularization), Elastic Net RegularizationML playlist @CampusX
Day322025-01-14Decision Tree Emtropy and Information GainML playlist @CampusX
Day332025-01-15Hyperparameters of Decision Tree with Scikit Learn, Regression TreesML playlist @CampusX , Visualize Yourself>>
Day342025-01-16Visualization Using DtreeViz(), Ensemble LearningGithub Repo: Dtreeviz, ML playlist @CampusX
Day352025-01-17Voting Ensemble >> Classification and RegressionML playlist @CampusX , Visualize Yourself
Day362025-01-18Bagging Ensemble > Classification and RegressionML playlist @CampusX
Day372025-01-19Random Forest: Intution, Working and difference with bagging, Random Forest HyperparametersML playlist @CampusX
Day382025-01-20Boosting Ensemble: Adaboost BoostingML playlist @CampusX
Day392025-01-21Understanding GradientBoosting with RegressionML playlist @CampusX
Day402025-01-22Gradient Boosting with ClassificationML playlist @CampusX , Vlog Link
Day412025-01-23XGboost IntroductionML playlist @CampusX
Day422025-01-24XGBoost for Regression and Classification, Catboost Vs XGboost Vs LightGBMML playlist @CampusX ,Research Paper
Day432025-01-25Stacking Ensemble, Understanding Blending and K foldML playlist @CampusX
Day442025-01-26K-Nearest Neighbor, Coding KNN from ScratchML playlist @CampusX
Day452025-01-27Support Vector MachineML playlist @CampusX
Day462025-01-28K-Means Clustering, DBSCANNotebook: K-Means clustering Demo , Notebook: DBSCAN demo
Day472025-01-29Hierarchical Clustering, Silhouette ScoreKaggle, ML playlist @CampusX
Day492025-01-30PCA (Principle Component Analysis), Implementing with MNIST datasetNotebook: Applying PCA on MNIST dataset
Day502025-02-01Visualizing and Comparing PCA, t-SNE, UMAP, and LDA + Revision with the course ML specializationMachine Learning Specialization
Day512025-02-02Anomaly DetectionMachine Learning Specialization, Notebook: Anomaly Detection
Day522025-02-03Collaborative FilteringMachine Learning Specialization
Day532025-02-04Project @ Football Players Market Value Prediction - Introduction and PlanningProject Plan
Day542025-02-05Project @ Football Players Market Value Prediction - Collecting Data (Scraping)Notebook
Day552025-02-06Project @ Football Players Market Value Prediction - Cleaning DataNotebook
Day562025-02-07Project @ Football Players Market Value Prediction - EDANotebook
Day572025-02-08Project @ Football Players Market Value Prediction - Feature Engineering: (Creating features, Transforming Features)Notebook
Day582025-02-09Project @ Football Players Market Value Prediction - ML: (Linear Regression with Refined Features and deploying with Streamlit)Notebook
Day592025-02-10Project @ Complete Streamlit setup for Linear RegressionStreamlit Documentation
Day602025-02-11Project @ Testing Ridge, Lasso, and Decision TreesProject @ Football Players Market Value Prediction
Day612025-02-12Project @ Had to hit reset from Feature EngineeringProject @ Football Players Market Value Prediction
Day622025-02-13Project @ Finalizing Project and Deploying itProject @ Football Players Market Value Prediction
Day632025-02-14Content-Based Movie Recommender System - PreprocessingNotebook
Day642025-02-15Content-Based Movie Recommender System - Building and DeploymentLive Demo
Day652025-02-16Diving into Deep LearningIntro to Deep Learning @MIT
Day662025-02-17PerceptronsDeep learning playlist @ CampusX
Day672025-02-18Perceptron, Loss function and gradient DescentDeep learning playlist @ CampusX , Grokking Deep Learning @Andrew W. Trask
Day682025-02-19Multilayer PerceptronDeep learning playlist @ CampusX
Day692025-02-20MLP notation, Forward PropagationDeep learning playlist @ CampusX
Day702025-02-21Loss Functions for deep learningDeep learning playlist @ CampusX
Day712025-02-22Backpropagation, deep diving this timeDeep learning playlist @ CampusX
Day722025-02-23Implementing Backpropagation for RegressionNotebook: Backpropagation Regression
Day732025-02-24Implementing Backpropagation for ClassificationNotebook: Implementation Backprop Classification
Day742025-02-25Revising old days, MemoizationDeep learning playlist @ CampusX
Day752025-02-26Vanishing Gradient, Exploding GradientDeep learning playlist @ CampusX
Day762025-02-27Implementing artificial neural networks (ann) for different datasetsDeep learning playlist @ CampusX
Day772025-02-28Improving Neural NetworksDeep learning playlist @ CampusX
Day782025-03-01Sequence Modeling / RNNs - Just OverviewIntro to Deep Learning @MIT
Day792025-03-02Transformers Attention - Just OverviewIntro to Deep Learning @MIT
Day802025-03-03CNNs - Just Overview Part 1Intro to Deep Learning @MIT
Day812025-03-04CNNs - Just Overview Part 2Intro to Deep Learning @MIT
Day822025-03-05Deep Generative Modeling - Just OverviewIntro to Deep Learning @MIT
Day832025-03-06Reinforcement Learning - Just OverviewIntro to Deep Learning @MIT
Day842025-03-07Deep Learning: Challenges & New Frontiers - Just OverviewIntro to Deep Learning @MIT
Day852025-03-08Early Stopping & Normalizing Inputs, DroputDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day862025-03-09Regularization, QuantizationDeep learning playlist @ CampusX
Day872025-03-10Activation Functions - RevisitedDeep learning playlist @ CampusX
Day882025-03-11Weight InitializationDeep learning playlist @ CampusX
Day892025-03-12Deep Learning OptimizersDeep learning playlist @ CampusX
Day902025-03-13Keras TunerDeep learning playlist @ CampusX
Day912025-03-14Deep Diving into CNNsDeep learning playlist @ CampusX
Day922025-03-15Understanding Paddings and StridesDeep learning playlist @ CampusX
Day932025-03-16Backpropagation in CNNs: A Quick BreakdownDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day942025-03-17LeNet5, Cat Vs Dog ClassificationDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day952025-03-18GPU slow than CPU - well in my case?Deep learning playlist @ CampusX
Day962025-03-19Data Augmentation, Pretrained ModelsDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day972025-03-20Visualizing Convolutional Layers, Transfer LearningDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day982025-03-21Keras Functional APIDeep learning playlist @ CampusX , Article
Day992025-03-21Finalizing Dog Cat Classifier ProjectProject - Live Demo
Day1002025-03-23Hidden Markov Model, Quantum Machine LearningMedium Article: Understanding Hidden Markov Models
Day1012025-03-24Exploring Pytorch SurfacelyDL with Pytorch - Datacamp
Day1022025-03-25Training a neural network with pytorchDL with Pytorch - Datacamp
Day1032025-03-26Evaluating and improving modelsDL with Pytorch - Datacamp
Day1042025-03-27Crawling through DL with pytrochDeep Learning for Coders -Fast.Ai, @CampusX
Day1052025-03-28Starting chapter 2 : from model to productionDeep Learning for Coders -Fast.Ai, @CampusX
Day1062025-03-29Exploring Autograd and Portfolio TweaksDeep Learning for Coders -Fast.Ai, @CampusX
Day1072025-03-30Refining Portfolio whole dayDeep Learning for Coders -Fast.Ai, @CampusX
Day1082025-03-31Autograd in PyTorch: Deeper UnderstandingDeep Learning for Coders -Fast.Ai, @CampusX
Day1092025-04-01PyTorch Training Pipeline (Manual + Using nn.Module)Deep Learning for Coders -Fast.Ai, @CampusX
Day1102025-04-02Dataset & DataLoader Class in PyTorchDeep Learning for Coders -Fast.Ai, @CampusX
Day1112025-04-03ANN on Fashion MNIST, GELUDeep Learning for Coders -Fast.Ai, @CampusX
Day1122025-04-04ANN on larger FMNIST dataset with GPU (local), GELU/SiLU historyDeep Learning for Coders -Fast.Ai, @CampusX
Day1132025-04-05Optimizing FMNIST NN using Dropouts, Regularization and Batch Normalization in PytorchNotebook
Day1142025-04-08RNNs revisited, Karpathy's blog, Project PlanningKarpathy Blog
Day1152025-07-08Classifying Footballers with their Eyes - Day 1Project Notebook
Day1162025-07-09Classifying Footballers with their Eyes – Day 2Project Notebook
Day1172025-07-10YOLO (You Only Look Once)YOLO Paper
Day1182025-07-11LSTM, GRU & Encoder-Decoder ArchitectureColah's Blog
Day1192025-07-12Bahdanau Attention and Luong AttentionBahdanau Paper
Day1202025-07-13Building a Seq2Seq Chatbot – Data Preparation & PreprocessingPyTorch Tutorial
Day1212025-07-14Building a Seq2Seq Chatbot - Defining Model (encoder, attention, decoder)Notebook
Day1222025-07-15Building a Seq2Seq Chatbot – Evaluation / DeploymentLive Demo
Day1232025-07-20Transformers – Deep Dive into Attention and ArchitectureAttention Paper
Day1242025-07-21Transformers – Vitals [understanding everything]Transformer Guide
Day1252025-07-22GPT from Scratch - Project SetupKarpathy's Tutorial
Day1262025-07-23GPT from Scratch - Bigram Language ModelNotebook
Day1272025-07-24GPT from Scratch - Self-AttentionNotebook
Day1282025-07-25GPT from Scratch – Complete TransformerNotebook
Day1292025-07-26Image Captioning – KickoffNotebook
Day1302025-07-27Image Captioning – Feature Extraction & PipelineNotebook
Day1312025‑07‑28Image Captioning – Training & DeploymentBuild a Large Language Model from Scratch
Day1322025‑07‑29Tokenizer Tricks (Subword, Byte-Pair Encoding) + Comparing ArchitecturesBuild a Large Language Model from Scratch
Day1332025‑07‑30Implementing seq2seq (Diff Approach of Visualizations)Build a Large Language Model from Scratch
Day1342025‑07‑31Implementing Transformer Encoder/Decoder AgainBuild a Large Language Model from Scratch
Day1352025‑08‑01Pretraining on Unlabeled Data, Evaluating, Loading Pretrained WeightsBuild a Large Language Model from Scratch
Day1362025‑08‑02Finetuning – ClassificationBuild a Large Language Model from Scratch
Day1372025‑08‑03Finetuning – Teaching LLMs to Follow Prompts and Perform Complex Tasks, VisualizationsBuild a Large Language Model from Scratch
Day1372025‑08‑03Finetuning – Teaching LLMs to Follow Prompts and Perform Complex Tasks, VisualizationsBuild a Large Language Model from Scratch
Day1382025‑08‑04LLM Fine-Tuning & EvaluationBuild a Large Language Model from Scratch
Day1392025‑08‑05Exploring Hugging Face TransformersBuild a Large Language Model from Scratch
Day1402025‑08‑06Project – Sentiment Analysis [Planning] + Exploring ViTNotebook, Kaggle Code
Day1412025‑08‑07Project – Sentiment Analysis [Preprocessing] + ViT ArchitectureKaggle Code
Day1422025‑08‑08Project – Sentiment Analysis [EDA + Testing GloVe]Notebook, Notebook, Kaggle Code
Day1432025‑08‑09Project – Sentiment Analysis [Advanced Architectures]Notebook, Notebook, Kaggle Code
Day1442025‑08‑10Project – Sentiment Analysis [App Deployment]Live Demo, Code
Day1452025‑08‑11Diving Deep into Vision Transformers (ViTs)Blog
Day1462025‑08‑12Diving into Diffusion ModelsVideo
Day1472025‑08‑13Diffusion Model Deep DivePaper
Day1482025‑08‑15Naive Bayes & Gaussian Mixture ModelsLangchain Playlist
Day1492025‑08‑17Introduction to LangchainLangchain Playlist
Day1502025‑08‑18Introduction to Langchain ComponentsLangchain Playlist
Day1512025‑08‑19Deep Dive into Langchain ModelsLangchain Playlist
Day1522025‑08‑20Langchain Prompts and Building a Simple ChatbotLangchain Playlist
Day1532025‑08‑21Structured Output with LangchainLangchain Playlist
Day1542025‑08‑23Langchain Output ParsersLangchain Playlist
Day1552025‑08‑24Langchain Chain Fundamentals (Simple, Sequential, Parallel, Conditional Chains)Langchain Playlist
Day1562025‑08‑25Langchain Runnables (Modular Components, Composable Workflows)Langchain Playlist
Day1572025‑08‑26Runnable Modules Deep Dive (Sequence, Parallel, Passthrough, Lambda, Branch)Langchain Playlist
Day1582025‑08‑27Document Loaders & Text Splitters (RAG Foundations)Langchain Playlist
Day1592025‑08‑28Vector Stores in Langchain (Chroma, CRUD Operations, Similarity Search)Langchain Playlist
Day1602025‑08‑29Retrievers & Few-Shot Learning (Wikipedia, Vector, MMR, MultiQuery, Contextual)Langchain Playlist
Day1612025‑08‑30RAG Application for UCL Draw (Text Splitting, Vector Embeddings, Response Generation)Langchain Playlist
Day1622025‑08‑31Langchain Tools (Built-in & Custom, Tool Calling & Binding)Langchain Playlist
Day1632025‑09‑01Langchain Agents (Zero-Shot, Conversational, ReAct DocStore, Self-Ask)Langchain Playlist
Day1642025‑09‑02Local Agent with Ollama & Langchain (ChromaDB, RAG)Langchain Playlist
Day1652025‑09‑03Introduction to LangGraph (Stateful Agent Workflows)Langgraph Playlist
Day1662025‑09‑04Agentic AI Fundamentals (Autonomy, Components, Planning, Memory)Langchain Playlist
Day1672025‑09‑05LangChain vs LangGraph Comparison (State Management, Chatbot Example)Langgraph Playlist
Day1682025-09-06Building a Branching ChatbotLanggraph Playlist
Day1692025-09-07Persistence with CheckpointsLanggraph Playlist
Day1702025-09-08Exploring LangSmithLanggraph Playlist
Day1712025-09-11Contextual Q&A with MemoryLangchain Playlist
Day1722025-09-12Bhagavad Gita Expert ChatbotLanggraph Playlist
Day1732025-09-13Multi-Agent Debating SystemLanggraph Playlist
Day1742025-09-14Debate Agent App CompletionLanggraph Playlist
Day1752025-09-15Introduction to FastAPI for MLFastAPI Documentation
Day1762025-09-16FastAPI ImplementationFastAPI Documentation
Day1772025-09-17HTTP Request Methods and REST ArchitectureFastAPI Documentation
Day1782025-09-18FastAPI Parameters and Request BodyFastAPI Documentation
Day1792025-09-19Mini Project with FastAPIFastAPI Documentation
Day1802025-09-20Building Industry-Ready APIs with FastAPIFastAPI Documentation
Day1812025-09-23Containerizing FastAPI ApplicationsFastAPI Documentation
Day1822025-09-24fastapi deployment on awsaws docs
Day1832025-09-25project setup – choose your own adventureproject repo
Day1842025-09-26database design and core componentsproject repo
Day1852025-09-27api implementation and background tasksproject repo
Day1862025-09-28backend completion and debuggingproject repo
Day1872025-09-29frontend integration and project completionproject repo
Day1882025-10-01concurrency patterns in fastapistarlette concurrency
Day1892025‑10‑02self-supervised learning – foundationsLil'Log SSL Blog
Day1902025‑10‑03mcp and lazy weekMasked Conditional Prediction
Day1912025‑10‑04ssrl – image & video approachesLil'Log SSRL
Day1922025‑10‑05wrapping up ssl + fun readsPostgres vs SQLite, GPT Speculations
Day1932025‑10‑06llama2 fine-tuning with qloraQLoRA Fine-Tuning Guide
Day1942025‑10‑12gemma 2 fine-tuning using unslothGemma2 Fine-tuning Notebook
Day1952025‑10‑13saving & loading lora adapters, explored "lora without regret"LoRA Blog Post
Day1962025‑10‑14attempted RAG evaluation, faced compatibility issuesMCP Documentation
Day1972025‑10‑16built a custom MCP server, integrated with CursorCustom Implementation
Day1982025‑10‑18studied RL fundamentals: policies, MDPs, rewardsHuggingFace RL Course
Day1992025‑10‑20explored Q-learning & Deep Q-learning, lunar lander envHuggingFace RL Course
Day2002025‑10‑21trained PPO agent on LunarLander-v3 using SB3HuggingFace RL Course



Day 01: Basics of Linear Algebra

linear algebra is used to represent data, perform matrix operations, and solve equations in algorithms like regression, pca, and neural networks.

  • Scalars, Vectors, Matrices, Tensors: Basic data structures for ML.

  • Linear Combination and Span: Representing data points as weighted sums. Used in Linear Regression and neural networks.

  • Determinants: Matrix invertibility, unique solutions in linear regression.

  • Dot and Cross Product: Similarity (e.g., in SVMs) and vector transformations.

Slow progress right?? but consistent wins the race!


Day 02: Decomposition, Derivation, Integration, and Gradient Descent

  • Identity and Inverse Matrices: Solving equations (e.g., linear regression) and optimization (e.g., gradient descent).

  • Eigenvalues and Eigenvectors: PCA, SVD, feature extraction; eigenvalues capture variance.

  • Singular Value Decomposition (SVD): PCA, image compression, and collaborative filtering.

Notes Here

Calculus Overview:

  • Functions & Graphs: Relationship between input (e.g., house size) and output (e.g., house price).

  • Derivatives: Adjust model parameters to minimize error in predictions (e.g., house price).

  • Partial Derivatives: Measure change with respect to one variable, used in neural networks for weight updates.

  • Gradient Descent: Optimization to minimize the cost function (error).

  • Optimization: Finding the best values (minima/maxima) of a function to improve predictions.

  • Integrals: Calculate area under a curve, used in probabilistic models (e.g., Naive Bayes).

Revised statistics and probability concepts. Ready for the ML Specialization course!


Day 03: Supervised Machine Learning: Regression and Classificaiton

  • Supervised Learning:
    Supervised Learning
  • Regression:
  • Classification:

Day 04: Unsupervised Learning: Clustering, dimensionality reduction

data only comes with input x, but not output labels y. Algorithm has to find structure in data. Unsupervised Learinging

  • Clustering: group similar data points together
    alt text
  • dimensionality reduction: compress data using fewer numbers eg image compression
  • anomaly detection: find unusual data points eg fraud detection

Day 05: Univariate Linear Regression:

  • Learned univariate linear regression and practiced building a model to predict house prices using size as input, including defining the hypothesis function, making predictions, and visualizing results.

Notebook: Model Representation

- Univariate Linear Regression Quiz


Day 06: Cost Function:

alt text Visualization of cost function: Visualization of cost function

  • manually reading these contour plot is not effective or correct, as the complexity increases, we need an algorithm which figures out the values w, b (parameters) to get the best fit time, minimizing cost function

Notebook: Model Representation

Gradient descent is an algorithm which does this task


Day 07: Gradient Descent

Notebook: Gradient descent

gradient descent learned the basics by assuming slope constant and with only the vertical shift. later learned GD with both the parameters w and b. alt text

alt text


Day 08: Effect of learning Rate, Cost function and Data on GD

  • learning rate on GD:Affects the step size; too high can overshoot, too low can slow convergence
- cost function on GD:Smooth, convex functions help faster convergence; complex ones may trap in local minima
  • Data on GD:Quality and scaling affect stability; more data improves gradient estimates

Notebook: gradient descent animation 3d


Day 09: Linear Regression with multiple features, Vectorization

Predicts target using multiple features, minimizing error.

  • Vectorization: Matrix operations replace loops for faster calculations.

alt text](image.png

Lab1: Vectorization


Day10: Feature Scaling

Lab2: Multiple Variable

Today, I learned about feature scaling and how it helps improve predictions. There are multiple methods for feature scaling, including

  • Min-Max Scaling
  • Mean Normalization
  • Z-Score Normalization alt text To ensure proper convergence: check the learning curve to confirm the loss is decreasing. Start with a small learning rate and gradually increase to find the optimal value.

alt text alt text


Day 11: Feature engineering and Polynomial Regression

feature engineering improves features to better predict the target.

eg If we need to predict the cost of flooring and have length and breadth of the room as features, we can use feature engineering to create a new feature, area (length × breadth), which directly impacts the flooring cost.

![alt text](01-Supervised-Learning/images/notes_featureengineering.jpg)

explored polynomial regression that models the relationship between variables as a polynomial curve instead of a straight line

Equation:
y = b₀ + b₁x + b₂x² + ... + bₙxⁿ It is useful for capturing nonlinear relationships in data. alt text

alt text

Lab1: Feature Scaling and Learning Rate
Lab2: Feature Engineering and PolyRegression


Day 12: Linear Regression using Scikit Learn

Had a productive session with linear regression in scikit learn. The lab helped me get a better grasp of the process, though I need more practice with tuning models. Also revisited the Scikit-Learn models ,more comfortable with them now

  • Scikit-learn is an open-source Python library used for machine learning that provides simple and efficient tools for data analysis, including algorithms for classification, regression, clustering, and dimensionality reduction.

alt text alt text alt text
Notebook: ScikitLearn GD


Day 13: Classification

Notebook: Graded Lab alt text
Notebook: Classification solution

  • Classification is the process of categorizing items into different groups based on shared characteristics, like classifying tumors into benign (non-cancerous) and malignant (cancerous) based on their growth behavior and potential to spread. The example above demonstrates that the linear model is insufficient to model categorical data. The model can be extended as described in the following lab.

Day 14: Logistic Regression, Sigmoid Function

  • Logistic Regression: A classification algorithm used to predict probabilities of binary outcomes. Logistic regression on categorical data
  • Sigmoid Function: A mathematical function that maps any input to a value between 0 and 1, used in logistic regression to model probabilities. sigmoid-function

Notebook: Sigmoid Function


Day 15: Decision Boundary, Cost Function

Notebook: Cost Function

  • Decision Boundary: A line or surface that separates different classes in a classification problem based on the model’s predictions.
  • cost function: formula cost function notes

Notebook: Decision boundary

Notebook Logistic Loss


Day 16: Gradient Descent for Logical Regression

Notebook: Gradient Descent Model implementation

Notebook: GD with Scikit-learn

Learned logistic regression cost, gradient descent, and sigmoid derivatives through step-by-step derivations and comparisons with linear regression.

notes_day16 alt text


Day 17: Underfitting, Overfitting

Today, explored teh concepts, overfitting (high variance), underfitting (high bias) and generalization(just right). Regularization to reduce Overfitting. Explored Regularized logistic regression.

  • If the data is in non linear behaviour then we have to appply Ml algos like decision tree, random forest and svm.

Explored hypermeters of logistic regression, and gained some knowledge. text text

Notebook:Overfitting Solution

text

Notebook:Regularization text


Day 18: Neurons, Layer, Neural netowrk, forward propagation

  • neural network: neural networks are machine learning algorithms that model complex patterns using multiple hidden layers and non-linear activation functions. they take inputs, pass them through hidden layers of neurons, and output a prediction. alt text

  • Neurons: a neuron takes weighted inputs, applies an activation function, and outputs a result. inputs can be features or outputs from previous neurons, with weights adjusting their influence. alt text fig: single neuron in action

  • Synapse: synapses connect neurons and carry the weighted inputs. each connection has a weight that adjusts during training.

  • weights: weights control the strength of connections between neurons. they are multiplied by inputs to influence the output, and are adjusted during training.

Popular activation functions include relu and sigmoid.

  • Bias: bias is a constant added to the weighted input before applying the activation function, helping the model represent patterns that don’t pass through the origin.

  • Layers: alt text

    • input layer: holds the data for the model, with each neuron representing an attribute.
    • hidden layer: applies activation functions to the inputs and passes results to the next layer.
    • output layer: receives input from the last hidden layer and returns the model’s prediction.

Day 19: Forward Propagation

  • Forward Propagation: Input data is “forward propagated” through the network layer by layer to the final layer which outputs a prediction.

Notes for today: alt text alt text

Matrix Representation: text How forward Prop works for digit classification?? text Notebook: Neurons and Layers Notebook: A small Neural Netowrk using tensoflow

  • tensorflow basics:
    • representation of data:numpy arrays used for input (e.g., 2D arrays).

      x = np.array([[1, 2, 3], [4, 5, 6]])
      
      
    • building a neural network:

      1. define layers:

        layer1 = dense(units=25, activation='sigmoid')
        layer2 = dense(units=15, activation='sigmoid')
        layer3 = dense(units=1, activation='sigmoid')
        
        
      2. stack layers in a model:

        model = sequential([layer1, layer2, layer3])
        
        
      3. compile and train:

        model.compile(optimizer='adam', loss='binary_crossentropy')
        model.fit(x, y, epochs=10)
        
        
    • visualization:neurons connect layer by layer, with weights and biases computed at each step (refer to attached gif).


Day 20: Python Implementation from Scratch

Implemented forward propagation to compute predictions and backpropagation to optimize weights for binary classification.

  • AGI: An advanced AI capable of generalizing across tasks like humans.

loss graph

model accuracy

Notebook: Building Models


Day 21: Vectorization, Model Training

Exploredd Vectorization for efficient computation

  • Tensorflow: model.compile, binary_crossentropy, model.fit trained a binary classification model and tested its accuracy

alt text Training Model with tensorflow: alt text Notes for today: Notes for today


Day 22: Activation Functions, Softmax

the universal approximation theorem explains that a neural network with enough hidden neurons and non-linear activations like sigmoid or relu can approximate almost any function, even complex patterns like wavy graphs.

Activation Functions:

commonly used activation functions include:

  • sigmoid: squashes values between 0 and 1, often used for binary classification.
  • relu: outputs 0 for negatives and the input itself for positives, commonly used in hidden layers.
  • tanh: outputs between -1 and 1, useful for centered data.

multiclass example

for multiclass classification, softmax is ideal in the output layer as it converts logits into probabilities that sum to 1. during training, the model adjusts weights to maximize the correct class probability, using categorical cross-entropy loss. softmax generalizes logistic regression, which is typically used for binary classification. in both, activation and loss functions differ based on the output type.

Logistic Vs softmax:

logistic vs softmax

Notes:

Notes Notes

Day 23: Implementation of Softmax Regression

NOTE: softmax regression is a classification algorithm that calculates probabilities for multiple classes using a linear combination of inputs and the softmax function. the class with the highest probability is chosen as the prediction

Improved Implementation of Softmax: Roundoff

Tensorflow implementation:

model = Sequential(
    [ 
        Dense(25, activation = 'relu'),
        Dense(15, activation = 'relu'),
        Dense(4, activation = 'softmax')    # < softmax activation here

        ##         Dense(4, activation = 'linear')   #<-- Note
    ]
)
model.compile(
    loss=tf.keras.losses.SparseCategoricalCrossentropy(),
    ##     loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),  #<-- Note ---- This is preferred model softmax and loss are combined for more accurate result.
    optimizer=tf.keras.optimizers.Adam(0.001),
)

model.fit(
    X_train,y_train,
    epochs=10
)
    

Day 24: Backpropagation - what and how ?

  • what? Backpropagation adjusts nneural network weights by propagating errors backward using the chain rule and optimizing them with gradient descent.
  • how? Forward pass to compute outputs, Calculate errors, propagate them backward, and update weights iteratively. Take a look at document below.

Backpropagation example with Neural Network: Backpropagation example with Neural Network

Notebook: Backprop

pdf- How to implement??

Notes:

alt text

Day 25: Backpropagation - Why? Advices for applying machine Learning

Today, I dived into the reasons behind backpropagation's effectiveness in training neural networks. It's not just about adjusting weights; it's the gradients that guide optimization, helping the model minimize error and improve predictions. The backpropagation process makes sure that the error gets distributed in a way that leads to better learning.

How????

  • Gradients: they ensure efficient error minimization by guiding weight updates.
  • Optimization: properly tuned gradients lead to smoother optimization and faster convergence.
  • Model Evaluation: evaluating model performance becomes easier with backpropagation because of the systematic error propagation and weight adjustments.

Notes:

alt text


Day 26: Model selection and training/cross validation/test sets, Bias and Variance

mnist dataset: Label and our prediction after training alt text Errors in our prediction: alt text

  • the importance of splitting data into training, validation, and test sets cross-validation and its role in hyperparameter tuning.
  • Diagnosing bias and variance with error trends: high bias = underfit, high variance = overfit.
  • regulariaztion to handle the tradeoff between bias and variance.
  • how learning curves reveal insights about model performance and whether gathering more data will help. summary of learning algorithm: summary of learning algorithm:

Notes: notes

Notebook: Practice Lab: Neural Networks for Handwritten Digit Recognition, Multiclass

Notebook: Diagnosing Bias and Variance

Notebook: Model Evaluation and selection


Day 27: Machine Learning Development Process, ML workflow

machine learning development process

  1. ml development is iterative, involving:
    • choosing model and data architecture.
    • training the model.
    • diagnosing bias, variance, and errors.

alt text 2. error analysis: identify and fix patterns in model failures.
3. adding data:

  • data augmentation: modify existing data (e.g., distortions).
  • data synthesis: create artificial data.
  1. transfer learning:

    • reuse pre-trained models for similar tasks.
    • fine-tune them with your own data.
  2. ml projects follow these steps:

    • data collection.
    • preprocessing.
    • modeling.
    • evaluation.
    • deployment.
      alt text
    • monitoring.
  3. ethics and fairness -
    ensure ethical use by:

    • avoiding biased decisions in loans, jobs, etc.
    • preventing harmful applications like deepfakes.

Notes: Notes


Day 28: Machine Learning Model: Error Analysis and Transfer Learning

Notebook: Code Implementation from Scratch

  • Confusion Matrix Analysis: the most frequent error is misclassifying 5 as 3. overall, the error rate is around 8%. alt text

  • Iterations Insight: after 200 iterations, the error rate does not decrease significantly, suggesting that 200 iterations are enough for the model to converge. alt text

  • Data Augmentation Insight: despite applying data augmentation, there was no improvement in accuracy. this is because the MNIST dataset is already preprocessed, with centered and normalized images, making the augmentation techniques less effective. in general, data augmentation works best when the dataset is smaller or images are not preprocessed. alt text

  • Transfer Learning with MobileNetV2:

    • Training set accuracy improved from 56.95% to 73.02%.
    • Validation accuracy increased from 70.52% to 74.06%.
    • Loss decreased consistently for both training and validation sets, signaling better learning and generalization.
    • The model shows significant improvement over epochs.
    • Using transfer learning, the model started with an initial accuracy of 74% (pre-trained on ImageNet). As training progressed, the model continued to adapt, improving with each epoch. transfer learning was effective, yielding good results with fewer epochs. alt text

Day 29: Error Metrices, Encoding of Categorical Data, Transoformers

Notebook: Lab week 3: Improving Model

Notebook: Error Metrics

Precision vs. Recall Trade-Off:

  • High Precision: Only hire candidates you’re sure are good. Result: Fewer bad hires, but you might miss some great ones. Example: You hire 5 people, all are good, but you missed 10 other good ones.
  • High Recall: Hire as many as possible to ensure no good candidate is missed. Result: You catch all great candidates but end up with some bad hires too. Example: You hire 50 people, 20 are good, but 30 are bad.
  • When to Focus on Each?
  1. Precision: When mistakes (bad hires) are costly. Example: Hiring a brain surgeon.

  2. Recall: When missing good candidates is worse. Example: Hiring for a customer service team.

Encoding of Categoriical Data:

Encoding TypeUse WhenExample
Label EncodingSmall, unordered categoriesColors: [Red, Blue]
Ordinal EncodingOrdered categoriesEducation: [Low, High]
One-Hot EncodingNominal data, fewer unique categoriesDays: [Mon, Tue, Wed]

Types of Transformers:

TransformerPurposeExample Use CaseInputOutput
Column TransformerApply different transformations to different columns (e.g., scaling and encoding).Scale age and one-hot encode city names.Age: [25, 35, 45], City: [NY, LA, CHI][-1.22, 0, 0, 1], [0, 1, 0, 0], [1.22, 0, 1, 0]
Function TransformerApply a mathematical function (e.g., log or sqrt) to all values.Apply logarithmic transformation to data.[1, 10, 100][0.69, 2.39, 4.61]
Power TransformerNormalize and reduce skewness in data, making it more Gaussian-like.Stabilize variance in highly skewed data.[1, 10, 100][-1.22, 0.0, 1.22]

Day 30: Scikit-Learn Pipelines & Ridge Regression (L2 Regularization)

I explored how to create pipelines in Scikit-learn to streamline the process of combining multiple steps (like preprocessing, model fitting, and regularization) into a single object. This simplifies workflows and ensures reproducibility.

There are three techniques of regularization:

  • Ridge (L2)
  • Lasso (L1)
  • Elastic Net (Combination)

Key Understanding of Ridge Regression:

  1. How the coefficient Get affected?
  • Regularization shrinks the coefficients, preventing them from becoming too large, which reduces overfitting. alt text
  1. Higher Values are impacted more
  • The larger the regularization value (alpha), the more the coefficients are reduced. alt text
  1. Impact on Bias Variance Tradeoff
  • Higher regularization increases bias but reduces variance, making the model more generalizable. alt text
  1. Effect on Loss Function
  • adds a penalty term that limits the magnitude of the coefficients. alt text
  1. Why Ridge Regression is called so?
  • Named after the concept of creating a "ridge" or constraint on the model’s coefficients, preventing them from growing too large.

Notes: alt text alt text

Notebook: Key Understandings of ridge Regression


Day 31: Lasso Regression (L1 Regularization), Elastic Net Regularization

Lasso Regression:

  1. How are coefficients affected by λ (alpha)? As λ increases, regularization strength grows, shrinking less important coefficients to exactly zero. Larger λ values lead to feature selection by removing irrelevant features. alt text

  2. Are higher coefficients affected more? No. Lasso affects smaller coefficients more, shrinking them to zero first. Larger coefficients remain relatively unaffected if they contribute significantly to the model. alt text

  3. Impact of λ on bias and variance:

    • Higher λ increases bias (simpler model) and reduces variance.
    • Lower λ reduces bias (complex model) but increases variance. alt text
  4. Effect of Regularization on Loss Function:

    • Lasso add λ sum(w_i) to the loss, promoting sparsity by penalizing the absolute magnitude of coefficients. Unlike Ridge, it can remove features entirely, improving interpretability. alt text

Elastic Net Regularization:

Elastic Net combines L1 (Lasso) and L2 (Ridge) penalties. It selects important features by shrinking some coefficients to zero (like Lasso) and handles correlated features by shrinking coefficients without setting them to zero (like Ridge). It’s useful when features are both highly correlated and some are irrelevant. Elastic Net is controlled by two parameters: α (mix of Lasso and Ridge) and λ (regularization strength). alt text

Notes: alt text


Day 32: Decision Tree: Entropy and Information Gain

A decision tree is a flowchart-like structure used for classification or regression, where data is split into branches based on conditions until a final decision (leaf) is reached. alt text

  • Decision Tree on Categorical Variables: Splits data based on categories like "Sunny" or "Rainy".
  • Decision Tree on Numerical Variables: Splits data using thresholds like Age > 30.
  • How Decision Tree Works: Repeatedly splits data into smaller groups based on conditions.
  • Terminology: Root (start), Branch (path), Leaf (decision).
  • Pro - Simple to understand; Con - Can overfit.
  • Entropy: Measures uncertainty; low entropy = purer data.
  • Entropy Calculation: ∑ p(x) * log2(p(x)).
  • Information Gain: Reduction in entropy after splitting data. IG=Entropy(before)−Weighted Entropy(after)
  • Gini Impurity: Measures group purity, faster than entropy.
  • Why Use Gini Over Entropy: Simpler and computationally faster.

Notes; alt text


Day 33: Hyperparameters of DT with sclearn, Regression Trees:

alt text

Studied hyperparameters of Decision Trees in Scikit-learn and techniques to handle overfitting and underfitting.

  • Criterion (gini, entropy, log loss): Determines the quality of a split.
  • Splitter: Helps reduce overfitting with better random splits.
  • Max Depth: Controls tree depth; too high causes overfitting, too low causes underfitting.
  • Min Sample Split: Sets the minimum samples required to split a node.
  • Min Sample Leaf: Sets the minimum samples per leaf.
  • Max Features: Determines how many features to use for splits.
  • Max Leaf Nodes: Limits the number of leaf nodes.
  • Min Impurity Decrease: Controls splitting based on impurity reduction.

A Regression Tree predicts continuous variables by splitting data to minimize variance. The best split is determined by maximizing variance reduction, calculated as the variance of the root node minus the weighted average variance of the leaf nodes.

Code: alt text Output: alt text


Day 34: Visualization Using dtreeviz, Ensemble Learning Overview

Notebook: Visualization of Decision Tree Official Notebook: Visualization of Decision Tree

  • started learning about ensemble learning: understood the concept of "wisdom of the crowd" in ensemble methods. types of ensemble methods: voting, bagging, boosting, and stacking. overview of random forest and bootstrapped aggregation (bagging). learned how boosting adjusts predictions iteratively.

Visuals of Bagging, boosting and stacking: Notes Notes; Notes


Day 35: Voting Ensemble > Classification and Regression

  • Soft voting: logistic regression: 60% fraud, random forest: 80% fraud, svm: 40% fraud → final probability: (60% + 80% + 40%) / 3 = 60%. alt text

  • Hard voting: majority wins, logistic regression predicts "fraud," random forest predicts "not fraud," and svm predicts "fraud" → final prediction: "fraud."

Conclusions from Notebook: voting Classifier

  • weighted voting: assigning weights to classifiers helps emphasize stronger models, further improving performance.

  • same algorithm, different hyperparameters: tweaking hyperparameters (e.g., kernel degree in SVM) can lead to significant accuracy changes, highlighting the value of hyperparameter optimization.

Classification:


voting_clf = VotingClassifier(
    estimators=[
        ('lr', log_reg),  # logistic regression: good for linear patterns
        ('rf', rand_forest),  # random forest: captures complex relationships
        ('svc', svm_clf)  # support vector machine: handles edge cases
    ],
    voting='soft',  # averages probabilities from all models for final prediction
    weights=[2, 1, 1],  # gives higher importance to logistic regression
    n_jobs=-1  # enables parallel processing for faster training
)


What to use? soft voting or hard voting, depends if possible use both and then try to find out: alt text alt text

Regression:

  • voting regressor combines multiple regression models to predict continuous values.

similarly, in regression, soft voting averages continuous predictions, and weighted voting helps models with higher performance contribute more.

Voting regressor:

Visualize voting regression here: link

voting_reg = VotingRegressor(
    estimators=[
        ('lr', lin_reg),  # linear regression: good for linear relationships
        ('rf', rand_forest_reg),  # random forest regressor: handles non-linear patterns
        ('svr', svr_clf)  # support vector regressor: captures complex relationships
    ],
    weights=[2, 1, 1],  # assigns higher importance to linear regression
    n_jobs=-1  # enables parallel processing for faster training
)


Day 36: Bagging Ensemble > Classification and Regression

  • bagging (bootstrap aggregating) is an ensemble learning technique that combines predictions from multiple models trained on different subsets of the data (created via bootstrapping) to improve accuracy and reduce variance.

  • Intution:

    alt text

    • random sampling: create multiple datasets by sampling with replacement from the original dataset.
    • train independently: train a model (single preferred) on each bootstrapped dataset (e.g., decision trees).
    • combine predictions: aggregate their predictions by majority voting (classification) or averaging (regression). this reduces overfitting and increases stability, especially for high-variance models.

Notebook: Bagging Intution

Classification:

  • Intution alt text
  • Code Demo
from sklearn.ensemble import BaggingClassifier
from sklearn.tree import DecisionTreeClassifier

# initialize the bagging classifier
bagging_model = BaggingClassifier(
    base_estimator=DecisionTreeClassifier(),  # base model for ensemble; here, decision trees
    n_estimators=10,                          # number of base models to train
    max_samples=1.0,                          # fraction of the dataset for each base model (1.0 = 100%)
    max_features=1.0,                         # fraction of features used in each bootstrap sample
    bootstrap=True,                           # sample datasets with replacement (True enables bootstrapping)
    bootstrap_features=False,                 # sample features with replacement (False = no feature bootstrapping)
    random_state=42                           # seed for reproducibility
)


Regression:

  • Intution: alt text
  • Code Demo:
bagging_model = BaggingRegressor(
    base_estimator=DecisionTreeRegressor(),  # base model for ensemble; here, decision trees
    n_estimators=10,                          # number of base models to train
    max_samples=1.0,                          # fraction of the dataset for each base model (1.0 = 100%)
    max_features=1.0,                         # fraction of features used in each bootstrap sample
    bootstrap=True,                           # sample datasets with replacement (True enables bootstrapping)
    bootstrap_features=False,                 # sample features with replacement (False = no feature bootstrapping)
    random_state=42                           # seed for reproducibility
)


Day 37: Random Forest: Intution, Working and difference with bagging, Random Forest Hyperparameters

random forest is like a bunch of decision trees making a group decision. each tree gets a vote on the outcome, and the most votes win. it's like asking a bunch of experts for advice and going with the majority.

Sampling Techniques:

  • row sampling
  • column sampling
  • combination

Notebook Random Forest

Random Forest vs bagging:

  • why random forest performs well: it reduces overfitting by averaging multiple decision trees trained on different random subsets of data and features, improving accuracy and robustness.

  • random forest vs bagging: both use multiple trees, but random forest adds feature randomness at each split, making trees less correlated and boosting performance.

alt text

Notebook: Random forest Vs bagging

Random forest Hyperparameters:


model = RandomForestClassifier(
    # core hyperparameters
    n_estimators=100,         # number of decision trees in the forest
    max_features="sqrt",      # max features to consider at each split
    max_depth=None,           # max depth of each tree; None = grow fully
    min_samples_split=2,      # min samples needed to split a node
    min_samples_leaf=1,       # min samples required in a leaf node
    bootstrap=True,           # with replacement or without replacement

    # advanced hyperparameters
    max_leaf_nodes=None,      # max number of leaf nodes per tree; None = unlimited
    min_weight_fraction_leaf=0.0,  # min fraction of total weight for a leaf node
    class_weight=None,        # weights for handling class imbalance (e.g., 'balanced')
    ccp_alpha=0.0,            # complexity parameter for pruning; trade-off between size and accuracy
    criterion="gini",         # metric to evaluate splits: "gini" (default) or "entropy"
    warm_start=False,         # reuse previous trees for incremental training; False = train from scratch
    oob_score=False,          # whether to use out-of-bag samples to estimate generalization accuracy
    verbose=0,                # verbosity of output (0 = silent)
    n_jobs=-1,                # number of CPU cores for parallel processing; -1 = use all cores
    random_state=42           # seed for reproducibility
)


Day 38: Boosting Ensemble: Adaboost Boosting

  • boosting is a sequential ensemble learning method where models correct the errors of previous models to reduce bias. combines weak learners (like decision stumps) to create a strong predictive model. Boosting

Notes: alt text

AdaBoost Exploration:

step-by-step understanding of adaboost’s workflow, including:

  • calculating sample weights and weak learner errors.
  • updating model contributions based on performance.
  • final prediction via weighted majority vote.
  • visualized how adaboost focuses on difficult samples.

implemented adaboost from scratch without using sklearn. Adaboost from scratch

Notebook: Adaboost Implementation


Day 39: Understanding GradientBoosting with Regression:

"a model that learns step-by-step by fixing the mistakes of the previous model." Algorithm: alt text

  • Comparision between Adaboost and Gradietn boost: alt text

boosting + gradients (gradients = direction to minimize error).

pseudo-  residual = actual - predicted
new prediction = old prediction + (learning_rate × residual)

gradient boosting variations:

  • xgboost: uses regularization and handles missing data efficiently.
  • lightgbm: faster with large datasets.
  • catboost: great for categorical data.
    gb = GradientBoostingRegressor(
        n_estimators=100,      # number of trees
        learning_rate=0.1,     # step size for updates
        max_depth=3,           # depth of each tree
        min_samples_split=2,   # min samples to split a node
        min_samples_leaf=1,    # min samples per leaf
        subsample=1.0,         # fraction of samples per tree
        max_features=None,     # use all features
        random_state=42        # ensures reproducibility
    )

Notes: alt text


Day 40: Gradient Boosting with Classification

Overview

  • Gradient boosting improves classification by minimizing log-loss iteratively.
  • Each tree focuses on correcting errors made by previous trees.
  • Uses gradients of the loss function to guide updates.

Algorithm

  1. Initialize predictions: F0(x) (log-odds for classification).
  2. For each iteration:
    • Compute pseudo-residuals:
      ri = - ∂(loss) / ∂F(xi)
    • Fit a weak learner (tree) to ri.
    • Update predictions:
      F(x) = F(x) + η * h(x)
      where η is the learning rate.

Loss Function

  • Binary classification: log-loss.
    Loss = -[y log(p) + (1 - y) log(1 - p)]
  • Multiclass classification: softmax loss.
gb = GradientBoostingClassifier(
    n_estimators=100,
    learning_rate=0.1,
    max_depth=3,
    random_state=42
)

Notes: alt text

Notebook: Gradient Boosting Classification

Blog Link

Visuals: Visualizations Visualizations


Day 41: Variations of Gradient Boosting: XGBoost - Introduction.

  • xgboost
  • lightgbm
  • catboost

XGBoost:

alt text xgboost (eXtreme Gradient Boosting) is an advanced implementation of gradient boosting that addresses some key limitations in traditional gradient boosting and adaboost. here's a quick rundown of what i've learned so far. Benchmark

  • Flexibility

    • cross-platform support means xgboost works on different operating systems without much hassle.
    • it supports multiple languages like python, c++, and r, which makes it easy to integrate with your preferred tech stack.
    • integrating with other libraries like scikit-learn or spark is a breeze.
    • it can handle all kinds of ml problems — whether it's classification, regression, or ranking.
  • Speed

    • parallel processing is key. it uses all your cores to speed things up, so you get faster results.
    • optimized data structures (like DMatrix) help reduce memory usage and make computations more efficient.
    • it’s cache-aware, so it knows how to make use of your cpu cache and speed up data retrieval.
    • xgboost handles datasets larger than memory through out-of-core computing.
    • distributed computing helps when your dataset is huge and you need to scale your work across multiple machines.
    • with gpu support, xgboost accelerates matrix calculations, making it much faster for large datasets.
  • Performance (Why this is different from other algos??)

    • regularization is a big win. it prevents overfitting with L1 and L2 regularization, keeping your model generalizable.
    • it automatically handles missing values, learning the best way to fill them in.
    • sparsity-aware split finding is great for sparse data (think text data or categorical features).
    • finding efficient splits in trees is another xgboost feature that speeds up training without compromising accuracy.
    • tree pruning helps reduce tree size after training, improving the model’s performance on unseen data.

Day 42: XGBoost for Regression and Classification, Catboost Vs XGboost Vs LightGBM

Notes on what i explored: Notes Notes Notes

  1. What if we have two or more than two features? We can scan all features , and do all possible splits for all features, then we will calculate gain and similarity score , and select feature which has max gain, algo use greedy search

  2. What if multiple feature and second feature is categorical (like binary: yes/no, male/female or muticlass: colors)? You have to encode them using OHE or other. or use othre variants of GDboost. idenntify unique values in catagorical columns , and consider both as a potential split point ., then calculate gain and similarity scores ., and select with maximum gain .

  3. What if the feature is binary categorical or multiclass categorical? older versions: encode them (e.g., one-hot, label, or target encoding). newer versions: native support for categorical features—no encoding needed.

Catboost Vs LightGBM Vs XGBoost:

alt text

  • categorical features
    • catboost handles them natively—no encoding needed.
    • lightgbm/xgboost require encoding, but xgboost supports label encoding in newer versions.
  • speed
    • lightgbm is the fastest, especially for large datasets.
    • catboost is slower but optimized for small/mid datasets.
    • xgboost is slower than both due to its exhaustive computations.
  • memory usage
    • lightgbm uses the least memory.
    • catboost and xgboost consume more, especially xgboost for large datasets.
  • overfitting prevention
    • all three handle overfitting well, but catboost excels in datasets prone to overfitting due to ordered boosting.
  • use case
    • catboost: small/mid datasets with many categorical features.
    • lightgbm: large datasets where speed and memory are critical.
    • xgboost: general-purpose, robust across various dataset types.

Day 43: Stacking Ensemble: Understanding Blending and K-fold methods

~ blending: splits the data into a training set and a holdout set to train base models and then a meta-model on the predictions of the base models.~

~ k-fold stacking: uses cross-validation to generate predictions for the meta-model by training base models on different training folds and predicting on the validation fold. ~

Notes: Notes:

!Hyperparameters Tuning:

  • identify hyperparameters: for decision trees: max_depth, min_samples_split, etc. for neural networks: learning rate, batch size, number of layers, etc.
  • choose a tuning method: grid search: try every combination of parameters (computationally expensive). random search: test random combinations (faster). bayesian optimization: automatically find the best parameters using probabilistic methods. automated tuning: libraries like optuna or hyperopt.
  • cross-validate: use k-fold cross-validation to test parameter combinations.
  • pick the best: finalize hyperparameters that minimize error metrics or maximize accuracy. alt text

1. manual tuning

  • tweak hyperparameters based on experience or trial-and-error.
  • works for simple models but not scalable.
  • systematically tests all combinations of hyperparameter values.

  • pros: exhaustive, finds the best combo (if time permits).

  • cons: computationally expensive, impractical for large spaces.

  • example:

    from sklearn.model_selection import GridSearchCV
    param_grid = {'max_depth': [3, 5, 10], 'min_samples_split': [2, 5, 10]}
    grid_search = GridSearchCV(DecisionTreeClassifier(), param_grid, cv=5)
    grid_search.fit(X_train, y_train)
    print(grid_search.best_params_)
    
    
  • samples random combinations of hyperparameters.

  • pros: faster, effective for large search spaces.

  • cons: might miss optimal combinations.

  • example:

    from sklearn.model_selection import RandomizedSearchCV
    from scipy.stats import randint
    param_dist = {'max_depth': randint(3, 20), 'min_samples_split': randint(2, 20)}
    random_search = RandomizedSearchCV(DecisionTreeClassifier(), param_dist, n_iter=100, cv=5)
    random_search.fit(X_train, y_train)
    print(random_search.best_params_)
    
    

4. bayesian optimization

  • predicts the best hyperparameters using probabilistic models (e.g., Gaussian processes).
  • pros: efficient in high-dimensional spaces.
  • cons: complex implementation.
  • libraries: Optuna, Hyperopt, BayesSearchCV.

5. evolutionary algorithms

  • optimizes hyperparameters using natural selection (e.g., genetic algorithms).
  • library: TPOT.

6. automated tools

  • auto-sklearn: automates model selection + hyperparameter tuning.
  • H2O.ai: distributed hyperparameter tuning.
  • Optuna: fast, user-friendly library for hyperparameter search.

Link to the blogpost


Day 44: K Nearest Neighbor, Coding KNN from scratch and applying on different datasets:

Predictions are based on the majority vote (classification) or average (regression) of the K closest data points in the training set.

Step-by-Step Workflow:

  1. Choose K: Number of neighbors to consider.
  2. Calculate Distance:
    • Common metrics: Euclidean (default), Manhattan, or Minkowski.
  3. Find K Nearest Neighbors: Identify the K points closest to the query.
  4. Make Prediction:
    • Classification: Majority class among neighbors.
    • Regression: Average value of neighbors.

3. Choosing the Right K

  • Small K (e.g., K=1): High variance, sensitive to noise (overfitting).
  • Large K (e.g., K=50): High bias, smoother boundaries (underfitting).
  • Rule of Thumb: Start with K=nK=n (where nn = number of samples) or use cross-validation.

Overfitting and Underfitting: alt text alt text

Code of KNN using Python: alt text

Finally After Hyperparameter tuning, KNN models imporves for California House price prediction: alt text

Notes: Notes:


Day 45: Support Vector Machines:

  1. What is SVM? Goal: SVM finds the "best" hyperplane to separate data into classes.

Key Idea: Maximize the margin (distance between the hyperplane and the nearest data points, called support vectors).

  • svm helps classify data by finding the best dividing line (hyperplane).
  • support vectors: closest points to the line that influence its position. kind of like the apples and oranges closest to the ruler.
  • margin: the gap between the line and the nearest points. svm tries to make this as wide as possible for better separation. Svm
  • hard margin svm: works only when data is clean and separable. not great for noisy or messy data.
    • example: classifying perfectly labeled "cat" vs. "dog" images where there’s no overlap.
  • soft margin svm: allows some mistakes for better flexibility with noisy/overlapping data.
    • example: separating spam and non-spam emails, where some emails are hard to classify.
  • kernel trick: useful when the data isn’t linearly separable. it projects data into a higher dimension to make it separable.
    • example: in handwriting recognition, svm can map curvy letters into a higher space to separate them more easily.

More Notes like optimization regularization: Notes


Day 46: K-Means Clustering, DBSCAN

K-Means (centeroid based):

Algorithm Steps

  1. Initialize Centroids: Randomly select K data points as initial centroids(Can lead to suboptimal clusters.) (or use k-means++(Distributes initial centroids to improve stability and speed) for smarter initialization).
  2. Assign Points to Clusters: For each data point, compute the Euclidean distance- used by default to all centroids. Assign the point to the nearest centroid.
  3. Recalculate Centroids: Compute the mean of all points in each cluster to update centroids.
  4. Repeat: Reassign points and update centroids until:
    • Centroids stabilize (change < tolerance threshold).
    • Maximum iterations are reached.

Kmeans

Applications

  • Customer Segmentation
  • Image Compression
  • Document Clustering
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

# Preprocess data
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Fit K-means
kmeans = KMeans(n_clusters=3, init='k-means++', n_init=10)
kmeans.fit(X_scaled)

# Get labels and centroids
labels = kmeans.labels_
centroids = scaler.inverse_transform(kmeans.cluster_centers_)

Choosing the Optimal K

  1. Elbow Method: Plot inertia vs. K; the "elbow" (point where inertia decline slows) suggests optimal K.
  2. Silhouette Score: Measures how similar a point is to its cluster vs others. Higher score (closer to 1) = better clustering.
  3. Domain Knowledge: Use prior understanding of the data to guide K selection (e.g., customer segments in marketing).

Notebook: K-Means clustering Demo

One limitation of K-means is that you must specify the number of clusters beforehand.

DBSCAN (density based clustering)

Visualize: DBSCAN here

Why DBSCAN

Notebook: DBSCAN demo

Groups data into density-based clusters (arbitrary shapes) and flags outliers.

  • Core Point: Has ≥ min_samples neighbors within radius eps.

  • Border Point: In a core point’s neighborhood but lacks enough neighbors.

  • Noise: Neither core nor border. alt text

  • eps: Use a k-distance plot (k = min_samples) to find the “knee” for optimal ε.

  • min_samples: Start with 2 * data dimensions (adjust for noise tolerance).

Steps

  1. Pick Parameters: eps (radius) and min_samples (density threshold).
  2. Expand Clusters:
    • For each unvisited point, check if it’s a core point (enough neighbors in eps).
    • If yes, form a cluster by adding all density-reachable points (core neighbors + their neighbors).
    • Mark non-core reachable points as border.
  3. Label Noise: Points not assigned to any cluster.
StrengthsWeaknesses
Finds any cluster shapeStruggles with varying densities
No need for KSensitive to ε and min_samples
Robust to outliersPoor performance in high dimensions

When to Use?

  • Data has noise or complex shapes (e.g., geospatial data, anomaly detection). dbscan vs k means
  • Avoid if clusters have highly varying densities (use HDBSCAN instead).

Day 47: Hierarchical clustering, Silhouette score

Hierarchical Clustering

Hierarchical clustering builds a tree-like hierarchy (dendrogram - a tree diagram showing how clusters merge/split. height shows the distance of merging, and cutting at a height defines the number of clusters.) of clusters. Two main approaches: alt text

  1. Agglomerative (Given below): Start with each point as its own cluster, iteratively merge closest clusters.
  2. Divisive (Reverse of Agglomerative): Start with all points in one cluster, recursively split into smaller clusters.

Agglomerative Clustering Steps:

  1. Initialize: Treat each data point as a singleton cluster.
  2. Compute Distance Matrix: Measure pairwise distances (e.g., Euclidean, Manhattan).
  3. Merge Clusters: Combine the two closest clusters; update the distance matrix.
  4. Repeat: Continue merging until one cluster remains.
ProsCons
No need to specify cluster count upfrontComputationally heavy (O(n³) time, O(n²) space)
Dendrograms aid visualizationSensitive to noise/outliers
Works well for small/mid-sized dataMerges are irreversible (local optima)

Linkage Criteria

Defines how distances between clusters are calculated:

  • Single Linkage: Minimum distance between clusters (prone to "chaining").
  • Complete Linkage: Maximum distance (creates compact clusters).
  • Average Linkage: Average distance between all pairs. alt text
  • Ward’s Method: Minimizes variance when merging (minimizes total within-cluster variance).

Silhouette Score

A metric to evaluate clustering quality by measuring how similar a data point is to its own cluster (cohesion) compared to other clusters (separation). silhouettescore

  • Range: −1 (poor clustering) to +1 (well-defined clusters).

Notes


Day 48: Dimensionality Reduction

Reduces the number of features in data while retaining meaningful patterns, addressing noise, computational cost, and visualization. Methods are hierarchically grouped into feature selection (keeping relevant features) and feature extraction (creating new features).

Dimensionality Redduction Hierarchy: alt text Feature Selection: alt text

Dimensionality Reduction of a Data: alt text

  • Curse of Dimensionality: The curse of dimensionality refers to the challenges and inefficiencies that arise when analyzing data in high-dimensional spaces, such as sparsity, computational complexity, and loss of meaningful patterns.

alt text fig: As the dimensionality of data increases, the feature space becomes sparser, and the data is easier to separate. This is the curse of dimensionality in a nutshell.

Notes:

Day 49: PCA (Principle Component analysis), Implementing with MNIST dataset

Goal: reduce dimensions while preserving maximum variance.

fig: pca projecting 2d data into 1d pc. project illustration

Example: Reducing 10D data to 2D → Use top 2 eigenvectors (highest eigenvalues).

Key Concepts

  • Variance: Spread of data along a feature.
  • Covariance: Measures how two variables vary together.
  • Eigenvectors: Directions of maximum variance (principal components).
  • Eigenvalues: Magnitude of variance along eigenvectors.
  • Transformation: Project data onto new axes (eigenvectors) to reduce dimensions.

Steps →

Steps PCA

How to choose the Optimal PC??

  • select the optimal number of pcs by checking the cumulative explained variance, aiming to retain 90%-95% of the total variance, or by using the elbow method on the explained variance plot where additional pcs add minimal value alt text Notes: Notes

Visualizing MNIST dataset on 2D and 3D:

alt text alt text

Notebook: Applying PCA on MNIST dataset

Conclusion: With about 100 PCs, our model predicts an accuracy of approximately 96%. In comparison, other models like KNN predict around 97% because KNN can capture more complex patterns in the data.

Day 50: Visualizing and Comparing PCA, t-SNE, UMAP, and LDA + Revision with the course ML specialization:

Today was a bit hectic as I tried to understand and visualize various dimensionality reduction techniques. Here's a summary of my learnings and comparisons:

Final Thoughts

  • t-SNE: Captures local similarities well but sometimes distorts the global structure. t-SNE Visuals
  • UMAP: Faster than t-SNE and preserves both local and global relationships better.
  • LDA: Supervised technique, works best when class separation is important.

Comparisons

Comparisons

We explored four dimensionality reduction techniques for data visualization: PCA, t-SNE, UMAP, and LDA. We used them to visualize a high-dimensional dataset in 2D and 3D plots.

Note: It's easy to fall into the trap of considering one technique better than the other. At the end of the day, there's no perfect way to map high-dimensional data into low dimensions while preserving the entire structure. There's always a trade-off in the qualities each technique offers.

Resources

Comparison of Results on MNIST

TechniqueLocal StructureGlobal StructureSupervised?Example Result
t-SNEPreservedNot preservedNoTight clusters of "2"s and "7"s, but arbitrary spacing between clusters.
UMAPPreservedPartially preservedNoTight clusters of "2"s and "7"s, with meaningful spacing between clusters.
LDAPreservedPreserved (class separation)YesDistinct groups for each digit, optimized for classification.

The learning is so hectic today, so I decided to revise the concepts we studied today with the help of Andrew NG. Here's the revision:

  1. What is clustering?
    • Clustering is grouping data points into clusters where points in the same cluster are more similar to each other than to those in other clusters.
  2. K-means?
    • K-means is a clustering algorithm that partitions data into K clusters by minimizing the distance between data points and their respective cluster centroids.
  3. optimization objective?
    • The goal is to minimize the sum of squared distances between data points and their nearest cluster centroid.
  4. Lab?? done// Notebook: Lab Assignment

Day 51: Anomaly Detection:

What is Anomaly Detection?

  • Identifying rare data points (anomalies) that deviate significantly from the majority of the data.
  • Fraud detection, system failure prediction, intrusion detection, healthcare monitoring. alt text

Approaches for Anomaly Detection:

  • Gaussian Distribution: Flag data points outside ±3σ (99.7% of data). alt text
  • Z-Score: Z=(x−μ)σZ=σ(x−μ); if ∣Z∣>3∣Z∣>3, mark as anomaly.
  1. Unsupervised:
    • Clustering (K-means, DBSCAN): Anomalies lie far from cluster centers.
    • Isolation Forest: Randomly splits data; anomalies are easier to isolate.
    • Autoencoders: Neural networks that reconstruct input. High reconstruction error = anomaly.

Evaluation Metrics

  • Precision: % of detected anomalies that are real.
  • Recall: % of true anomalies detected.
  • F1-Score: Harmonic mean of precision and recall.
  • ROC-AUC: Measures trade-off between TPR (recall) and FPR.

Choose Algorithm:

  • Small data? Use statistical methods (Z-score).
  • Large data? Try Isolation Forest or Autoencoders.

Notes:

alt text alt text

  • Assuming data is normally distributed.
  • Ignoring temporal/spatial context (e.g., seasonal trends).
  • Not updating models as data evolves.

Day 52: collaborative filtering

Best Article (Working behind CF): https://medium.com/@ashmi_banerjee/understanding-collaborative-filtering-f1f496c673fd

Collaborative Filtering is a recommendation algorithm that considers the similarities between different users when recommending an item to another user.

Making Recommendations

  • Predict what a user might like based on similar users/items. alt text
  1. User-User CF: Find users like you → recommend what they liked.
  2. Item-Item CF: Find items similar to what you liked → recommend those.

the approach minimizes a regularized cost function using gradient descent, adjusting user and item parameters iteratively. the learning rate (α) controls step size, balancing convergence speed and stability. alt text

while effective, it suffers from the cold start problem (new users/items lack data) and sparsity issues (many missing ratings). despite these challenges, it remains a widely used technique in recommender systems.

Mean Normalization

  • Why? Handle users who rate everything too high/low.

Collaborative Filtering vs Content-Based Filtering

alt text

CFContent-Based
Uses user-item interactionsUses item features (e.g., text, genre)
Example: Netflix recommendationsExample: News articles recommended based on text keywords

Day 53: Project @ Football Players Market Value Prediction - Introduction and Planning

Inspired by the work of Youla Sozen

Every project starts with a problem or question. However, this project is different. it's all about having fun. As a football enthusiast, creating these kinds of projects is always enjoyable. The plan is straightforward, and I will implement it step by step.

Project plan

Although this is a fun project, i aim to ensure the following:

  • accurate player valuation: estimating market values to benefit clubs, agents, and investors.
  • transfer market efficiency: preventing overpayment, aiding negotiation, and optimizing resource allocation.
  • risk assessment & player development: evaluating investments while identifying young talents with high growth potential.
  • data-driven insights: supporting fairer contract negotiations and improving decision-making in fantasy football & betting.

this is a future plan, and i will work towards achieving these goals in the coming days.

Plan of Project

Tips for the project (crafted with deepseek):

Tips


Day 54: Project @ Football Players Market Value Prediction - Collecting Data (Scraping)

web scraping is a technique to collect data from the internet and convert it into a meaningful format, like a data frame, when direct downloads aren't available. in this project, i used the sofifa dataset. here's the main page of sofifa main page

Code to scrape data: alt text

Plan for cleaning Data: alt text


Day 55: Project @ Football Players Market Value Prediction - Cleaning Data

just finished a major data cleaning session for my project. went through steps like handling missing values, converting currencies, splitting combined columns (height/weight), and ensuring consistent data types. cleaned up outliers, removed duplicates, and made sure everything’s ready for the next phase: EDA and model training. feeling good with the progress 😁

Here's a final look: Notebook: Data Cleaning

Code: alt text

Day 56: Project @ Football Players Market Value Prediction - EDA

Notebook: EDA (With Complete Documentation)

I have mostly used Plotly to visualize as it is interactive and for such beginners like me, the visualization impact is powerful. as well as the codes are also easy to write.

overall dataset insights

  1. what are the top 10 most valuable players? alt text
  2. how does market value vary by position (e.g., are strikers more expensive than defenders)? alt text
  3. which teams have the highest average market value? alt text
  4. what’s the distribution of market values (is it skewed towards a few expensive players)? alt text
  5. how does age correlate with market value (are younger players generally worth more)? alt text

player attributes vs. market value

  1. how does a player’s overall rating affect their market value? alt text
  2. which individual attributes (e.g., pace, stamina, strength) correlate the most with market value?
  3. does international reputation (1-5 stars) impact market value?
  4. how do potential ratings compare to market value (are high-potential players priced higher)? alt text
  5. do physical attributes (height, weight, strength) play a role in market value? alt text

position-specific insights

  1. are attacking midfielders (CAM/CM) more valuable than defensive midfielders (CDM/CM)? alt text
  2. how does pace affect wingers' (LW/RW) market value?
  3. do goalkeepers follow the same market trends as outfield players? alt text

contract & transfer market impact

  1. does a player's contract end year affect their market value (e.g., do players with 1 year left have lower values)? alt text
  2. are players on loan priced differently compared to permanent squad members?

Day 57: Project @ Football Players Market Value Prediction - Feature Engineering: (Creating features, Transforming Features)

Notebook: Creating and Transforming Features

spent 6+ hours experimenting with feature engineering. ran into some challenges, but made progress:

  1. position-based features: grouped players into categories (attackers, midfielders, defenders, goalkeepers) with scores. will refine tomorrow.
  2. club-based features: switched from one-hot encoding (curse of dimensionality) to target encoding using the mean market value for each club.
  3. contract-based features: correlation with market value was neutral. most features, except age, overall, and potential, seemed less important.

pairplot

alt text


Day 58: Project @ Football Players Market Value Prediction - ML : (Linear Regression with Refined Features and deploying with Streamlit)

Just for fun: fun

Pairplots: alt text

So, before diving into feature engineering after cleaning the data, i tried out linear regression and got an r2 score of around 0.52. then, after applying some feature engineering and playing around with features, i ran the same model and got the r2 score up to 0.96 with only numerical features
today, i wasn’t fully happy with the result and got confused about feature selection and engineering. so i decided to convert all features into numerical, applied scaling and transformation, and reran the model , r2 score shot up to 0.97
Saved the model immediately, then tried deploying it with streamlit locally, with a user input form and inverse transformations (MOST HATED PART) . while the model’s still a work in progress and not perfect, i'm proud of what i’ve learned so far. next steps are all about finding the best model and getting it deployed with some solid predictions and managing the form with the backend properly.

alt text

Streamlit Preview: https://www.linkedin.com/posts/paudelsamir_day-58365-linear-regression-with-refined-activity-7294398498777501697-2vWi?utm_source=share&utm_medium=member_desktop


Day 59: Project @ Football Players Market Value Prediction - Complete Streamlit Setup for our first Model - Linear Regression

Yesterday, I set up streamlit for my project with some help from ai tools. deploiyng isn't my strong suit, and it got pretty hectic trying to nail down the format and inputs and converting to model inputs. had to leave it unclear


So todayt's goal was get the ui sorted and code the input transformation for the model within streamlit. i could've tinkered with other models and tuned them, but finishing what i started felt right. a simple linear regression is doing surprisingly well for my needs. of course, i'll explore other algorithms and fine tune for better accuracy soon. planning to scale up from 5,000 to 20,000 rows in my dataset. let's see
Here's the demo where linear regression predicts player market values quite accurately. grabbed data from the site i scraped so that to visuailze properly. and its around 90 percent accurate for all the positions. that's already great !! Loving it

Here are some previewsL:

  • KDB (Real Vs Predicted) alt text alt text

  • Lamine (Real Vs Predicted) alt text alt text

  • Oblak (Real Vs Predicted) alt text alt text


Day 60: Project @ Football Players Market Value Prediction - Testing Ridge, Lasso, and Decision Trees

Today was fun! Started by handling outliers for the linear regression model, but didn’t see any improvement, so no luck there. Then I dove into applying PCA for dimensionality reduction. After converting everything to numerical features and applying all the feature engineering, I reduced the features from 50 to 40, and guess what? Model accuracy jumped to 99%! But here’s the twist, I can’t use this model for my project since I’m limited with deployment knowledge, and reverse transforming features while predicting is still the most hectic part of the process.

Applying PCA with 40 features

Then I tried Ridge and Lasso regression. Ridge performed the best and outperformed Linear and Lasso, so I’ll stick with Ridge for now until a simpler model comes along.

Ridge Regression Lasso Regression

Next up was Decision Tree Regressor. Applied it, and without hyperparameter tuning, I got around a 0.98 R2 score. I know Decision Trees are prone to overfitting, so I visualized, but couldn’t predict by myself. Decided to try hyperparameter tuning, but the results weren’t drastically different.

Decision Tree Regressor

At this point, the Decision Tree is the best model for the project. Let's see what’s coming next. Good luck, city. 𝐒𝐡𝐮𝐯𝐚𝐫𝐚𝐭𝐫𝐢 🌙 GridSearchCV Decision Tree

Decision Tree Insights

Did i just wasted 2 hours?? 😅😅Fun

Notebook: Experimentation 1


Day 61: Project @ Football Players Market Value Prediction -Had to hit reset from Feature Engineering

Just 20 minutes ago, i realized i’ve been making a huge mistake since day 4 with feature engineering. i found out today that as a beginner, it’s easy to mess up, but it's all part of the learning process. the mistkae was thinking about how to transform features back for deployment without realizing that features like overall rating, best oiverall, and potential are actually super correlated with market value. i was happy with the 99% accuracy, but i didn’t see the problem until now. i knew about overfitting can cause it and tried to fix it, but i never thought about visualizing feature importance and how it can affect the model.


So, iwas almost done with the project and deployed it to my local server using streamlit. feeling soooo dumb at the moment. the day before yesterday, i tested it with real players like KDB, Lamine Yamal, and Oblak, and 𝐭𝐡𝐞 𝐩𝐫𝐞𝐝𝐢𝐜𝐭𝐢𝐨𝐧𝐬 𝐥𝐨𝐨𝐤𝐞𝐝 𝐠𝐨𝐨𝐝. 𝐰𝐡𝐲?? 𝐛𝐞𝐜𝐚𝐮𝐬𝐞 𝐢𝐭 𝐭𝐨𝐭𝐚𝐥𝐥𝐲 𝐫𝐞𝐥𝐲𝐢𝐧𝐠 𝐨𝐧 𝐣𝐮𝐬𝐭 𝐭𝐡𝐫𝐞𝐞 𝐢𝐧𝐩𝐮𝐭𝐬 𝐁𝐞𝐬𝐭 𝐎𝐯𝐞𝐫𝐚𝐥𝐥, 𝐎𝐯𝐞𝐫𝐚𝐥𝐥 𝐑𝐚𝐭𝐢𝐧𝐠, 𝐚𝐧𝐝 𝐏𝐨𝐭𝐞𝐧𝐭𝐢𝐚𝐥. 𝐲𝐨𝐮 𝐜𝐚𝐧 𝐞𝐯𝐞𝐧 𝐩𝐫𝐞𝐝𝐢𝐜𝐭 𝐰𝐡𝐨𝐥𝐞 𝐭𝐡𝐢𝐧𝐠𝐬 𝐰𝐢𝐭𝐡 𝐣𝐮𝐬𝐭 𝐭𝐡𝐞𝐬𝐞 𝐢𝐧𝐩𝐮𝐭𝐬. The impact of other features is literally minimal. now, i have to rebuild the model from scratch again with better feature engineering.
Even after this dumbest mistake, total wasted grinding, wasted time, wasted energy---for the first time in my learning journey, it feels like now i'm actually learning something.

Context


Day 62: Project @ Football Players Market Value Prediction - Finalizing Project and Deploying it

Finally, i’m able to go live with my first end-to-end ml project! 🎉 you can check it out here: https://paudelsamir.streamlit.app/

  • Today was all about polishing things i messed up, after some solid feature engineering, i was able to hit 85% accuracy, pretty solid start. then, i played around a hour with hyperparameter tuning, which bumped it up to 89%.

actual vs predicted

alt text

but the real magic happened when i tried ensemble learning techniques. after a bit of back and forth, gradient boosting took me all the way to 94% accuracy. and will be using the same algorithm for deployment too.

And with that, after 10 days of nonstop grinding, i’m officially closing this project. it’s been a fun ride, full of learning and surprises !!

𝐆𝐢𝐭𝐇𝐮𝐛 𝐑𝐞𝐩𝐨 For the Project: https://github.com/paudelsamir/ML-Based-Football-Players-Market-Value-Prediction

Day 63: Content-Based Movie Recommender System - Preprocessing

Today, I focused on the preprocessing phase of building a content-based movie recommender system. I created a tags feature by combining key keywords from columns like genres, descriptions, top 3 cast members, and crew, especially the director. This step was crucial to ensure that the recommendation engine has a rich set of features to work with. Notes:


Day 64: Content-Based Movie Recommender System - Building and Deployment

Today, I built the recommendation engine based on movie content similarity using vectorization (bag of words). I also deployed it with Streamlit, so now you can input a movie name and get the top 5 similar movies based on the similarity matrix. Additionally, I integrated an API to pull movie posters in real-time from the website TMDB!

𝐂𝐡𝐞𝐜𝐤 𝐨𝐮𝐭 𝐡𝐞 𝐥𝐢𝐯𝐞 𝐝𝐞𝐦𝐨 𝐡𝐞𝐫𝐞: https://lnkd.in/d7R3Wsnk


Day 65: Diving into Deep Learning

Explored deep learning concepts, including its significance, how it differs from machine learning, and whether it will replace ML. Covered key architectures like Feedforward Neural Networks (FNNs), Convolutional Neural Networks (CNNs) for image processing, Recurrent Neural Networks (RNNs) for sequential data, Autoencoders for feature learning, and Generative Adversarial Networks (GANs) for data generation. Notes from the day: Notes Notes

  • Types of Neural network: alt text

Day 66: Perceptrons

Today, I dived into the concept of perceptrons, which are the building blocks of neural networks. I explored the perceptron algorithm, its working mechanism, and how it can be used for binary classification tasks. A supervised learning algorithm used for binary classifiers.

  • Steps in Prceptron Algorithm: alt text

  • Perceptron from Scratch: alt text

Visuals:

  • Training Data: alt text
  • Perceptron Training: alt text

I explored the fundamentals of MLOps with this paper: Machine Learning Operations (MLOps): Overview, Definition, and Architecture : https://arxiv.org/pdf/2205.02302


Day 67: Perceptron Loss Function and Gradient Descent

Today, I explored the Perceptron Loss Function, which helps adjust weights when misclassification occurs, ensuring better decision boundaries. I learned how the perceptron updates its weights using the weight update rule and how Gradient Descent optimizes the loss function by iteratively moving in the direction of the negative gradient.

alt text alt text

Notes: alt text


Day 68: Multilayer Perceptron

The problem with Perceptrons lies in their limitation to learn complex patterns and functions, especially those that are not linearly separable. A Perceptron is a single-layer neural network with binary outputs, and it can only solve problems where the data points are linearly separable. If the data is not linearly separable, a Perceptron cannot converge and find a solution. alt text

So the solution is Multilayer Perceptron: GIF

There's a website named: https://playground.tensorflow.org/

I practiced different optimizations there some of them are,

  • Adding nodes to hidden layer
  • Adding nodes to input layer
  • Adding nodes to output layer (for multicalss)
  • Adding nuber of hidden layer

The conclusion is you can classify any type of problem within regression and classifcion by optimizing those nodes and others like activation and regularization.

Batch and Gradient Descent:

number of samples processed before updating model weights

alt text

  1. SGD (batch size = 1) → updates after each individual sample (one row at a time). noisy but good for escaping local minima.
  2. Mini-batch gradient descent (batch size = 16, 32, 64, etc.) → updates after a small group of samples (e.g., rows 1-100). balances speed and stability.
  3. Full-batch gradient descent (batch size = all samples) → updates after seeing the entire dataset. very stable but slow and memory-heavy.

each batch in mini-batch or full-batch contains multiple rows, and the loss is computed over those samples before updating weights.

so in SGD, you're updating weights after every single row (which makes it very random and noisy). in mini-batch, you take a chunk of rows, calculate gradients over that group, then update weights. in full-batch, you process all the rows at once and then update.

Notes: alt text


Day 69: MLP notation, Forward Propagation

today, i deepened my understanding of multi-layer perceptrons (MLPs), including their formal notation and the calculation of weights and biases for each layer. i also explored forward propagation and practiced matrix multiplication by manually constructing and multiplying matrices to intuitively follow the perceptron’s computations. alt text

additionally, i studied MLP training with pytorch from the book deep learning with python by françois chollet.

additionally, i explored Image processing with datacamp, here's image representation of what i learned today alt text


Day 70: Loss Functions for Deep Learning

  • Regression:

    • MSE : squares errors, punishes big mistakes more.
    • MAE : takes absolute difference, treats all errors equally.
    • Huber Loss: mix of mse & mae, good for outliers. alt text
  • Classificaiton:

    • Binary cross entropy : for yes/no classification (spam or not spam). alt text
    • Categorical Cross entropy : for multiple classes (yes/no/ maybe). alt text
    • Sparse Categorical cross entropy - same as categorical but works with integer labels.
    • Hinge Loss: used in SVMs, pushes correct class far from the wrong ones. alt text
  • Autoencoders / VAE loss:

    • KL divergence
  • GANs:

    • Discriminator loss : helps the discriminator tell real from fake.
    • Minmax Loss : generator tries to fool discriminator by minimizing its best-case performance.
  • Object Detection and segmentation loss:

    • interseciton over union loss
    • smoooth l1 loss
    • dice loss
  • Reinforcement loss:

    • policy gradient - rewards good actions.
    • Q-Learning loss : teaches agent to choose best long-term rewards.
    • proximal policy optimization
  • Custom losss function:

    • Perceptrual loss
    • combined loss functions: mix of different losses for better results (e.g., cross-entropy + dice loss).

Notes:

alt text alt text


Day 71: Deep Diving Backpropagation

I already explored backpropagation in andrew ng’s ml specialization course, but that was more of a surface level explanation just the mechanics of how it works.

Today, i’m diving deep. like, REALLY deep. i want an intuitive, mathematical understanding of backpropagation, not just the algorithmic steps. all my tracing and derivations are going into my handwritten notes. this is for intutive approach to understand Regression part.

Notes: alt text alt text

Tomorrow, i’ll probably implement backprop from scratch, test it on a proper dataset, and try to visualize what’s actually happening. the key question: how?

Also, i might challenge myself to explain why backprop works in my own words. maybe even turn it into an article.


Day 72: Implementing backpropagation for regression

Today was all about applying what i learned yesterday to code and visualizing backpropagation.

first, i created a toy dataset that looks like this:
alt text

Then, i wrote functions to implement backpropagation from scratch. after running the training loop, here’s what the final parameters looked like—no keras, no tensorflow, just raw python:
alt text

all the code is in my notebook:
Notebook: Backpropagation Regression

I also tried using keras' sequential api to train the same model. after around 700 epochs, the error dropped significantly.
Keras

final weights with keras:
alt text

PS: intentionally chose a confusing dataset to mess with my own head.


Day 73: Implementing Backpropagation for Classification

i already implemented backprop for regression, both handwritten and in code. today, i'm tweaking it for classification.

few things to change:

  • loss function is binary cross-entropy instead of MSE
  • activation function is sigmoid instead of linear

but the backprop algo stays the same. all derivatives are now based on the new log function. since i already get the intuition, i'm skipping the math and just coding it.

Notebook: implementation backprop classification

sample data looks like this:
sample data

final parameters after training:
final parameters

function to update parameters:
parameter update function

at last, i tried implementing the same using tensorflow to see how it compares to my scratch implementation:
tensorflow code


Day 74: Revising old days, Memoization

Today, i revised concept of gradient and derivatives, focusing on how subtracting the gradient term is helping minimize loss. gradient relies on derivatives to find the optimal weights, and the learning rate controls the step size too high can cause overshooting, while too low leads to slow convergence. another key takeaway was memoization, a technique to store previously computed values to optimize calculations. in neural networks, repeated derivative computations can slow down training, and memoization helps speed things up by avoiding redundant calculations. this approach is widely used in dynamic programming and can improve efficiency in deep learning models.

Notes: alt text alt text


Day 75: Vanishing Gradient, Exploding Gradient

Today i first revised Gradient descent in NN:

  • Batch : faster to complete epochs ( batch size = all)

  • Stochastic : faster to converge ( batch size = 1)

  • Mini- Batch : mostly suitable ( batch size around center)

  • vanishing gradient: when gradients become too small, causing early layers to learn very slowly or not at all. example: in deep networks using sigmoid activation, earlier layers stop updating because gradients shrink to near zero.

  • exploding gradient: when gradients become too large, leading to unstable updates and divergence. example: in rnn training, weights keep multiplying large gradients, causing values to explode to infinity.

problems

How to Handle VGD problem:

  • use better activation functions – replace sigmoid/tanh with relu, leaky relu, or elu. Using Sigmoid: alt text Using ReLU: alt text This is the final weights comparision: alt text
  • use proper weight initialization – xavier/glorot for sigmoid/tanh, he initialization for relu.
  • batch normalization – normalizes activations to maintain stable gradients.
  • residual connections (skip connections) – used in resnets to allow gradients to flow easily.
  • gradient clipping – caps gradients to prevent them from becoming too small.

Day 76: Implementing artificial neural networks (ann) for different datasets

  • Experimented with ann on two datasets: mnist for handwritten digit classification and a gre dataset for graduate admission prediction. the goal was to train models and analyze performance across different domains.

Notebook: GRE prediction
Notebook: MNIST Classification

  1. trained an ann on mnist to classify handwritten digits (0-9) using a sequential model with dense layers. training ran for 30 epochs with promising results.
    • training results (30 epochs):
      mnist training
    • model summary:
      mnist model summary
    • sample predictions:
      mnist prediction
  • applied ann to predict graduate school admission chances based on gre scores, gpa, and other factors. dataset required preprocessing before feeding into the model. trained for 100 epochs.
    • sample dataset:
      gre dataset sample
    • training results (100 epochs):
      gre training
    • model summary:
      gre model summary
    • model accuracy evaluation:
      gre model accuracy

next steps: hyperparameter tuning, dropout layers for regularization, and testing on additional datasets.

Day 77: Improving Neural Networks

Notes: notes

Fine-tuning Neural Network Hyperparameters

Fine-tuning neural network hyperparameters is about adjusting key settings to improve learning

  • Learning rate - controls how fast the model updates. Too high = unstable, too low = slow learning.
  • Batch size - number of samples processed before an update. Small = noisy but frequent updates, large = stable but slow.
  • Epochs - full passes through data. Too few = underfitting, too many = overfitting.
  • Layers & neurons - more can improve learning but make training harder.
  • Activation function - decides neuron output; common ones are ReLU, sigmoid, tanh. alt text
  • Dropout - turns off some neurons randomly to prevent overfitting.
  • Optimizer - algorithm that adjusts weights efficiently (Adam, SGD, etc.).

Problem Solving Strategies

  • Vanishing/exploding gradient – gradients shrink or blow up, stopping learning → use ReLU, batch norm, weight init, or residual connections.
  • Not enough data – small datasets cause poor generalization → apply data augmentation, transfer learning, or synthetic data generation.
  • Slow training – long training times due to large models or bad optimizers → use mini-batches, better optimizers, mixed precision, and GPUs.
  • Overfitting – model memorizes training data but fails on new data → apply dropout, regularization, early stopping, and data augmentation.
  • Underfitting – model is too simple and fails to learn patterns → increase complexity, train longer, improve features, and reduce regularization.
  • Imbalanced data – one class dominates, leading to biased predictions → use class weighting, oversampling, or synthetic data (SMOTE/GANs).
  • Poor generalization – model does well in training but fails on real-world data → ensure diverse data, reduce leakage, use domain adaptation, or adversarial training.
Transfer Learning

alt text alt text

Day 78: Sequence Modeling / RNNs - Just Overview

I just thought ki Before diving deeper into neural network improvement techniques, I should first gain a surface-level understanding of deep learning concepts I'll be tackling in the future as this learning technique is helping me alot. For this purpose, I found an excellent YouTube playlist: MIT 6.S191: Introduction to Deep Learning. There are approximately 10 to 15 videos that I plan to watch to build a foundational overview and i too will deep dive into these later

alt text alt text

Implementation Preview alt text alt text


Day 79: Transformers , Attention - Just Overview

Transformers replace RNNs by using self-attention, enabling parallel processing and handling long-range dependencies efficiently. Introduced in "Attention Is All You Need" (2017), they power models like BERT and GPT.

Self-attention computes Query (Q), Key (K), and Value (V) matrices to determine word relationships. Multi-head attention allows the model to capture different contextual meanings.

The transformer consists of encoder-decoder blocks with self-attention, feed-forward layers, and normalization. Encoders learn representations, while decoders generate sequences.

Transformers are used in chatbots, translation, search engines, and AI coding assistants. Key models include BERT (bi-directional understanding), GPT (text generation), and T5 (text-to-text tasks).

alt text alt text alt text


Day 80: CNNs - Just Overview Part 1

  • Computer vision enables machines to interpret visual data.
  • Used in self-driving cars, medical imaging, surveillance, AR.

What Computers "See"

  • Images are matrices of pixel values (grayscale: single matrix, RGB: 3 matrices).
    alt text

Feature Extraction & Convolution

  • CNNs learn features like edges, textures, and shapes automatically.
  • Convolution: uses filters (kernels) to extract patterns from images.
    alt text
  • Example Kernel (Edge Detection - traditional filter): alt text

CNN Architecture

  1. Convolution Layers: detect patterns.
  2. ReLU Activation: makes model non-linear.
  3. Pooling Layers: reduce size, keep key features.
  4. Fully Connected Layers: classify objects.
    alt text

Object Detection

use yolo when speed matters more than precision (e.g., real-time apps). use rcnn when accuracy is critical and speed isn’t a constraint. use faster rcnn for a balance between accuracy and speed. alt text

Self-Driving Cars

  • CNNs help detect lanes, pedestrians, and traffic signs.
  • End-to-end models predict steering angles using video input.
    alt text alt text

Day 81: CNNs - Just Overview Part 2

yesterday, i watched a video on cnn. the goal was just to explore it for a day, but i feel like this is an interesting topic. so today, i want to learn more—like, in-depth—about the feature extraction part, which i find the most interesting aspect of cnn.

i learned to use relu and understood convolution layers and how they work yesterday. but today, i learned about pooling and how it reduces the size. i also explored the classification process in more depth.

notes: alt text note: cnn by itself doesn't handle rotation and scaling well. for that, use data augmentation.


Day 82: Deep Generative modeling - Just overview

Watched this single video : https://www.youtube.com/watch?v=Dmm4UG-6jxA&t=3242s

today’s deep dive into generative models gave me a solid grasp of how ai can not only recognize patterns but also create new data from scratch. these models are the backbone of modern generative ai, and understanding them is key to keeping up with the field.

Generative models: what they do

  • they don’t just classify or analyze data; they generate entirely new data that resembles what they learned from.
  • examples include autoencoders, variational autoencoders (VAEs), GANs, and diffusion models—each with its own strengths. alt text

Latent variable models: finding the hidden structure

  • autoencoders learn to compress data into lower-dimensional representations and then reconstruct it.
  • vaes take this further by adding randomness, making them better at generating diverse outputs instead of just memorizing patterns.
  • plato’s cave analogy clicked here—observed data is like shadows on a wall, while latent variables represent the actual objects casting those shadows. alt text

Generative adversarial networks (GANs): competition makes better results

  • a generator creates fake data, while a discriminator tries to catch the fakes. alt text
  • through constant feedback, both improve, leading to highly realistic outputs.
  • training is tricky—imbalanced learning can cause mode collapse, where the generator keeps making similar outputs instead of diverse ones. alt text

Regularization & structure in vaes

  • forcing the latent space into a structured distribution (often gaussian) helps ensure smoothness and continuity, making vaes more useful for controlled generation. alt text

Applications beyond images

  • these models aren't just for ai art—they power speech synthesis, text generation, and even domain adaptation (like CycleGAN for translating images without paired data).

Diffusion models: a new generative powerhouse

  • instead of generating images all at once (like GANs), diffusion models gradually refine noise into meaningful images. alt text
  • they’re more stable and produce higher-quality results, making them the future of generative ai.

Day 83: Reinforcement Learning - Just Overview

with the help of this video: https://www.youtube.com/watch?v=8JVRbHAVCws&t=3242s

  • explored key concepts of reinforcement learning (RL), including Q-learning and policy learning algorithms. alt text

  • dived into Deep Q Networks (DQN) and their role in handling complex environments, like Atari games. alt text

  • understood the difference between discrete and continuous actions and how they impact RL models. alt text

  • learned about real-world applications of RL, from robotics to game AI, and cutting-edge technologies like AlphaGo and MuZero. alt text

  • explored training techniques like policy gradients and how they improve decision-making in RL agents. alt text alt text alt text alt text


Day 84: Deep learning: challenges & new frontiers - Just Overvview

Video Link: https://www.youtube.com/watch?v=N1fbskTpwZ0&t=3021s

alt text

  • neural network failure modes – ai can fail unpredictably, often overconfident in wrong answers. alt text
  • uncertainty in deep learning – models lack awareness of their own mistakes, leading to unreliable predictions. alt text
  • adversarial attacks – tiny changes in input can trick ai into making completely wrong decisions. alt text
  • algorithmic bias – ai inherits and amplifies biases from training data, leading to unfair outcomes. alt text

generative ai & diffusion models

  • the landscape – diffusion models are replacing GANs in high-quality image generation.
  • diffusion process – gradually add noise to data and train a model to reverse it.
  • noising & denoising – AI learns to reconstruct images by stepwise noise removal. alt text
  • text to image, beyond images – models like DALL·E generate images, but AI is expanding into text, music, and 3D.

large language models (LLMs)

  • using llms to generate text – models predict words to generate human-like text.
  • limitations – hallucinations, bias, and lack of real understanding. alt text
  • more parameters = better performance, but higher computational cost.
  • how they work – transformers + attention mechanism process and generate context-aware text

Day 85: Early Stopping, Normalizing Inputs, Dropout

These are the techniques i will cover in upcoming days:

  1. Vanishing Gradients

    • Activation Functions
    • Weight Initialization
  2. Overfitting

    • Reduce Complexity/Increase Data
    • Dropout Layers
    • Regularization (L1 & L2)
    • Early Stopping
  3. Normalization

    • Normalizing inputs
    • Batch Normalization
    • Normalizing Activations
  4. Gradient Checking and Clipping

  5. Optimizers

    • Momentum
    • Adagrad
    • RMSprop
  6. Learning rate scheduling

  7. Hyperparameter Tuning

    • No. of hidden layers
    • Nodes/layer
    • Batch size

Today, I explored Early Stopping, Normalizing Inputs, and Dropout techniques for improving neural network performance.

  • Early Stopping: Prevents overfitting by halting training when validation performance drops.
  • Normalizing Inputs: Scales features to a consistent range, aiding in faster and more stable learning.
  • Dropout: Reduces overfitting by randomly deactivating neurons during training, forcing the model to generalize better.

Notebook: Dropout on Regression alt text Notebook: Dropout on Classification alt text


Day 86: Regularization, Quantization

L1 and L2 regularization are typically used for smaller networks. For larger networks, it is better to use neural network-specific regularization which is dropout regularization. alt text

An evaluation procedure must be used when using a regularizer to monitor that regularization process. For this, we can plot model performance against the number of epochs during the training process.

Regularization Regularization

Notebook: without Regularization vs Applying Regularization alt text

Quantization:

quantization in deep learning reduces the precision of numbers in a model to save memory and speed up processing. it can be done in two ways: post-training quantization (ptq), which converts the model to lower precision after training for faster performance but may lose some accuracy, and quantization-aware training (qat), where quantization is simulated during training, resulting in better accuracy but requiring more time. frameworks like TensorFlow provide tools for both methods to help deploy lighter and faster models. alt text

Day 87 - Activation Functions - Revisited

why needed? introduce non-linearity to capture complex patterns.

alt text alt text ideal properties: non-linear, differentiable, computationally inexpensive, zero-centered, non-saturating.

  • use relu + he init, sigmoid/tanh + xavier init.
  • relu in early layers, tanh or sigmoid in later/output layers.
  • prefer gelu for transformers.
  • avoid dying relu with leaky relu or prelu.

relu: fast, simple but can die. leaky relu: small slope for negatives, avoids dead neurons. prelu: learnable slope, more flexible. elu: better generalization, but expensive. selu: self-normalizing, good for deep nets. alt text


Day 88: Weight Initialization

Zero Init: All weights as zero → no learning (same gradients).

weights = np.zeros((input_size, output_size))

alt text

One Init: All weights as one → same issue, no symmetry breaking.

weights = np.ones((input_size, output_size))

✅ Random Init: Small random values.

weights = np.random.randn(input_size, output_size) * 0.01

✅ Xavier Init (for tanh/sigmoid):

weights = np.random.randn(input_size, output_size) * np.sqrt(1 / input_size)

alt text

✅ He Init (for ReLU):

weights = np.random.randn(input_size, output_size) * np.sqrt(2 / input_si

alt text

Notebook: Weight Initialization

Notebook: Xavier and He initialization

Notes: alt text alt text


Day 89: Deeep Learning Optimizers

  • Gradient Descent: Fundamental optimizer that updates weights by moving in the opposite direction of the gradient of the loss function w.r.t. the weights. It's basic but effective for many problems.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01)

bgd

  • Stochastic Gradient Descent: Optimizes with each sample instead of the entire dataset, which is faster but more noisy.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01)
  • Momentum: Momentum helps accelerate SGD by moving along relevant directions and dampening oscillations. It adds a fraction of the previous update to the current one.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01, momentum=0.9)

momentum

  • NAG: An improvement over momentum. It looks ahead by calculating the gradient not just at the current position but slightly ahead, giving better convergence.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01, momentum=0.9, nesterov=True)

nag

  • AdaGrad: Adapts the learning rate to the parameters, performing larger updates for infrequent parameters and smaller updates for frequent ones.
from tensorflow.keras.optimizers import Adagrad
optimizer = Adagrad(learning_rate=0.01)

adagrad

  • RMSProp: Divides the learning rate by an exponentially decaying average of squared gradients. Useful for non-stationary objectives.
from tensorflow.keras.optimizers import RMSprop
optimizer = RMSprop(learning_rate=0.001)

rmsprop

  • ADAM (Adaptive Moment Estimation): Adaptive learning rate method that combines the advantages of RMSprop and momentum. Popular for its efficiency.
from tensorflow.keras.optimizers import Adam
optimizer = Adam(learning_rate=0.001)

Adam

Overall: overall Notes: notes notes notes

Interactive Visualization of Optimization Algorithms in Deep Learning: https://emiliendupont.github.io/2018/01/24/optimization-visualization/


Day 90: Keras Tuner

hyperparameters (like learning rate, batch size, optimizer) directly impact model performance. tuning helps optimize accuracy and generalization.

in my case, i worked with the diabetes dataset and focused on tuning key hyperparameters like learning rate, batch size, optimizer, number of neurons, and dropout rate. each of these parameters influences different aspects of training—for example, the learning rate affects how quickly the model converges, while dropout helps prevent overfitting. alt text

to streamline the tuning process, i used keras tuner’s RandomSearch. i defined a hypermodel where parameters like the number of layers, neurons, dropout rates, learning rates, and optimizer types were set as tunable. the objective was to maximize validation accuracy. i also configured settings like max_trials to control the search space and executions_per_trial to ensure consistent evaluation. alt text alt text after running the tuning process, the model achieved around 79% accuracy. the tuning helped balance model complexity and performance, reducing overfitting and improving generalization. alt text for further improvements, i could fine-tune hyperparameters like the learning rate and dropout in smaller increments, try advanced optimizers like adamw, or implement early stopping to avoid unnecessary training once the model stops improving.


Day 91: Deep Diving into CNNs:

Focused on planning a deep dive into cnns. explored why anns fall short for cnn tasks and uncovered some fascinating cnn applications. along the way, stumbled upon some surprisingly cool ideas for future projects. fueled by that curiosity, i tried something basic today—simple, but a solid starting point.

today, i explored opencv from scratch, debugged image loading issues, applied basic filters using custom convolution kernels in pure python as well as with Opencv, and created amazingly undefinable custom filters with the excitement. alt text Notebook:Trying OpenCV for the first time Watch out notebook what i did with this lovely image: alt text

Notes: alt text

architectures

  • early arch
    • lenet-5
    • alexnet
    • vgg net (16/19)
  • modern arch
    • resnet (skip connection)
    • inception net (google net)
    • mobilenet
    • efficientnet

Plan after Architecture

  1. data augmentation (rotation, scaling, flipping)
  2. transfer learning (using pretrained models like vgg, resnet)
  3. object detection (yolo, r-cnn, ssd, etc.)
  4. image segmentation (u-net, mask r-cnn)
  5. attention mechanisms (squeeze, spatial channel-wise)
  6. self-supervised learning (contrastive, ssl architectures)

Day 92: Understanding Paddings and Strides

How CNNs working with Grayscale and Rgb images?? alt text alt text

padding and strides are important in convolutional neural networks (cnns) because they affect feature extraction, output size, and computational efficiency.

Padding is used to prevent reduction in spatial dimensions and retain edge information. for example, a 5x5 image with a 3x3 filter produces a 3x3 feature map, which keeps shrinking with more layers. adding padding helps maintain the size. alt text there are two common types of padding:

  • valid: no padding, meaning the output size shrinks.
  • same: pads the input so the output size remains the same.

the output size with padding is calculated as:

(n + 2p - f + 1) × (n + 2p - f + 1), where n is the input size, f is the filter size, and p is the padding amount.

in keras, padding is applied like this Demo for MNIST:

# Importing necessary libraries
import tensorflow as tf
from tensorflow.keras.layers import Dense, Conv2D, Flatten
from tensorflow.keras.models import Sequential
from tensorflow.keras.datasets import mnist

# Loading MNIST dataset
(X_train, y_train), (X_test, y_test) = mnist.load_data()

# Creating a Sequential model
model = Sequential()

# Adding Convolutional layers with valid padding
model.add(Conv2D(32, kernel_size=(3,3), padding='valid', activation='relu', input_shape=(28,28,1)))
model.add(Conv2D(32, kernel_size=(3,3), padding='valid', activation='relu'))
model.add(Conv2D(32, kernel_size=(3,3), padding='valid', activation='relu'))

model.add(Flatten())

model.add(Dense(128, activation='relu'))
model.add(Dense(10, activation='softmax'))

# Printing model summary
model.summary()

strides control how far the filter moves across the image at each step. a stride of (1,1) moves one pixel at a time, while higher strides skip pixels, reducing spatial dimensions and computation time. stride = 2 output size with strides is calculated as:

((n + 2p - f) / s + 1) × ((n + 2p - f) / s + 1), where s is the stride value.

higher strides help capture larger patterns but reduce spatial resolution. for example, a stride of (2,2) makes the filter shift 2 pixels at a time:

# Importing necessary libraries
import tensorflow as tf
from tensorflow.keras.layers import Dense, Conv2D, Flatten
from tensorflow.keras.models import Sequential
from tensorflow.keras.datasets import mnist

# Loading MNIST dataset
(X_train, y_train), (X_test, y_test) = mnist.load_data()

# Creating a Sequential model with strides
model = Sequential()

model.add(Conv2D(32, kernel_size=(3,3), padding='same', strides=(2,2), activation='relu', input_shape=(28,28,1)))
model.add(Conv2D(32, kernel_size=(3,3), padding='same', strides=(2,2), activation='relu'))
model.add(Conv2D(32, kernel_size=(3,3), padding='same', strides=(2,2), activation='relu'))

model.add(Flatten())

model.add(Dense(128, activation='relu'))
model.add(Dense(10, activation='softmax'))

# Printing model summary
model.summary()

in summary, padding helps retain image details and control output size, while strides affect computational efficiency and feature abstraction. tuning these parameters is key to optimizing cnns.

Pooling:

pooling is used to downsample feature maps, reducing their size while retaining important information. it helps prevent overfitting and reduces computation.

common types of pooling: alt text max pooling: selects the maximum value in a region. average pooling: takes the average of values in a region. for example, a 2x2 max pooling layer with stride 2 reduces feature maps to half their original size.


# importing necessary libraries
import tensorflow as tf
from tensorflow.keras.layers import Dense, Conv2D, Flatten, MaxPooling2D
from tensorflow.keras.models import Sequential
from tensorflow.keras.datasets import mnist

# loading mnist dataset
(X_train, y_train), (X_test, y_test) = mnist.load_data()

# reshaping input data
X_train = X_train.reshape(-1, 28, 28, 1).astype('float32') / 255
X_test = X_test.reshape(-1, 28, 28, 1).astype('float32') / 255

# creating a sequential model
model = Sequential()

# adding convolutional layers with padding
model.add(Conv2D(32, kernel_size=(3,3), padding='same', activation='relu', input_shape=(28,28,1)))
model.add(MaxPooling2D(pool_size=(2,2), strides=(2,2)))

model.add(Conv2D(64, kernel_size=(3,3), padding='same', activation='relu'))
model.add(MaxPooling2D(pool_size=(2,2), strides=(2,2)))

model.add(Conv2D(128, kernel_size=(3,3), padding='same', activation='relu'))
model.add(MaxPooling2D(pool_size=(2,2), strides=(2,2)))

# flattening and adding dense layers
model.add(Flatten())
model.add(Dense(128, activation='relu'))
model.add(Dense(10, activation='softmax'))

# printing model summary
model.summary()


Day 93: Backpropagation in cnns: a quick breakdown

cnns learn by adjusting their parameters through backpropagation. here's how:

  1. forward propagation:
    • input passes through convolution → relu → max pooling → flatten → fully connected layers → output.
  2. backpropagation:
    • calculate loss and propagate errors backward.
    • use the chain rule to compute gradients for weights, biases, and activations.
  3. layer-wise backprop:
    • convolutional layers: errors backpropagate through filters.
    • max pooling: only max-selected neurons contribute gradients.
    • flattening: reshapes gradients before sending them back.

Notes: alt text alt text


Day 94: LeNet5, Cat Vs Dog Classification

LeNet-5, introduced by Yann LeCun in 1989, is one of the earliest convolutional neural networks (CNNs), designed for handwritten character recognition. It consists of seven layers:

alt text

  1. Input Layer: 32×32 grayscale image.
  2. Conv Layer 1 (C1): 6 filters (5×5), output size 28×28×6. alt text
  3. Pooling Layer 1 (S2): Average pooling (2×2), output 14×14×6. alt text
  4. Conv Layer 2 (C3): 16 filters (5×5), selectively connected, output 10×10×16. alt text
  5. Pooling Layer 2 (S4): Average pooling (2×2), output 5×5×16. alt text
  6. Fully Connected (C5): 120 neurons, connected to all 5×5×16 inputs. alt text
  7. Fully Connected (F6): 84 neurons. alt text
  8. Output Layer: Softmax with 10 classes (digits 0-9). alt text

Implementation:

def build_lenet(input_shape):
  # Define Sequential Model
  model = tf.keras.Sequential()
  
  # C1 Convolution Layer
  model.add(tf.keras.layers.Conv2D(filters=6, strides=(1,1), kernel_size=(5,5), activation='tanh', input_shape=input_shape))
  
  # S2 SubSampling Layer
  model.add(tf.keras.layers.AveragePooling2D(pool_size=(2,2), strides=(2,2)))

  # C3 Convolution Layer
  model.add(tf.keras.layers.Conv2D(filters=6, strides=(1,1), kernel_size=(5,5), activation='tanh'))

  # S4 SubSampling Layer
  model.add(tf.keras.layers.AveragePooling2D(pool_size=(2,2), strides=(2,2)))

  # C5 Fully Connected Layer
  model.add(tf.keras.layers.Dense(units=120, activation='tanh'))

  # Flatten the output so that we can connect it with the fully connected layers by converting it into a 1D Array
  model.add(tf.keras.layers.Flatten())

  # FC6 Fully Connected Layers
  model.add(tf.keras.layers.Dense(units=84, activation='tanh'))

  # Output Layer
  model.add(tf.keras.layers.Dense(units=10, activation='softmax'))

  # Compile the Model
  model.compile(loss='categorical_crossentropy', optimizer=tf.keras.optimizers.SGD(lr=0.1, momentum=0.0, decay=0.0), metrics=['accuracy'])

  return model


Day 95: GPU slow than CPU - well in my case?

Today, I trained an MNIST model on both CPU (Ryzen 7 6000) and GPU (RTX 3050 Ti), expecting a significant speedup with the GPU. Instead, the CPU performed slightly faster, and when I tried adding a small CNN, my GPU environment crashed, while the CPU handled it fine (but slower).

  • The GPU initially took longer due to kernel warm up and data transfer overhead.
  • Batch Size Impact – GPUs perform best with larger batch sizes (e.g., 512+), while I used a smaller batch.
  • Data Bottlenecks – My CPU handled data preloading better, while the GPU might have suffered from inefficient memory access.
  • GPU Utilization – The GPU wasn’t fully utilized, likely due to suboptimal parallelism in my setup.

And When I added even a small convo layer, my GPU environment crashed, but the CPU ran it (slowly). here’s what possible reason i estimated?

  • My RTX 3050 Ti has 4GB VRAM, and it might be running out.
  • Despite setting up a dedicated GPU environment (with proper CUDA, cuDNN, and TensorFlow versions), I couldn’t debug this issue today.
  • There might be a compatibility issue between TensorFlow, CUDA, and my system's architecture.

Could it be a CUDA/cuDNN issue or VRAM exhaustion?

I used two separate virtual environments:

  • CPU environment (normal setup)
  • GPU environment (proper versions of CUDA, TensorFlow, and cuDNN) Still, I couldn’t debug the crashes today. Any suggestions on debugging the bit architecture issue or possible TensorFlow config problems? and is it abnormal CPU performing better than GPU in my optimization??

Here's the comparision: GPU vs CPU GPU vs CPU

Some rough notebooks: Notebook: CPU Notebook: GPU


Day 96: Data Augmentation, Pretrained Models

Data augmentation is crucial in machine learning, especially for tasks like computer vision, to enhance model performance and prevent overfitting. It involves applying various transformations to existing data, such as rotation, translation, scaling, flipping, shearing, zooming, and adjusting brightness and contrast. These techniques help in creating a larger and more diverse training dataset, thereby improving model generalization. alt text

Why Use Data Augmentation?

  1. Increased Dataset Size: Enhances model training with more examples.
  2. Regularization: Adds noise to prevent overfitting.
  3. Improved Generalization: Helps models perform better on unseen data.

Notebook: data augmentation on cifar10 frog

Pretrained models in CNN:

Notes: Notes

before deep learning took over, imagenet models relied on classical ml methods like svm, decision trees, and hand-crafted features (think hog, sift, and lbp). this worked, but scaling to millions of images? a nightmare. then alexnet (2012) happened—deep cnns trained with relu activations and dropout on gpus. it crushed traditional methods, slashing classification error rates by half. alt text alt text

vgg (2014) pushed deeper with 3x3 convolutions, proving that simplicity + depth = power. same year, googlenet (inception v1) introduced inception modules—parallel conv layers reducing parameter overhead while boosting efficiency. resnet (2015) then solved the vanishing gradient problem with skip connections, making ultra-deep networks (152 layers!) trainable.

today, pretrained models built on imagenet—resnet, vgg, inception, efficientnet—are the backbone of modern deep learning. explored these today, and yeah, standing on the shoulders of giants makes life easier.

PaperLink: ImageNet Classification with Deep Convolutional Neural Networks

Medium Article on alexnet

Keras pretrained models: https://keras.io/api/applications/ alt text

I will test the examples from that website using their code there in following notebook: Notebook: Pretrained Model Testing

side syb side predictions


Day 97: Visualizing Convolutional Layers, Transfer learning

Today, i'll be using the article : https://machinelearningmastery.com/how-to-visualize-filters-and-feature-maps-in-convolutional-neural-networks/ with respect to following topics.

  • Visualizing Convolutional Layers
  • Pre-fit VGG Model
  • How to Visualize Filters
  • How to Visualize Feature Maps

Notebook: visualizing Layers with elephant image

notes: Notes

Transfer learning:

alt text transfer learning helps train deep learning models efficiently by leveraging pre-trained networks like vgg, resnet, or mobilenet. instead of starting from scratch, we use the convolutional base (which extracts features) and replace the fully connected layers with our own classifier.

two main approaches:

  1. feature extraction – freeze the convolutional layers and train only the new classifier. useful when the target dataset is similar to the original dataset.
  2. fine-tuning – unfreeze some deeper layers and retrain them along with the classifier. this helps when the target dataset is quite different.

Resource: https://www.tensorflow.org/tutorials/images/transfer_learning

Day 98: Keras functional API

How to use keras functional api for deep learning?

Today i went through an article: https://machinelearningmastery.com/keras-functional-api-deep-learning/

The Sequential model API is great for developing deep learning models in most situations, but it also has some limitations.

I started with the Sequential API to build familiarity:

  • Architecture: Simple CNN with Conv2D → MaxPooling → Flatten → Dense → Output.
  • Workflow:
    • Loaded data, normalized pixels (0-1), reshaped images (28x28x1), and one-hot encoded labels.

    • Built a linear stack of layers:

      model = Sequential([
          Conv2D(32, (3,3),
          MaxPooling2D(),
          Flatten(),
          Dense(128),
          Dense(10, activation='softmax')
      ])
      
    • Trained with model.fit(), achieving ~91% validation accuracy in 10 epochs.


2. Transition to Functional API

alt text

I rebuilt the same model using the Functional API to see the syntax shift:

  • Input Layer: Explicitly defined with Input(shape=(28,28,1)).

  • Layer Connections: Layers are chained like functions:

    x = Conv2D(32, (3,3)(input_layer)
    x = MaxPooling2D()(x)
    ...
    
  • Model Definition: Declared inputs/outputs explicitly:

    model_func = Model(inputs=input_layer, outputs=output
    

Truncated — view the full README on GitHub.

100daysofcode
100daysofml
365dayschallenge
365daysofdata
ai
dailylearning
data
datascience
deeplearning
machinelearning
ml

paudelsamir/365DaysOfData

A year long journey with ai from data, exploring adjacent techs

Jupyter Notebook

50

960 commits

updated Dec 26, 2025

See the code

README

Last Updated Repo Size

365DaysOfData cover

[!note]
I'll share progress and demos on linkedin and twitter.
I won’t post daily or raw learns, updates will be for specific topics, concise, and focused on what i actually built or explored. plan is 4–5 posts per week.
This journey is about AI from scratch with data, not my entire learning history. I’ll keep building in public while also learning other adjacent techs beyond AI.

Projects Completed

ProjectsDescriptionDeployment
Football Players Market Value PredictionA 10-day end-to-end machine learning capstone project involving data scraping, cleaning, feature engineering, model training, and deployment. Achieved 94% accuracy using gradient boosting algorithms.Live Demo 👆🏽
Movie Recommender SystemAn end-to-end content-based movie recommender system leveraging a dataset of 5000 movies from Kaggle. Built with cosine similarity and TF-IDF vectorization.Live Demo 👆🏽
Cat vs Dog ClassifierA deep learning model leveraging VGG16 architecture, trained on an RTX 3050 Ti for 30 epochs, achieving 95% accuracy using the Kaggle Dogs vs Cats dataset.Live Demo 👆🏽
Guess The Footballer By EyesAn interactive game where users compete against AI to recognize 25 famous footballers by their eyes alone. Built with ResNet18 achieving ~70% accuracy. Features scoring system and streak tracking.Demo 👆🏽
Seq2Seq ChatbotA sequence-to-sequence chatbot trained on Cornell Movie-Dialogs Corpus using encoder-decoder architecture with Luong attention mechanism. Built from scratch in PyTorch.Live Demo 👆🏽
GPT from ScratchComplete implementation of GPT transformer architecture from scratch following Karpathy's tutorial. Includes bigram model, self-attention, multi-head attention, and complete transformer blocks.Notebook 📓
Image CaptioningAn end-to-end image captioning project using the Flickr8k dataset. Explored the "Show, Attend & Tell" paper, built vocabulary, extracted features with ResNet-18, and trained a transformer decoder. Achieved a BLEU-4 score of 0.18 and deployed a Streamlit demo app.Live Demo 👆🏽
cineRank - A movie ranker appCommunity-driven movie leaderboard app with trending picks, sentiment reviews, and personal watchlists. Built using IMDb reviews, advanced text cleaning, EDA, vectorization (BoW, TF-IDF, GloVe, BERT), and BERT fine-tuning for sentiment classification. Features leaderboard, watchlists, and real-time updates.Live Demo 👆🏽
Choose Your Own AdventureInspired by interactive fiction like AI Dungeon, this app lets you become the protagonist in a personalized adventure story. Enter any theme—haunted mansions, space exploration, and more—and AI generates a unique branching narrative with multiple paths and endings. Features include an interactive visual map, concise story nodes (~40 words), meaningful choices, and a clean black-and-white interface for all devices. Explore different decision paths and control your own dynamic storytelling experience.Project Demo 👆🏽
Projects-Based-GenAIHands-on GenAI projects including text generation, multimodal models, and advanced LLM fine-tuning. Explore practical implementations of state-of-the-art generative AI techniques.Project folder
Project-Based-AgenticAIApplied agentic AI projects focusing on autonomous agents, multi-agent systems, and real-world agentic workflows using LangGraph and LangChain.Project folder

Resources

Progress

DaysDateTopicsResources
Day12024‑12‑14Basics of Linear Algebra3blue1brown
Day22024-12-15Decomposition, Derivation, Integration, and Gradient Descent3blue1brown
Day32024-12-16Supervised Learning, Regression and classificationMachine Learning Specialization
Day42024-12-17Unsupervised Learning: Clustering and dimensionality reductionMachine Learning Specialization
Day52024-12-18Univariate linear RegressionMachine Learning Specialization
Day62024-12-19Cost FunctionsMachine Learning Specialization
Day72024-12-20Gradient DescentCampusX, Machine Learning Specialization
Day82024-12-21Effect of learning Rate, Cost function and Data on GDCampusX, Machine Learning Specialization
Day92024-12-22Linear Regression with multiple features, VectorizationMachine Learning Specialization
Day102024-12-23Feature Scaling, Visualization of Multiple Regression and Polynomial RegressionMachine Learning Specialization
Day112024-12-24Feature Engineering, Polynomial RegressionMachine Learning Specialization
Day122024-12-25Scikit-Learn revision, Linear Regression using Scikit LearnMachine Learning Specialization
Day132024-12-26LR lab, ClassificationMachine Learning Specialization
Day142024-12-27Logistic Regression, Sigmoid FunctionMachine Learning Specialization , CampusX
Day152024-12-28Decision Boundary, Cost FunctionMachine Learning Specialization , CampusX
Day162024-12-29Gradient Descent for logical regressionMachine Learning Specialization , CampusX
Day172024-12-30Underfitting, Overfitting, Regularization Polynomial Features, HyperparametersMachine Learning Specialization
Day182024-12-31Neurons, Neural Netowrk, Forward PropagationMachine Learning Specialization
Day192025-01-01Forward Propagation, Tensorflow implementationsMachine Learning Specialization
Day202025-01-02Building and comparing models (Binary Classification)Machine Learning Specialization
Day212025-01-03Vectorization, Model training using TensoflowMachine Learning Specialization
Day222025-01-04Activation Functions, Softmax IntutionMachine Learning Specialization
Day232025-01-05Implementing SoftmaxMachine Learning Specialization
Day242025-01-06Backpropagaton, What and how??Machine Learning Specialization
Day252025-01-07Backpropagation - Why? Advices for applying machine LearningMachine Learning Specialization
Day262025-01-08Model selection, training test, cross validation, Bias and Variance, Learning curvesMachine Learning Specialization
Day272025-01-09Machine Learning Development Process, ML workflowMachine Learning Specialization
Day282025-01-10Implementing ML model: Error Analysis and Transfer LearningNotebook: Implementation, Machine Learning Specialization
Day292025-01-11Error Metrices, Encoding of Categorical Data, TransoformersMachine Learning Specialization , CampusX
Day302025-01-12Scikit-Learn Pipelines & Ridge Regression (L2 Regularization)Documentation: Scikit-Learn , CampusX
Day312025-01-13Lasso Regression (L1 Regularization), Elastic Net RegularizationML playlist @CampusX
Day322025-01-14Decision Tree Emtropy and Information GainML playlist @CampusX
Day332025-01-15Hyperparameters of Decision Tree with Scikit Learn, Regression TreesML playlist @CampusX , Visualize Yourself>>
Day342025-01-16Visualization Using DtreeViz(), Ensemble LearningGithub Repo: Dtreeviz, ML playlist @CampusX
Day352025-01-17Voting Ensemble >> Classification and RegressionML playlist @CampusX , Visualize Yourself
Day362025-01-18Bagging Ensemble > Classification and RegressionML playlist @CampusX
Day372025-01-19Random Forest: Intution, Working and difference with bagging, Random Forest HyperparametersML playlist @CampusX
Day382025-01-20Boosting Ensemble: Adaboost BoostingML playlist @CampusX
Day392025-01-21Understanding GradientBoosting with RegressionML playlist @CampusX
Day402025-01-22Gradient Boosting with ClassificationML playlist @CampusX , Vlog Link
Day412025-01-23XGboost IntroductionML playlist @CampusX
Day422025-01-24XGBoost for Regression and Classification, Catboost Vs XGboost Vs LightGBMML playlist @CampusX ,Research Paper
Day432025-01-25Stacking Ensemble, Understanding Blending and K foldML playlist @CampusX
Day442025-01-26K-Nearest Neighbor, Coding KNN from ScratchML playlist @CampusX
Day452025-01-27Support Vector MachineML playlist @CampusX
Day462025-01-28K-Means Clustering, DBSCANNotebook: K-Means clustering Demo , Notebook: DBSCAN demo
Day472025-01-29Hierarchical Clustering, Silhouette ScoreKaggle, ML playlist @CampusX
Day492025-01-30PCA (Principle Component Analysis), Implementing with MNIST datasetNotebook: Applying PCA on MNIST dataset
Day502025-02-01Visualizing and Comparing PCA, t-SNE, UMAP, and LDA + Revision with the course ML specializationMachine Learning Specialization
Day512025-02-02Anomaly DetectionMachine Learning Specialization, Notebook: Anomaly Detection
Day522025-02-03Collaborative FilteringMachine Learning Specialization
Day532025-02-04Project @ Football Players Market Value Prediction - Introduction and PlanningProject Plan
Day542025-02-05Project @ Football Players Market Value Prediction - Collecting Data (Scraping)Notebook
Day552025-02-06Project @ Football Players Market Value Prediction - Cleaning DataNotebook
Day562025-02-07Project @ Football Players Market Value Prediction - EDANotebook
Day572025-02-08Project @ Football Players Market Value Prediction - Feature Engineering: (Creating features, Transforming Features)Notebook
Day582025-02-09Project @ Football Players Market Value Prediction - ML: (Linear Regression with Refined Features and deploying with Streamlit)Notebook
Day592025-02-10Project @ Complete Streamlit setup for Linear RegressionStreamlit Documentation
Day602025-02-11Project @ Testing Ridge, Lasso, and Decision TreesProject @ Football Players Market Value Prediction
Day612025-02-12Project @ Had to hit reset from Feature EngineeringProject @ Football Players Market Value Prediction
Day622025-02-13Project @ Finalizing Project and Deploying itProject @ Football Players Market Value Prediction
Day632025-02-14Content-Based Movie Recommender System - PreprocessingNotebook
Day642025-02-15Content-Based Movie Recommender System - Building and DeploymentLive Demo
Day652025-02-16Diving into Deep LearningIntro to Deep Learning @MIT
Day662025-02-17PerceptronsDeep learning playlist @ CampusX
Day672025-02-18Perceptron, Loss function and gradient DescentDeep learning playlist @ CampusX , Grokking Deep Learning @Andrew W. Trask
Day682025-02-19Multilayer PerceptronDeep learning playlist @ CampusX
Day692025-02-20MLP notation, Forward PropagationDeep learning playlist @ CampusX
Day702025-02-21Loss Functions for deep learningDeep learning playlist @ CampusX
Day712025-02-22Backpropagation, deep diving this timeDeep learning playlist @ CampusX
Day722025-02-23Implementing Backpropagation for RegressionNotebook: Backpropagation Regression
Day732025-02-24Implementing Backpropagation for ClassificationNotebook: Implementation Backprop Classification
Day742025-02-25Revising old days, MemoizationDeep learning playlist @ CampusX
Day752025-02-26Vanishing Gradient, Exploding GradientDeep learning playlist @ CampusX
Day762025-02-27Implementing artificial neural networks (ann) for different datasetsDeep learning playlist @ CampusX
Day772025-02-28Improving Neural NetworksDeep learning playlist @ CampusX
Day782025-03-01Sequence Modeling / RNNs - Just OverviewIntro to Deep Learning @MIT
Day792025-03-02Transformers Attention - Just OverviewIntro to Deep Learning @MIT
Day802025-03-03CNNs - Just Overview Part 1Intro to Deep Learning @MIT
Day812025-03-04CNNs - Just Overview Part 2Intro to Deep Learning @MIT
Day822025-03-05Deep Generative Modeling - Just OverviewIntro to Deep Learning @MIT
Day832025-03-06Reinforcement Learning - Just OverviewIntro to Deep Learning @MIT
Day842025-03-07Deep Learning: Challenges & New Frontiers - Just OverviewIntro to Deep Learning @MIT
Day852025-03-08Early Stopping & Normalizing Inputs, DroputDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day862025-03-09Regularization, QuantizationDeep learning playlist @ CampusX
Day872025-03-10Activation Functions - RevisitedDeep learning playlist @ CampusX
Day882025-03-11Weight InitializationDeep learning playlist @ CampusX
Day892025-03-12Deep Learning OptimizersDeep learning playlist @ CampusX
Day902025-03-13Keras TunerDeep learning playlist @ CampusX
Day912025-03-14Deep Diving into CNNsDeep learning playlist @ CampusX
Day922025-03-15Understanding Paddings and StridesDeep learning playlist @ CampusX
Day932025-03-16Backpropagation in CNNs: A Quick BreakdownDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day942025-03-17LeNet5, Cat Vs Dog ClassificationDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day952025-03-18GPU slow than CPU - well in my case?Deep learning playlist @ CampusX
Day962025-03-19Data Augmentation, Pretrained ModelsDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day972025-03-20Visualizing Convolutional Layers, Transfer LearningDeep learning playlist @ CampusX, Grokking Deep Learning @Andrew W. Trask
Day982025-03-21Keras Functional APIDeep learning playlist @ CampusX , Article
Day992025-03-21Finalizing Dog Cat Classifier ProjectProject - Live Demo
Day1002025-03-23Hidden Markov Model, Quantum Machine LearningMedium Article: Understanding Hidden Markov Models
Day1012025-03-24Exploring Pytorch SurfacelyDL with Pytorch - Datacamp
Day1022025-03-25Training a neural network with pytorchDL with Pytorch - Datacamp
Day1032025-03-26Evaluating and improving modelsDL with Pytorch - Datacamp
Day1042025-03-27Crawling through DL with pytrochDeep Learning for Coders -Fast.Ai, @CampusX
Day1052025-03-28Starting chapter 2 : from model to productionDeep Learning for Coders -Fast.Ai, @CampusX
Day1062025-03-29Exploring Autograd and Portfolio TweaksDeep Learning for Coders -Fast.Ai, @CampusX
Day1072025-03-30Refining Portfolio whole dayDeep Learning for Coders -Fast.Ai, @CampusX
Day1082025-03-31Autograd in PyTorch: Deeper UnderstandingDeep Learning for Coders -Fast.Ai, @CampusX
Day1092025-04-01PyTorch Training Pipeline (Manual + Using nn.Module)Deep Learning for Coders -Fast.Ai, @CampusX
Day1102025-04-02Dataset & DataLoader Class in PyTorchDeep Learning for Coders -Fast.Ai, @CampusX
Day1112025-04-03ANN on Fashion MNIST, GELUDeep Learning for Coders -Fast.Ai, @CampusX
Day1122025-04-04ANN on larger FMNIST dataset with GPU (local), GELU/SiLU historyDeep Learning for Coders -Fast.Ai, @CampusX
Day1132025-04-05Optimizing FMNIST NN using Dropouts, Regularization and Batch Normalization in PytorchNotebook
Day1142025-04-08RNNs revisited, Karpathy's blog, Project PlanningKarpathy Blog
Day1152025-07-08Classifying Footballers with their Eyes - Day 1Project Notebook
Day1162025-07-09Classifying Footballers with their Eyes – Day 2Project Notebook
Day1172025-07-10YOLO (You Only Look Once)YOLO Paper
Day1182025-07-11LSTM, GRU & Encoder-Decoder ArchitectureColah's Blog
Day1192025-07-12Bahdanau Attention and Luong AttentionBahdanau Paper
Day1202025-07-13Building a Seq2Seq Chatbot – Data Preparation & PreprocessingPyTorch Tutorial
Day1212025-07-14Building a Seq2Seq Chatbot - Defining Model (encoder, attention, decoder)Notebook
Day1222025-07-15Building a Seq2Seq Chatbot – Evaluation / DeploymentLive Demo
Day1232025-07-20Transformers – Deep Dive into Attention and ArchitectureAttention Paper
Day1242025-07-21Transformers – Vitals [understanding everything]Transformer Guide
Day1252025-07-22GPT from Scratch - Project SetupKarpathy's Tutorial
Day1262025-07-23GPT from Scratch - Bigram Language ModelNotebook
Day1272025-07-24GPT from Scratch - Self-AttentionNotebook
Day1282025-07-25GPT from Scratch – Complete TransformerNotebook
Day1292025-07-26Image Captioning – KickoffNotebook
Day1302025-07-27Image Captioning – Feature Extraction & PipelineNotebook
Day1312025‑07‑28Image Captioning – Training & DeploymentBuild a Large Language Model from Scratch
Day1322025‑07‑29Tokenizer Tricks (Subword, Byte-Pair Encoding) + Comparing ArchitecturesBuild a Large Language Model from Scratch
Day1332025‑07‑30Implementing seq2seq (Diff Approach of Visualizations)Build a Large Language Model from Scratch
Day1342025‑07‑31Implementing Transformer Encoder/Decoder AgainBuild a Large Language Model from Scratch
Day1352025‑08‑01Pretraining on Unlabeled Data, Evaluating, Loading Pretrained WeightsBuild a Large Language Model from Scratch
Day1362025‑08‑02Finetuning – ClassificationBuild a Large Language Model from Scratch
Day1372025‑08‑03Finetuning – Teaching LLMs to Follow Prompts and Perform Complex Tasks, VisualizationsBuild a Large Language Model from Scratch
Day1372025‑08‑03Finetuning – Teaching LLMs to Follow Prompts and Perform Complex Tasks, VisualizationsBuild a Large Language Model from Scratch
Day1382025‑08‑04LLM Fine-Tuning & EvaluationBuild a Large Language Model from Scratch
Day1392025‑08‑05Exploring Hugging Face TransformersBuild a Large Language Model from Scratch
Day1402025‑08‑06Project – Sentiment Analysis [Planning] + Exploring ViTNotebook, Kaggle Code
Day1412025‑08‑07Project – Sentiment Analysis [Preprocessing] + ViT ArchitectureKaggle Code
Day1422025‑08‑08Project – Sentiment Analysis [EDA + Testing GloVe]Notebook, Notebook, Kaggle Code
Day1432025‑08‑09Project – Sentiment Analysis [Advanced Architectures]Notebook, Notebook, Kaggle Code
Day1442025‑08‑10Project – Sentiment Analysis [App Deployment]Live Demo, Code
Day1452025‑08‑11Diving Deep into Vision Transformers (ViTs)Blog
Day1462025‑08‑12Diving into Diffusion ModelsVideo
Day1472025‑08‑13Diffusion Model Deep DivePaper
Day1482025‑08‑15Naive Bayes & Gaussian Mixture ModelsLangchain Playlist
Day1492025‑08‑17Introduction to LangchainLangchain Playlist
Day1502025‑08‑18Introduction to Langchain ComponentsLangchain Playlist
Day1512025‑08‑19Deep Dive into Langchain ModelsLangchain Playlist
Day1522025‑08‑20Langchain Prompts and Building a Simple ChatbotLangchain Playlist
Day1532025‑08‑21Structured Output with LangchainLangchain Playlist
Day1542025‑08‑23Langchain Output ParsersLangchain Playlist
Day1552025‑08‑24Langchain Chain Fundamentals (Simple, Sequential, Parallel, Conditional Chains)Langchain Playlist
Day1562025‑08‑25Langchain Runnables (Modular Components, Composable Workflows)Langchain Playlist
Day1572025‑08‑26Runnable Modules Deep Dive (Sequence, Parallel, Passthrough, Lambda, Branch)Langchain Playlist
Day1582025‑08‑27Document Loaders & Text Splitters (RAG Foundations)Langchain Playlist
Day1592025‑08‑28Vector Stores in Langchain (Chroma, CRUD Operations, Similarity Search)Langchain Playlist
Day1602025‑08‑29Retrievers & Few-Shot Learning (Wikipedia, Vector, MMR, MultiQuery, Contextual)Langchain Playlist
Day1612025‑08‑30RAG Application for UCL Draw (Text Splitting, Vector Embeddings, Response Generation)Langchain Playlist
Day1622025‑08‑31Langchain Tools (Built-in & Custom, Tool Calling & Binding)Langchain Playlist
Day1632025‑09‑01Langchain Agents (Zero-Shot, Conversational, ReAct DocStore, Self-Ask)Langchain Playlist
Day1642025‑09‑02Local Agent with Ollama & Langchain (ChromaDB, RAG)Langchain Playlist
Day1652025‑09‑03Introduction to LangGraph (Stateful Agent Workflows)Langgraph Playlist
Day1662025‑09‑04Agentic AI Fundamentals (Autonomy, Components, Planning, Memory)Langchain Playlist
Day1672025‑09‑05LangChain vs LangGraph Comparison (State Management, Chatbot Example)Langgraph Playlist
Day1682025-09-06Building a Branching ChatbotLanggraph Playlist
Day1692025-09-07Persistence with CheckpointsLanggraph Playlist
Day1702025-09-08Exploring LangSmithLanggraph Playlist
Day1712025-09-11Contextual Q&A with MemoryLangchain Playlist
Day1722025-09-12Bhagavad Gita Expert ChatbotLanggraph Playlist
Day1732025-09-13Multi-Agent Debating SystemLanggraph Playlist
Day1742025-09-14Debate Agent App CompletionLanggraph Playlist
Day1752025-09-15Introduction to FastAPI for MLFastAPI Documentation
Day1762025-09-16FastAPI ImplementationFastAPI Documentation
Day1772025-09-17HTTP Request Methods and REST ArchitectureFastAPI Documentation
Day1782025-09-18FastAPI Parameters and Request BodyFastAPI Documentation
Day1792025-09-19Mini Project with FastAPIFastAPI Documentation
Day1802025-09-20Building Industry-Ready APIs with FastAPIFastAPI Documentation
Day1812025-09-23Containerizing FastAPI ApplicationsFastAPI Documentation
Day1822025-09-24fastapi deployment on awsaws docs
Day1832025-09-25project setup – choose your own adventureproject repo
Day1842025-09-26database design and core componentsproject repo
Day1852025-09-27api implementation and background tasksproject repo
Day1862025-09-28backend completion and debuggingproject repo
Day1872025-09-29frontend integration and project completionproject repo
Day1882025-10-01concurrency patterns in fastapistarlette concurrency
Day1892025‑10‑02self-supervised learning – foundationsLil'Log SSL Blog
Day1902025‑10‑03mcp and lazy weekMasked Conditional Prediction
Day1912025‑10‑04ssrl – image & video approachesLil'Log SSRL
Day1922025‑10‑05wrapping up ssl + fun readsPostgres vs SQLite, GPT Speculations
Day1932025‑10‑06llama2 fine-tuning with qloraQLoRA Fine-Tuning Guide
Day1942025‑10‑12gemma 2 fine-tuning using unslothGemma2 Fine-tuning Notebook
Day1952025‑10‑13saving & loading lora adapters, explored "lora without regret"LoRA Blog Post
Day1962025‑10‑14attempted RAG evaluation, faced compatibility issuesMCP Documentation
Day1972025‑10‑16built a custom MCP server, integrated with CursorCustom Implementation
Day1982025‑10‑18studied RL fundamentals: policies, MDPs, rewardsHuggingFace RL Course
Day1992025‑10‑20explored Q-learning & Deep Q-learning, lunar lander envHuggingFace RL Course
Day2002025‑10‑21trained PPO agent on LunarLander-v3 using SB3HuggingFace RL Course



Day 01: Basics of Linear Algebra

linear algebra is used to represent data, perform matrix operations, and solve equations in algorithms like regression, pca, and neural networks.

  • Scalars, Vectors, Matrices, Tensors: Basic data structures for ML.

  • Linear Combination and Span: Representing data points as weighted sums. Used in Linear Regression and neural networks.

  • Determinants: Matrix invertibility, unique solutions in linear regression.

  • Dot and Cross Product: Similarity (e.g., in SVMs) and vector transformations.

Slow progress right?? but consistent wins the race!


Day 02: Decomposition, Derivation, Integration, and Gradient Descent

  • Identity and Inverse Matrices: Solving equations (e.g., linear regression) and optimization (e.g., gradient descent).

  • Eigenvalues and Eigenvectors: PCA, SVD, feature extraction; eigenvalues capture variance.

  • Singular Value Decomposition (SVD): PCA, image compression, and collaborative filtering.

Notes Here

Calculus Overview:

  • Functions & Graphs: Relationship between input (e.g., house size) and output (e.g., house price).

  • Derivatives: Adjust model parameters to minimize error in predictions (e.g., house price).

  • Partial Derivatives: Measure change with respect to one variable, used in neural networks for weight updates.

  • Gradient Descent: Optimization to minimize the cost function (error).

  • Optimization: Finding the best values (minima/maxima) of a function to improve predictions.

  • Integrals: Calculate area under a curve, used in probabilistic models (e.g., Naive Bayes).

Revised statistics and probability concepts. Ready for the ML Specialization course!


Day 03: Supervised Machine Learning: Regression and Classificaiton

  • Supervised Learning:
    Supervised Learning
  • Regression:
  • Classification:

Day 04: Unsupervised Learning: Clustering, dimensionality reduction

data only comes with input x, but not output labels y. Algorithm has to find structure in data. Unsupervised Learinging

  • Clustering: group similar data points together
    alt text
  • dimensionality reduction: compress data using fewer numbers eg image compression
  • anomaly detection: find unusual data points eg fraud detection

Day 05: Univariate Linear Regression:

  • Learned univariate linear regression and practiced building a model to predict house prices using size as input, including defining the hypothesis function, making predictions, and visualizing results.

Notebook: Model Representation

- Univariate Linear Regression Quiz


Day 06: Cost Function:

alt text Visualization of cost function: Visualization of cost function

  • manually reading these contour plot is not effective or correct, as the complexity increases, we need an algorithm which figures out the values w, b (parameters) to get the best fit time, minimizing cost function

Notebook: Model Representation

Gradient descent is an algorithm which does this task


Day 07: Gradient Descent

Notebook: Gradient descent

gradient descent learned the basics by assuming slope constant and with only the vertical shift. later learned GD with both the parameters w and b. alt text

alt text


Day 08: Effect of learning Rate, Cost function and Data on GD

  • learning rate on GD:Affects the step size; too high can overshoot, too low can slow convergence
- cost function on GD:Smooth, convex functions help faster convergence; complex ones may trap in local minima
  • Data on GD:Quality and scaling affect stability; more data improves gradient estimates

Notebook: gradient descent animation 3d


Day 09: Linear Regression with multiple features, Vectorization

Predicts target using multiple features, minimizing error.

  • Vectorization: Matrix operations replace loops for faster calculations.

alt text](image.png

Lab1: Vectorization


Day10: Feature Scaling

Lab2: Multiple Variable

Today, I learned about feature scaling and how it helps improve predictions. There are multiple methods for feature scaling, including

  • Min-Max Scaling
  • Mean Normalization
  • Z-Score Normalization alt text To ensure proper convergence: check the learning curve to confirm the loss is decreasing. Start with a small learning rate and gradually increase to find the optimal value.

alt text alt text


Day 11: Feature engineering and Polynomial Regression

feature engineering improves features to better predict the target.

eg If we need to predict the cost of flooring and have length and breadth of the room as features, we can use feature engineering to create a new feature, area (length × breadth), which directly impacts the flooring cost.

![alt text](01-Supervised-Learning/images/notes_featureengineering.jpg)

explored polynomial regression that models the relationship between variables as a polynomial curve instead of a straight line

Equation:
y = b₀ + b₁x + b₂x² + ... + bₙxⁿ It is useful for capturing nonlinear relationships in data. alt text

alt text

Lab1: Feature Scaling and Learning Rate
Lab2: Feature Engineering and PolyRegression


Day 12: Linear Regression using Scikit Learn

Had a productive session with linear regression in scikit learn. The lab helped me get a better grasp of the process, though I need more practice with tuning models. Also revisited the Scikit-Learn models ,more comfortable with them now

  • Scikit-learn is an open-source Python library used for machine learning that provides simple and efficient tools for data analysis, including algorithms for classification, regression, clustering, and dimensionality reduction.

alt text alt text alt text
Notebook: ScikitLearn GD


Day 13: Classification

Notebook: Graded Lab alt text
Notebook: Classification solution

  • Classification is the process of categorizing items into different groups based on shared characteristics, like classifying tumors into benign (non-cancerous) and malignant (cancerous) based on their growth behavior and potential to spread. The example above demonstrates that the linear model is insufficient to model categorical data. The model can be extended as described in the following lab.

Day 14: Logistic Regression, Sigmoid Function

  • Logistic Regression: A classification algorithm used to predict probabilities of binary outcomes. Logistic regression on categorical data
  • Sigmoid Function: A mathematical function that maps any input to a value between 0 and 1, used in logistic regression to model probabilities. sigmoid-function

Notebook: Sigmoid Function


Day 15: Decision Boundary, Cost Function

Notebook: Cost Function

  • Decision Boundary: A line or surface that separates different classes in a classification problem based on the model’s predictions.
  • cost function: formula cost function notes

Notebook: Decision boundary

Notebook Logistic Loss


Day 16: Gradient Descent for Logical Regression

Notebook: Gradient Descent Model implementation

Notebook: GD with Scikit-learn

Learned logistic regression cost, gradient descent, and sigmoid derivatives through step-by-step derivations and comparisons with linear regression.

notes_day16 alt text


Day 17: Underfitting, Overfitting

Today, explored teh concepts, overfitting (high variance), underfitting (high bias) and generalization(just right). Regularization to reduce Overfitting. Explored Regularized logistic regression.

  • If the data is in non linear behaviour then we have to appply Ml algos like decision tree, random forest and svm.

Explored hypermeters of logistic regression, and gained some knowledge. text text

Notebook:Overfitting Solution

text

Notebook:Regularization text


Day 18: Neurons, Layer, Neural netowrk, forward propagation

  • neural network: neural networks are machine learning algorithms that model complex patterns using multiple hidden layers and non-linear activation functions. they take inputs, pass them through hidden layers of neurons, and output a prediction. alt text

  • Neurons: a neuron takes weighted inputs, applies an activation function, and outputs a result. inputs can be features or outputs from previous neurons, with weights adjusting their influence. alt text fig: single neuron in action

  • Synapse: synapses connect neurons and carry the weighted inputs. each connection has a weight that adjusts during training.

  • weights: weights control the strength of connections between neurons. they are multiplied by inputs to influence the output, and are adjusted during training.

Popular activation functions include relu and sigmoid.

  • Bias: bias is a constant added to the weighted input before applying the activation function, helping the model represent patterns that don’t pass through the origin.

  • Layers: alt text

    • input layer: holds the data for the model, with each neuron representing an attribute.
    • hidden layer: applies activation functions to the inputs and passes results to the next layer.
    • output layer: receives input from the last hidden layer and returns the model’s prediction.

Day 19: Forward Propagation

  • Forward Propagation: Input data is “forward propagated” through the network layer by layer to the final layer which outputs a prediction.

Notes for today: alt text alt text

Matrix Representation: text How forward Prop works for digit classification?? text Notebook: Neurons and Layers Notebook: A small Neural Netowrk using tensoflow

  • tensorflow basics:
    • representation of data:numpy arrays used for input (e.g., 2D arrays).

      x = np.array([[1, 2, 3], [4, 5, 6]])
      
      
    • building a neural network:

      1. define layers:

        layer1 = dense(units=25, activation='sigmoid')
        layer2 = dense(units=15, activation='sigmoid')
        layer3 = dense(units=1, activation='sigmoid')
        
        
      2. stack layers in a model:

        model = sequential([layer1, layer2, layer3])
        
        
      3. compile and train:

        model.compile(optimizer='adam', loss='binary_crossentropy')
        model.fit(x, y, epochs=10)
        
        
    • visualization:neurons connect layer by layer, with weights and biases computed at each step (refer to attached gif).


Day 20: Python Implementation from Scratch

Implemented forward propagation to compute predictions and backpropagation to optimize weights for binary classification.

  • AGI: An advanced AI capable of generalizing across tasks like humans.

loss graph

model accuracy

Notebook: Building Models


Day 21: Vectorization, Model Training

Exploredd Vectorization for efficient computation

  • Tensorflow: model.compile, binary_crossentropy, model.fit trained a binary classification model and tested its accuracy

alt text Training Model with tensorflow: alt text Notes for today: Notes for today


Day 22: Activation Functions, Softmax

the universal approximation theorem explains that a neural network with enough hidden neurons and non-linear activations like sigmoid or relu can approximate almost any function, even complex patterns like wavy graphs.

Activation Functions:

commonly used activation functions include:

  • sigmoid: squashes values between 0 and 1, often used for binary classification.
  • relu: outputs 0 for negatives and the input itself for positives, commonly used in hidden layers.
  • tanh: outputs between -1 and 1, useful for centered data.

multiclass example

for multiclass classification, softmax is ideal in the output layer as it converts logits into probabilities that sum to 1. during training, the model adjusts weights to maximize the correct class probability, using categorical cross-entropy loss. softmax generalizes logistic regression, which is typically used for binary classification. in both, activation and loss functions differ based on the output type.

Logistic Vs softmax:

logistic vs softmax

Notes:

Notes Notes

Day 23: Implementation of Softmax Regression

NOTE: softmax regression is a classification algorithm that calculates probabilities for multiple classes using a linear combination of inputs and the softmax function. the class with the highest probability is chosen as the prediction

Improved Implementation of Softmax: Roundoff

Tensorflow implementation:

model = Sequential(
    [ 
        Dense(25, activation = 'relu'),
        Dense(15, activation = 'relu'),
        Dense(4, activation = 'softmax')    # < softmax activation here

        ##         Dense(4, activation = 'linear')   #<-- Note
    ]
)
model.compile(
    loss=tf.keras.losses.SparseCategoricalCrossentropy(),
    ##     loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),  #<-- Note ---- This is preferred model softmax and loss are combined for more accurate result.
    optimizer=tf.keras.optimizers.Adam(0.001),
)

model.fit(
    X_train,y_train,
    epochs=10
)
    

Day 24: Backpropagation - what and how ?

  • what? Backpropagation adjusts nneural network weights by propagating errors backward using the chain rule and optimizing them with gradient descent.
  • how? Forward pass to compute outputs, Calculate errors, propagate them backward, and update weights iteratively. Take a look at document below.

Backpropagation example with Neural Network: Backpropagation example with Neural Network

Notebook: Backprop

pdf- How to implement??

Notes:

alt text

Day 25: Backpropagation - Why? Advices for applying machine Learning

Today, I dived into the reasons behind backpropagation's effectiveness in training neural networks. It's not just about adjusting weights; it's the gradients that guide optimization, helping the model minimize error and improve predictions. The backpropagation process makes sure that the error gets distributed in a way that leads to better learning.

How????

  • Gradients: they ensure efficient error minimization by guiding weight updates.
  • Optimization: properly tuned gradients lead to smoother optimization and faster convergence.
  • Model Evaluation: evaluating model performance becomes easier with backpropagation because of the systematic error propagation and weight adjustments.

Notes:

alt text


Day 26: Model selection and training/cross validation/test sets, Bias and Variance

mnist dataset: Label and our prediction after training alt text Errors in our prediction: alt text

  • the importance of splitting data into training, validation, and test sets cross-validation and its role in hyperparameter tuning.
  • Diagnosing bias and variance with error trends: high bias = underfit, high variance = overfit.
  • regulariaztion to handle the tradeoff between bias and variance.
  • how learning curves reveal insights about model performance and whether gathering more data will help. summary of learning algorithm: summary of learning algorithm:

Notes: notes

Notebook: Practice Lab: Neural Networks for Handwritten Digit Recognition, Multiclass

Notebook: Diagnosing Bias and Variance

Notebook: Model Evaluation and selection


Day 27: Machine Learning Development Process, ML workflow

machine learning development process

  1. ml development is iterative, involving:
    • choosing model and data architecture.
    • training the model.
    • diagnosing bias, variance, and errors.

alt text 2. error analysis: identify and fix patterns in model failures.
3. adding data:

  • data augmentation: modify existing data (e.g., distortions).
  • data synthesis: create artificial data.
  1. transfer learning:

    • reuse pre-trained models for similar tasks.
    • fine-tune them with your own data.
  2. ml projects follow these steps:

    • data collection.
    • preprocessing.
    • modeling.
    • evaluation.
    • deployment.
      alt text
    • monitoring.
  3. ethics and fairness -
    ensure ethical use by:

    • avoiding biased decisions in loans, jobs, etc.
    • preventing harmful applications like deepfakes.

Notes: Notes


Day 28: Machine Learning Model: Error Analysis and Transfer Learning

Notebook: Code Implementation from Scratch

  • Confusion Matrix Analysis: the most frequent error is misclassifying 5 as 3. overall, the error rate is around 8%. alt text

  • Iterations Insight: after 200 iterations, the error rate does not decrease significantly, suggesting that 200 iterations are enough for the model to converge. alt text

  • Data Augmentation Insight: despite applying data augmentation, there was no improvement in accuracy. this is because the MNIST dataset is already preprocessed, with centered and normalized images, making the augmentation techniques less effective. in general, data augmentation works best when the dataset is smaller or images are not preprocessed. alt text

  • Transfer Learning with MobileNetV2:

    • Training set accuracy improved from 56.95% to 73.02%.
    • Validation accuracy increased from 70.52% to 74.06%.
    • Loss decreased consistently for both training and validation sets, signaling better learning and generalization.
    • The model shows significant improvement over epochs.
    • Using transfer learning, the model started with an initial accuracy of 74% (pre-trained on ImageNet). As training progressed, the model continued to adapt, improving with each epoch. transfer learning was effective, yielding good results with fewer epochs. alt text

Day 29: Error Metrices, Encoding of Categorical Data, Transoformers

Notebook: Lab week 3: Improving Model

Notebook: Error Metrics

Precision vs. Recall Trade-Off:

  • High Precision: Only hire candidates you’re sure are good. Result: Fewer bad hires, but you might miss some great ones. Example: You hire 5 people, all are good, but you missed 10 other good ones.
  • High Recall: Hire as many as possible to ensure no good candidate is missed. Result: You catch all great candidates but end up with some bad hires too. Example: You hire 50 people, 20 are good, but 30 are bad.
  • When to Focus on Each?
  1. Precision: When mistakes (bad hires) are costly. Example: Hiring a brain surgeon.

  2. Recall: When missing good candidates is worse. Example: Hiring for a customer service team.

Encoding of Categoriical Data:

Encoding TypeUse WhenExample
Label EncodingSmall, unordered categoriesColors: [Red, Blue]
Ordinal EncodingOrdered categoriesEducation: [Low, High]
One-Hot EncodingNominal data, fewer unique categoriesDays: [Mon, Tue, Wed]

Types of Transformers:

TransformerPurposeExample Use CaseInputOutput
Column TransformerApply different transformations to different columns (e.g., scaling and encoding).Scale age and one-hot encode city names.Age: [25, 35, 45], City: [NY, LA, CHI][-1.22, 0, 0, 1], [0, 1, 0, 0], [1.22, 0, 1, 0]
Function TransformerApply a mathematical function (e.g., log or sqrt) to all values.Apply logarithmic transformation to data.[1, 10, 100][0.69, 2.39, 4.61]
Power TransformerNormalize and reduce skewness in data, making it more Gaussian-like.Stabilize variance in highly skewed data.[1, 10, 100][-1.22, 0.0, 1.22]

Day 30: Scikit-Learn Pipelines & Ridge Regression (L2 Regularization)

I explored how to create pipelines in Scikit-learn to streamline the process of combining multiple steps (like preprocessing, model fitting, and regularization) into a single object. This simplifies workflows and ensures reproducibility.

There are three techniques of regularization:

  • Ridge (L2)
  • Lasso (L1)
  • Elastic Net (Combination)

Key Understanding of Ridge Regression:

  1. How the coefficient Get affected?
  • Regularization shrinks the coefficients, preventing them from becoming too large, which reduces overfitting. alt text
  1. Higher Values are impacted more
  • The larger the regularization value (alpha), the more the coefficients are reduced. alt text
  1. Impact on Bias Variance Tradeoff
  • Higher regularization increases bias but reduces variance, making the model more generalizable. alt text
  1. Effect on Loss Function
  • adds a penalty term that limits the magnitude of the coefficients. alt text
  1. Why Ridge Regression is called so?
  • Named after the concept of creating a "ridge" or constraint on the model’s coefficients, preventing them from growing too large.

Notes: alt text alt text

Notebook: Key Understandings of ridge Regression


Day 31: Lasso Regression (L1 Regularization), Elastic Net Regularization

Lasso Regression:

  1. How are coefficients affected by λ (alpha)? As λ increases, regularization strength grows, shrinking less important coefficients to exactly zero. Larger λ values lead to feature selection by removing irrelevant features. alt text

  2. Are higher coefficients affected more? No. Lasso affects smaller coefficients more, shrinking them to zero first. Larger coefficients remain relatively unaffected if they contribute significantly to the model. alt text

  3. Impact of λ on bias and variance:

    • Higher λ increases bias (simpler model) and reduces variance.
    • Lower λ reduces bias (complex model) but increases variance. alt text
  4. Effect of Regularization on Loss Function:

    • Lasso add λ sum(w_i) to the loss, promoting sparsity by penalizing the absolute magnitude of coefficients. Unlike Ridge, it can remove features entirely, improving interpretability. alt text

Elastic Net Regularization:

Elastic Net combines L1 (Lasso) and L2 (Ridge) penalties. It selects important features by shrinking some coefficients to zero (like Lasso) and handles correlated features by shrinking coefficients without setting them to zero (like Ridge). It’s useful when features are both highly correlated and some are irrelevant. Elastic Net is controlled by two parameters: α (mix of Lasso and Ridge) and λ (regularization strength). alt text

Notes: alt text


Day 32: Decision Tree: Entropy and Information Gain

A decision tree is a flowchart-like structure used for classification or regression, where data is split into branches based on conditions until a final decision (leaf) is reached. alt text

  • Decision Tree on Categorical Variables: Splits data based on categories like "Sunny" or "Rainy".
  • Decision Tree on Numerical Variables: Splits data using thresholds like Age > 30.
  • How Decision Tree Works: Repeatedly splits data into smaller groups based on conditions.
  • Terminology: Root (start), Branch (path), Leaf (decision).
  • Pro - Simple to understand; Con - Can overfit.
  • Entropy: Measures uncertainty; low entropy = purer data.
  • Entropy Calculation: ∑ p(x) * log2(p(x)).
  • Information Gain: Reduction in entropy after splitting data. IG=Entropy(before)−Weighted Entropy(after)
  • Gini Impurity: Measures group purity, faster than entropy.
  • Why Use Gini Over Entropy: Simpler and computationally faster.

Notes; alt text


Day 33: Hyperparameters of DT with sclearn, Regression Trees:

alt text

Studied hyperparameters of Decision Trees in Scikit-learn and techniques to handle overfitting and underfitting.

  • Criterion (gini, entropy, log loss): Determines the quality of a split.
  • Splitter: Helps reduce overfitting with better random splits.
  • Max Depth: Controls tree depth; too high causes overfitting, too low causes underfitting.
  • Min Sample Split: Sets the minimum samples required to split a node.
  • Min Sample Leaf: Sets the minimum samples per leaf.
  • Max Features: Determines how many features to use for splits.
  • Max Leaf Nodes: Limits the number of leaf nodes.
  • Min Impurity Decrease: Controls splitting based on impurity reduction.

A Regression Tree predicts continuous variables by splitting data to minimize variance. The best split is determined by maximizing variance reduction, calculated as the variance of the root node minus the weighted average variance of the leaf nodes.

Code: alt text Output: alt text


Day 34: Visualization Using dtreeviz, Ensemble Learning Overview

Notebook: Visualization of Decision Tree Official Notebook: Visualization of Decision Tree

  • started learning about ensemble learning: understood the concept of "wisdom of the crowd" in ensemble methods. types of ensemble methods: voting, bagging, boosting, and stacking. overview of random forest and bootstrapped aggregation (bagging). learned how boosting adjusts predictions iteratively.

Visuals of Bagging, boosting and stacking: Notes Notes; Notes


Day 35: Voting Ensemble > Classification and Regression

  • Soft voting: logistic regression: 60% fraud, random forest: 80% fraud, svm: 40% fraud → final probability: (60% + 80% + 40%) / 3 = 60%. alt text

  • Hard voting: majority wins, logistic regression predicts "fraud," random forest predicts "not fraud," and svm predicts "fraud" → final prediction: "fraud."

Conclusions from Notebook: voting Classifier

  • weighted voting: assigning weights to classifiers helps emphasize stronger models, further improving performance.

  • same algorithm, different hyperparameters: tweaking hyperparameters (e.g., kernel degree in SVM) can lead to significant accuracy changes, highlighting the value of hyperparameter optimization.

Classification:


voting_clf = VotingClassifier(
    estimators=[
        ('lr', log_reg),  # logistic regression: good for linear patterns
        ('rf', rand_forest),  # random forest: captures complex relationships
        ('svc', svm_clf)  # support vector machine: handles edge cases
    ],
    voting='soft',  # averages probabilities from all models for final prediction
    weights=[2, 1, 1],  # gives higher importance to logistic regression
    n_jobs=-1  # enables parallel processing for faster training
)


What to use? soft voting or hard voting, depends if possible use both and then try to find out: alt text alt text

Regression:

  • voting regressor combines multiple regression models to predict continuous values.

similarly, in regression, soft voting averages continuous predictions, and weighted voting helps models with higher performance contribute more.

Voting regressor:

Visualize voting regression here: link

voting_reg = VotingRegressor(
    estimators=[
        ('lr', lin_reg),  # linear regression: good for linear relationships
        ('rf', rand_forest_reg),  # random forest regressor: handles non-linear patterns
        ('svr', svr_clf)  # support vector regressor: captures complex relationships
    ],
    weights=[2, 1, 1],  # assigns higher importance to linear regression
    n_jobs=-1  # enables parallel processing for faster training
)


Day 36: Bagging Ensemble > Classification and Regression

  • bagging (bootstrap aggregating) is an ensemble learning technique that combines predictions from multiple models trained on different subsets of the data (created via bootstrapping) to improve accuracy and reduce variance.

  • Intution:

    alt text

    • random sampling: create multiple datasets by sampling with replacement from the original dataset.
    • train independently: train a model (single preferred) on each bootstrapped dataset (e.g., decision trees).
    • combine predictions: aggregate their predictions by majority voting (classification) or averaging (regression). this reduces overfitting and increases stability, especially for high-variance models.

Notebook: Bagging Intution

Classification:

  • Intution alt text
  • Code Demo
from sklearn.ensemble import BaggingClassifier
from sklearn.tree import DecisionTreeClassifier

# initialize the bagging classifier
bagging_model = BaggingClassifier(
    base_estimator=DecisionTreeClassifier(),  # base model for ensemble; here, decision trees
    n_estimators=10,                          # number of base models to train
    max_samples=1.0,                          # fraction of the dataset for each base model (1.0 = 100%)
    max_features=1.0,                         # fraction of features used in each bootstrap sample
    bootstrap=True,                           # sample datasets with replacement (True enables bootstrapping)
    bootstrap_features=False,                 # sample features with replacement (False = no feature bootstrapping)
    random_state=42                           # seed for reproducibility
)


Regression:

  • Intution: alt text
  • Code Demo:
bagging_model = BaggingRegressor(
    base_estimator=DecisionTreeRegressor(),  # base model for ensemble; here, decision trees
    n_estimators=10,                          # number of base models to train
    max_samples=1.0,                          # fraction of the dataset for each base model (1.0 = 100%)
    max_features=1.0,                         # fraction of features used in each bootstrap sample
    bootstrap=True,                           # sample datasets with replacement (True enables bootstrapping)
    bootstrap_features=False,                 # sample features with replacement (False = no feature bootstrapping)
    random_state=42                           # seed for reproducibility
)


Day 37: Random Forest: Intution, Working and difference with bagging, Random Forest Hyperparameters

random forest is like a bunch of decision trees making a group decision. each tree gets a vote on the outcome, and the most votes win. it's like asking a bunch of experts for advice and going with the majority.

Sampling Techniques:

  • row sampling
  • column sampling
  • combination

Notebook Random Forest

Random Forest vs bagging:

  • why random forest performs well: it reduces overfitting by averaging multiple decision trees trained on different random subsets of data and features, improving accuracy and robustness.

  • random forest vs bagging: both use multiple trees, but random forest adds feature randomness at each split, making trees less correlated and boosting performance.

alt text

Notebook: Random forest Vs bagging

Random forest Hyperparameters:


model = RandomForestClassifier(
    # core hyperparameters
    n_estimators=100,         # number of decision trees in the forest
    max_features="sqrt",      # max features to consider at each split
    max_depth=None,           # max depth of each tree; None = grow fully
    min_samples_split=2,      # min samples needed to split a node
    min_samples_leaf=1,       # min samples required in a leaf node
    bootstrap=True,           # with replacement or without replacement

    # advanced hyperparameters
    max_leaf_nodes=None,      # max number of leaf nodes per tree; None = unlimited
    min_weight_fraction_leaf=0.0,  # min fraction of total weight for a leaf node
    class_weight=None,        # weights for handling class imbalance (e.g., 'balanced')
    ccp_alpha=0.0,            # complexity parameter for pruning; trade-off between size and accuracy
    criterion="gini",         # metric to evaluate splits: "gini" (default) or "entropy"
    warm_start=False,         # reuse previous trees for incremental training; False = train from scratch
    oob_score=False,          # whether to use out-of-bag samples to estimate generalization accuracy
    verbose=0,                # verbosity of output (0 = silent)
    n_jobs=-1,                # number of CPU cores for parallel processing; -1 = use all cores
    random_state=42           # seed for reproducibility
)


Day 38: Boosting Ensemble: Adaboost Boosting

  • boosting is a sequential ensemble learning method where models correct the errors of previous models to reduce bias. combines weak learners (like decision stumps) to create a strong predictive model. Boosting

Notes: alt text

AdaBoost Exploration:

step-by-step understanding of adaboost’s workflow, including:

  • calculating sample weights and weak learner errors.
  • updating model contributions based on performance.
  • final prediction via weighted majority vote.
  • visualized how adaboost focuses on difficult samples.

implemented adaboost from scratch without using sklearn. Adaboost from scratch

Notebook: Adaboost Implementation


Day 39: Understanding GradientBoosting with Regression:

"a model that learns step-by-step by fixing the mistakes of the previous model." Algorithm: alt text

  • Comparision between Adaboost and Gradietn boost: alt text

boosting + gradients (gradients = direction to minimize error).

pseudo-  residual = actual - predicted
new prediction = old prediction + (learning_rate × residual)

gradient boosting variations:

  • xgboost: uses regularization and handles missing data efficiently.
  • lightgbm: faster with large datasets.
  • catboost: great for categorical data.
    gb = GradientBoostingRegressor(
        n_estimators=100,      # number of trees
        learning_rate=0.1,     # step size for updates
        max_depth=3,           # depth of each tree
        min_samples_split=2,   # min samples to split a node
        min_samples_leaf=1,    # min samples per leaf
        subsample=1.0,         # fraction of samples per tree
        max_features=None,     # use all features
        random_state=42        # ensures reproducibility
    )

Notes: alt text


Day 40: Gradient Boosting with Classification

Overview

  • Gradient boosting improves classification by minimizing log-loss iteratively.
  • Each tree focuses on correcting errors made by previous trees.
  • Uses gradients of the loss function to guide updates.

Algorithm

  1. Initialize predictions: F0(x) (log-odds for classification).
  2. For each iteration:
    • Compute pseudo-residuals:
      ri = - ∂(loss) / ∂F(xi)
    • Fit a weak learner (tree) to ri.
    • Update predictions:
      F(x) = F(x) + η * h(x)
      where η is the learning rate.

Loss Function

  • Binary classification: log-loss.
    Loss = -[y log(p) + (1 - y) log(1 - p)]
  • Multiclass classification: softmax loss.
gb = GradientBoostingClassifier(
    n_estimators=100,
    learning_rate=0.1,
    max_depth=3,
    random_state=42
)

Notes: alt text

Notebook: Gradient Boosting Classification

Blog Link

Visuals: Visualizations Visualizations


Day 41: Variations of Gradient Boosting: XGBoost - Introduction.

  • xgboost
  • lightgbm
  • catboost

XGBoost:

alt text xgboost (eXtreme Gradient Boosting) is an advanced implementation of gradient boosting that addresses some key limitations in traditional gradient boosting and adaboost. here's a quick rundown of what i've learned so far. Benchmark

  • Flexibility

    • cross-platform support means xgboost works on different operating systems without much hassle.
    • it supports multiple languages like python, c++, and r, which makes it easy to integrate with your preferred tech stack.
    • integrating with other libraries like scikit-learn or spark is a breeze.
    • it can handle all kinds of ml problems — whether it's classification, regression, or ranking.
  • Speed

    • parallel processing is key. it uses all your cores to speed things up, so you get faster results.
    • optimized data structures (like DMatrix) help reduce memory usage and make computations more efficient.
    • it’s cache-aware, so it knows how to make use of your cpu cache and speed up data retrieval.
    • xgboost handles datasets larger than memory through out-of-core computing.
    • distributed computing helps when your dataset is huge and you need to scale your work across multiple machines.
    • with gpu support, xgboost accelerates matrix calculations, making it much faster for large datasets.
  • Performance (Why this is different from other algos??)

    • regularization is a big win. it prevents overfitting with L1 and L2 regularization, keeping your model generalizable.
    • it automatically handles missing values, learning the best way to fill them in.
    • sparsity-aware split finding is great for sparse data (think text data or categorical features).
    • finding efficient splits in trees is another xgboost feature that speeds up training without compromising accuracy.
    • tree pruning helps reduce tree size after training, improving the model’s performance on unseen data.

Day 42: XGBoost for Regression and Classification, Catboost Vs XGboost Vs LightGBM

Notes on what i explored: Notes Notes Notes

  1. What if we have two or more than two features? We can scan all features , and do all possible splits for all features, then we will calculate gain and similarity score , and select feature which has max gain, algo use greedy search

  2. What if multiple feature and second feature is categorical (like binary: yes/no, male/female or muticlass: colors)? You have to encode them using OHE or other. or use othre variants of GDboost. idenntify unique values in catagorical columns , and consider both as a potential split point ., then calculate gain and similarity scores ., and select with maximum gain .

  3. What if the feature is binary categorical or multiclass categorical? older versions: encode them (e.g., one-hot, label, or target encoding). newer versions: native support for categorical features—no encoding needed.

Catboost Vs LightGBM Vs XGBoost:

alt text

  • categorical features
    • catboost handles them natively—no encoding needed.
    • lightgbm/xgboost require encoding, but xgboost supports label encoding in newer versions.
  • speed
    • lightgbm is the fastest, especially for large datasets.
    • catboost is slower but optimized for small/mid datasets.
    • xgboost is slower than both due to its exhaustive computations.
  • memory usage
    • lightgbm uses the least memory.
    • catboost and xgboost consume more, especially xgboost for large datasets.
  • overfitting prevention
    • all three handle overfitting well, but catboost excels in datasets prone to overfitting due to ordered boosting.
  • use case
    • catboost: small/mid datasets with many categorical features.
    • lightgbm: large datasets where speed and memory are critical.
    • xgboost: general-purpose, robust across various dataset types.

Day 43: Stacking Ensemble: Understanding Blending and K-fold methods

~ blending: splits the data into a training set and a holdout set to train base models and then a meta-model on the predictions of the base models.~

~ k-fold stacking: uses cross-validation to generate predictions for the meta-model by training base models on different training folds and predicting on the validation fold. ~

Notes: Notes:

!Hyperparameters Tuning:

  • identify hyperparameters: for decision trees: max_depth, min_samples_split, etc. for neural networks: learning rate, batch size, number of layers, etc.
  • choose a tuning method: grid search: try every combination of parameters (computationally expensive). random search: test random combinations (faster). bayesian optimization: automatically find the best parameters using probabilistic methods. automated tuning: libraries like optuna or hyperopt.
  • cross-validate: use k-fold cross-validation to test parameter combinations.
  • pick the best: finalize hyperparameters that minimize error metrics or maximize accuracy. alt text

1. manual tuning

  • tweak hyperparameters based on experience or trial-and-error.
  • works for simple models but not scalable.
  • systematically tests all combinations of hyperparameter values.

  • pros: exhaustive, finds the best combo (if time permits).

  • cons: computationally expensive, impractical for large spaces.

  • example:

    from sklearn.model_selection import GridSearchCV
    param_grid = {'max_depth': [3, 5, 10], 'min_samples_split': [2, 5, 10]}
    grid_search = GridSearchCV(DecisionTreeClassifier(), param_grid, cv=5)
    grid_search.fit(X_train, y_train)
    print(grid_search.best_params_)
    
    
  • samples random combinations of hyperparameters.

  • pros: faster, effective for large search spaces.

  • cons: might miss optimal combinations.

  • example:

    from sklearn.model_selection import RandomizedSearchCV
    from scipy.stats import randint
    param_dist = {'max_depth': randint(3, 20), 'min_samples_split': randint(2, 20)}
    random_search = RandomizedSearchCV(DecisionTreeClassifier(), param_dist, n_iter=100, cv=5)
    random_search.fit(X_train, y_train)
    print(random_search.best_params_)
    
    

4. bayesian optimization

  • predicts the best hyperparameters using probabilistic models (e.g., Gaussian processes).
  • pros: efficient in high-dimensional spaces.
  • cons: complex implementation.
  • libraries: Optuna, Hyperopt, BayesSearchCV.

5. evolutionary algorithms

  • optimizes hyperparameters using natural selection (e.g., genetic algorithms).
  • library: TPOT.

6. automated tools

  • auto-sklearn: automates model selection + hyperparameter tuning.
  • H2O.ai: distributed hyperparameter tuning.
  • Optuna: fast, user-friendly library for hyperparameter search.

Link to the blogpost


Day 44: K Nearest Neighbor, Coding KNN from scratch and applying on different datasets:

Predictions are based on the majority vote (classification) or average (regression) of the K closest data points in the training set.

Step-by-Step Workflow:

  1. Choose K: Number of neighbors to consider.
  2. Calculate Distance:
    • Common metrics: Euclidean (default), Manhattan, or Minkowski.
  3. Find K Nearest Neighbors: Identify the K points closest to the query.
  4. Make Prediction:
    • Classification: Majority class among neighbors.
    • Regression: Average value of neighbors.

3. Choosing the Right K

  • Small K (e.g., K=1): High variance, sensitive to noise (overfitting).
  • Large K (e.g., K=50): High bias, smoother boundaries (underfitting).
  • Rule of Thumb: Start with K=nK=n (where nn = number of samples) or use cross-validation.

Overfitting and Underfitting: alt text alt text

Code of KNN using Python: alt text

Finally After Hyperparameter tuning, KNN models imporves for California House price prediction: alt text

Notes: Notes:


Day 45: Support Vector Machines:

  1. What is SVM? Goal: SVM finds the "best" hyperplane to separate data into classes.

Key Idea: Maximize the margin (distance between the hyperplane and the nearest data points, called support vectors).

  • svm helps classify data by finding the best dividing line (hyperplane).
  • support vectors: closest points to the line that influence its position. kind of like the apples and oranges closest to the ruler.
  • margin: the gap between the line and the nearest points. svm tries to make this as wide as possible for better separation. Svm
  • hard margin svm: works only when data is clean and separable. not great for noisy or messy data.
    • example: classifying perfectly labeled "cat" vs. "dog" images where there’s no overlap.
  • soft margin svm: allows some mistakes for better flexibility with noisy/overlapping data.
    • example: separating spam and non-spam emails, where some emails are hard to classify.
  • kernel trick: useful when the data isn’t linearly separable. it projects data into a higher dimension to make it separable.
    • example: in handwriting recognition, svm can map curvy letters into a higher space to separate them more easily.

More Notes like optimization regularization: Notes


Day 46: K-Means Clustering, DBSCAN

K-Means (centeroid based):

Algorithm Steps

  1. Initialize Centroids: Randomly select K data points as initial centroids(Can lead to suboptimal clusters.) (or use k-means++(Distributes initial centroids to improve stability and speed) for smarter initialization).
  2. Assign Points to Clusters: For each data point, compute the Euclidean distance- used by default to all centroids. Assign the point to the nearest centroid.
  3. Recalculate Centroids: Compute the mean of all points in each cluster to update centroids.
  4. Repeat: Reassign points and update centroids until:
    • Centroids stabilize (change < tolerance threshold).
    • Maximum iterations are reached.

Kmeans

Applications

  • Customer Segmentation
  • Image Compression
  • Document Clustering
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

# Preprocess data
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Fit K-means
kmeans = KMeans(n_clusters=3, init='k-means++', n_init=10)
kmeans.fit(X_scaled)

# Get labels and centroids
labels = kmeans.labels_
centroids = scaler.inverse_transform(kmeans.cluster_centers_)

Choosing the Optimal K

  1. Elbow Method: Plot inertia vs. K; the "elbow" (point where inertia decline slows) suggests optimal K.
  2. Silhouette Score: Measures how similar a point is to its cluster vs others. Higher score (closer to 1) = better clustering.
  3. Domain Knowledge: Use prior understanding of the data to guide K selection (e.g., customer segments in marketing).

Notebook: K-Means clustering Demo

One limitation of K-means is that you must specify the number of clusters beforehand.

DBSCAN (density based clustering)

Visualize: DBSCAN here

Why DBSCAN

Notebook: DBSCAN demo

Groups data into density-based clusters (arbitrary shapes) and flags outliers.

  • Core Point: Has ≥ min_samples neighbors within radius eps.

  • Border Point: In a core point’s neighborhood but lacks enough neighbors.

  • Noise: Neither core nor border. alt text

  • eps: Use a k-distance plot (k = min_samples) to find the “knee” for optimal ε.

  • min_samples: Start with 2 * data dimensions (adjust for noise tolerance).

Steps

  1. Pick Parameters: eps (radius) and min_samples (density threshold).
  2. Expand Clusters:
    • For each unvisited point, check if it’s a core point (enough neighbors in eps).
    • If yes, form a cluster by adding all density-reachable points (core neighbors + their neighbors).
    • Mark non-core reachable points as border.
  3. Label Noise: Points not assigned to any cluster.
StrengthsWeaknesses
Finds any cluster shapeStruggles with varying densities
No need for KSensitive to ε and min_samples
Robust to outliersPoor performance in high dimensions

When to Use?

  • Data has noise or complex shapes (e.g., geospatial data, anomaly detection). dbscan vs k means
  • Avoid if clusters have highly varying densities (use HDBSCAN instead).

Day 47: Hierarchical clustering, Silhouette score

Hierarchical Clustering

Hierarchical clustering builds a tree-like hierarchy (dendrogram - a tree diagram showing how clusters merge/split. height shows the distance of merging, and cutting at a height defines the number of clusters.) of clusters. Two main approaches: alt text

  1. Agglomerative (Given below): Start with each point as its own cluster, iteratively merge closest clusters.
  2. Divisive (Reverse of Agglomerative): Start with all points in one cluster, recursively split into smaller clusters.

Agglomerative Clustering Steps:

  1. Initialize: Treat each data point as a singleton cluster.
  2. Compute Distance Matrix: Measure pairwise distances (e.g., Euclidean, Manhattan).
  3. Merge Clusters: Combine the two closest clusters; update the distance matrix.
  4. Repeat: Continue merging until one cluster remains.
ProsCons
No need to specify cluster count upfrontComputationally heavy (O(n³) time, O(n²) space)
Dendrograms aid visualizationSensitive to noise/outliers
Works well for small/mid-sized dataMerges are irreversible (local optima)

Linkage Criteria

Defines how distances between clusters are calculated:

  • Single Linkage: Minimum distance between clusters (prone to "chaining").
  • Complete Linkage: Maximum distance (creates compact clusters).
  • Average Linkage: Average distance between all pairs. alt text
  • Ward’s Method: Minimizes variance when merging (minimizes total within-cluster variance).

Silhouette Score

A metric to evaluate clustering quality by measuring how similar a data point is to its own cluster (cohesion) compared to other clusters (separation). silhouettescore

  • Range: −1 (poor clustering) to +1 (well-defined clusters).

Notes


Day 48: Dimensionality Reduction

Reduces the number of features in data while retaining meaningful patterns, addressing noise, computational cost, and visualization. Methods are hierarchically grouped into feature selection (keeping relevant features) and feature extraction (creating new features).

Dimensionality Redduction Hierarchy: alt text Feature Selection: alt text

Dimensionality Reduction of a Data: alt text

  • Curse of Dimensionality: The curse of dimensionality refers to the challenges and inefficiencies that arise when analyzing data in high-dimensional spaces, such as sparsity, computational complexity, and loss of meaningful patterns.

alt text fig: As the dimensionality of data increases, the feature space becomes sparser, and the data is easier to separate. This is the curse of dimensionality in a nutshell.

Notes:

Day 49: PCA (Principle Component analysis), Implementing with MNIST dataset

Goal: reduce dimensions while preserving maximum variance.

fig: pca projecting 2d data into 1d pc. project illustration

Example: Reducing 10D data to 2D → Use top 2 eigenvectors (highest eigenvalues).

Key Concepts

  • Variance: Spread of data along a feature.
  • Covariance: Measures how two variables vary together.
  • Eigenvectors: Directions of maximum variance (principal components).
  • Eigenvalues: Magnitude of variance along eigenvectors.
  • Transformation: Project data onto new axes (eigenvectors) to reduce dimensions.

Steps →

Steps PCA

How to choose the Optimal PC??

  • select the optimal number of pcs by checking the cumulative explained variance, aiming to retain 90%-95% of the total variance, or by using the elbow method on the explained variance plot where additional pcs add minimal value alt text Notes: Notes

Visualizing MNIST dataset on 2D and 3D:

alt text alt text

Notebook: Applying PCA on MNIST dataset

Conclusion: With about 100 PCs, our model predicts an accuracy of approximately 96%. In comparison, other models like KNN predict around 97% because KNN can capture more complex patterns in the data.

Day 50: Visualizing and Comparing PCA, t-SNE, UMAP, and LDA + Revision with the course ML specialization:

Today was a bit hectic as I tried to understand and visualize various dimensionality reduction techniques. Here's a summary of my learnings and comparisons:

Final Thoughts

  • t-SNE: Captures local similarities well but sometimes distorts the global structure. t-SNE Visuals
  • UMAP: Faster than t-SNE and preserves both local and global relationships better.
  • LDA: Supervised technique, works best when class separation is important.

Comparisons

Comparisons

We explored four dimensionality reduction techniques for data visualization: PCA, t-SNE, UMAP, and LDA. We used them to visualize a high-dimensional dataset in 2D and 3D plots.

Note: It's easy to fall into the trap of considering one technique better than the other. At the end of the day, there's no perfect way to map high-dimensional data into low dimensions while preserving the entire structure. There's always a trade-off in the qualities each technique offers.

Resources

Comparison of Results on MNIST

TechniqueLocal StructureGlobal StructureSupervised?Example Result
t-SNEPreservedNot preservedNoTight clusters of "2"s and "7"s, but arbitrary spacing between clusters.
UMAPPreservedPartially preservedNoTight clusters of "2"s and "7"s, with meaningful spacing between clusters.
LDAPreservedPreserved (class separation)YesDistinct groups for each digit, optimized for classification.

The learning is so hectic today, so I decided to revise the concepts we studied today with the help of Andrew NG. Here's the revision:

  1. What is clustering?
    • Clustering is grouping data points into clusters where points in the same cluster are more similar to each other than to those in other clusters.
  2. K-means?
    • K-means is a clustering algorithm that partitions data into K clusters by minimizing the distance between data points and their respective cluster centroids.
  3. optimization objective?
    • The goal is to minimize the sum of squared distances between data points and their nearest cluster centroid.
  4. Lab?? done// Notebook: Lab Assignment

Day 51: Anomaly Detection:

What is Anomaly Detection?

  • Identifying rare data points (anomalies) that deviate significantly from the majority of the data.
  • Fraud detection, system failure prediction, intrusion detection, healthcare monitoring. alt text

Approaches for Anomaly Detection:

  • Gaussian Distribution: Flag data points outside ±3σ (99.7% of data). alt text
  • Z-Score: Z=(x−μ)σZ=σ(x−μ); if ∣Z∣>3∣Z∣>3, mark as anomaly.
  1. Unsupervised:
    • Clustering (K-means, DBSCAN): Anomalies lie far from cluster centers.
    • Isolation Forest: Randomly splits data; anomalies are easier to isolate.
    • Autoencoders: Neural networks that reconstruct input. High reconstruction error = anomaly.

Evaluation Metrics

  • Precision: % of detected anomalies that are real.
  • Recall: % of true anomalies detected.
  • F1-Score: Harmonic mean of precision and recall.
  • ROC-AUC: Measures trade-off between TPR (recall) and FPR.

Choose Algorithm:

  • Small data? Use statistical methods (Z-score).
  • Large data? Try Isolation Forest or Autoencoders.

Notes:

alt text alt text

  • Assuming data is normally distributed.
  • Ignoring temporal/spatial context (e.g., seasonal trends).
  • Not updating models as data evolves.

Day 52: collaborative filtering

Best Article (Working behind CF): https://medium.com/@ashmi_banerjee/understanding-collaborative-filtering-f1f496c673fd

Collaborative Filtering is a recommendation algorithm that considers the similarities between different users when recommending an item to another user.

Making Recommendations

  • Predict what a user might like based on similar users/items. alt text
  1. User-User CF: Find users like you → recommend what they liked.
  2. Item-Item CF: Find items similar to what you liked → recommend those.

the approach minimizes a regularized cost function using gradient descent, adjusting user and item parameters iteratively. the learning rate (α) controls step size, balancing convergence speed and stability. alt text

while effective, it suffers from the cold start problem (new users/items lack data) and sparsity issues (many missing ratings). despite these challenges, it remains a widely used technique in recommender systems.

Mean Normalization

  • Why? Handle users who rate everything too high/low.

Collaborative Filtering vs Content-Based Filtering

alt text

CFContent-Based
Uses user-item interactionsUses item features (e.g., text, genre)
Example: Netflix recommendationsExample: News articles recommended based on text keywords

Day 53: Project @ Football Players Market Value Prediction - Introduction and Planning

Inspired by the work of Youla Sozen

Every project starts with a problem or question. However, this project is different. it's all about having fun. As a football enthusiast, creating these kinds of projects is always enjoyable. The plan is straightforward, and I will implement it step by step.

Project plan

Although this is a fun project, i aim to ensure the following:

  • accurate player valuation: estimating market values to benefit clubs, agents, and investors.
  • transfer market efficiency: preventing overpayment, aiding negotiation, and optimizing resource allocation.
  • risk assessment & player development: evaluating investments while identifying young talents with high growth potential.
  • data-driven insights: supporting fairer contract negotiations and improving decision-making in fantasy football & betting.

this is a future plan, and i will work towards achieving these goals in the coming days.

Plan of Project

Tips for the project (crafted with deepseek):

Tips


Day 54: Project @ Football Players Market Value Prediction - Collecting Data (Scraping)

web scraping is a technique to collect data from the internet and convert it into a meaningful format, like a data frame, when direct downloads aren't available. in this project, i used the sofifa dataset. here's the main page of sofifa main page

Code to scrape data: alt text

Plan for cleaning Data: alt text


Day 55: Project @ Football Players Market Value Prediction - Cleaning Data

just finished a major data cleaning session for my project. went through steps like handling missing values, converting currencies, splitting combined columns (height/weight), and ensuring consistent data types. cleaned up outliers, removed duplicates, and made sure everything’s ready for the next phase: EDA and model training. feeling good with the progress 😁

Here's a final look: Notebook: Data Cleaning

Code: alt text

Day 56: Project @ Football Players Market Value Prediction - EDA

Notebook: EDA (With Complete Documentation)

I have mostly used Plotly to visualize as it is interactive and for such beginners like me, the visualization impact is powerful. as well as the codes are also easy to write.

overall dataset insights

  1. what are the top 10 most valuable players? alt text
  2. how does market value vary by position (e.g., are strikers more expensive than defenders)? alt text
  3. which teams have the highest average market value? alt text
  4. what’s the distribution of market values (is it skewed towards a few expensive players)? alt text
  5. how does age correlate with market value (are younger players generally worth more)? alt text

player attributes vs. market value

  1. how does a player’s overall rating affect their market value? alt text
  2. which individual attributes (e.g., pace, stamina, strength) correlate the most with market value?
  3. does international reputation (1-5 stars) impact market value?
  4. how do potential ratings compare to market value (are high-potential players priced higher)? alt text
  5. do physical attributes (height, weight, strength) play a role in market value? alt text

position-specific insights

  1. are attacking midfielders (CAM/CM) more valuable than defensive midfielders (CDM/CM)? alt text
  2. how does pace affect wingers' (LW/RW) market value?
  3. do goalkeepers follow the same market trends as outfield players? alt text

contract & transfer market impact

  1. does a player's contract end year affect their market value (e.g., do players with 1 year left have lower values)? alt text
  2. are players on loan priced differently compared to permanent squad members?

Day 57: Project @ Football Players Market Value Prediction - Feature Engineering: (Creating features, Transforming Features)

Notebook: Creating and Transforming Features

spent 6+ hours experimenting with feature engineering. ran into some challenges, but made progress:

  1. position-based features: grouped players into categories (attackers, midfielders, defenders, goalkeepers) with scores. will refine tomorrow.
  2. club-based features: switched from one-hot encoding (curse of dimensionality) to target encoding using the mean market value for each club.
  3. contract-based features: correlation with market value was neutral. most features, except age, overall, and potential, seemed less important.

pairplot

alt text


Day 58: Project @ Football Players Market Value Prediction - ML : (Linear Regression with Refined Features and deploying with Streamlit)

Just for fun: fun

Pairplots: alt text

So, before diving into feature engineering after cleaning the data, i tried out linear regression and got an r2 score of around 0.52. then, after applying some feature engineering and playing around with features, i ran the same model and got the r2 score up to 0.96 with only numerical features
today, i wasn’t fully happy with the result and got confused about feature selection and engineering. so i decided to convert all features into numerical, applied scaling and transformation, and reran the model , r2 score shot up to 0.97
Saved the model immediately, then tried deploying it with streamlit locally, with a user input form and inverse transformations (MOST HATED PART) . while the model’s still a work in progress and not perfect, i'm proud of what i’ve learned so far. next steps are all about finding the best model and getting it deployed with some solid predictions and managing the form with the backend properly.

alt text

Streamlit Preview: https://www.linkedin.com/posts/paudelsamir_day-58365-linear-regression-with-refined-activity-7294398498777501697-2vWi?utm_source=share&utm_medium=member_desktop


Day 59: Project @ Football Players Market Value Prediction - Complete Streamlit Setup for our first Model - Linear Regression

Yesterday, I set up streamlit for my project with some help from ai tools. deploiyng isn't my strong suit, and it got pretty hectic trying to nail down the format and inputs and converting to model inputs. had to leave it unclear


So todayt's goal was get the ui sorted and code the input transformation for the model within streamlit. i could've tinkered with other models and tuned them, but finishing what i started felt right. a simple linear regression is doing surprisingly well for my needs. of course, i'll explore other algorithms and fine tune for better accuracy soon. planning to scale up from 5,000 to 20,000 rows in my dataset. let's see
Here's the demo where linear regression predicts player market values quite accurately. grabbed data from the site i scraped so that to visuailze properly. and its around 90 percent accurate for all the positions. that's already great !! Loving it

Here are some previewsL:

  • KDB (Real Vs Predicted) alt text alt text

  • Lamine (Real Vs Predicted) alt text alt text

  • Oblak (Real Vs Predicted) alt text alt text


Day 60: Project @ Football Players Market Value Prediction - Testing Ridge, Lasso, and Decision Trees

Today was fun! Started by handling outliers for the linear regression model, but didn’t see any improvement, so no luck there. Then I dove into applying PCA for dimensionality reduction. After converting everything to numerical features and applying all the feature engineering, I reduced the features from 50 to 40, and guess what? Model accuracy jumped to 99%! But here’s the twist, I can’t use this model for my project since I’m limited with deployment knowledge, and reverse transforming features while predicting is still the most hectic part of the process.

Applying PCA with 40 features

Then I tried Ridge and Lasso regression. Ridge performed the best and outperformed Linear and Lasso, so I’ll stick with Ridge for now until a simpler model comes along.

Ridge Regression Lasso Regression

Next up was Decision Tree Regressor. Applied it, and without hyperparameter tuning, I got around a 0.98 R2 score. I know Decision Trees are prone to overfitting, so I visualized, but couldn’t predict by myself. Decided to try hyperparameter tuning, but the results weren’t drastically different.

Decision Tree Regressor

At this point, the Decision Tree is the best model for the project. Let's see what’s coming next. Good luck, city. 𝐒𝐡𝐮𝐯𝐚𝐫𝐚𝐭𝐫𝐢 🌙 GridSearchCV Decision Tree

Decision Tree Insights

Did i just wasted 2 hours?? 😅😅Fun

Notebook: Experimentation 1


Day 61: Project @ Football Players Market Value Prediction -Had to hit reset from Feature Engineering

Just 20 minutes ago, i realized i’ve been making a huge mistake since day 4 with feature engineering. i found out today that as a beginner, it’s easy to mess up, but it's all part of the learning process. the mistkae was thinking about how to transform features back for deployment without realizing that features like overall rating, best oiverall, and potential are actually super correlated with market value. i was happy with the 99% accuracy, but i didn’t see the problem until now. i knew about overfitting can cause it and tried to fix it, but i never thought about visualizing feature importance and how it can affect the model.


So, iwas almost done with the project and deployed it to my local server using streamlit. feeling soooo dumb at the moment. the day before yesterday, i tested it with real players like KDB, Lamine Yamal, and Oblak, and 𝐭𝐡𝐞 𝐩𝐫𝐞𝐝𝐢𝐜𝐭𝐢𝐨𝐧𝐬 𝐥𝐨𝐨𝐤𝐞𝐝 𝐠𝐨𝐨𝐝. 𝐰𝐡𝐲?? 𝐛𝐞𝐜𝐚𝐮𝐬𝐞 𝐢𝐭 𝐭𝐨𝐭𝐚𝐥𝐥𝐲 𝐫𝐞𝐥𝐲𝐢𝐧𝐠 𝐨𝐧 𝐣𝐮𝐬𝐭 𝐭𝐡𝐫𝐞𝐞 𝐢𝐧𝐩𝐮𝐭𝐬 𝐁𝐞𝐬𝐭 𝐎𝐯𝐞𝐫𝐚𝐥𝐥, 𝐎𝐯𝐞𝐫𝐚𝐥𝐥 𝐑𝐚𝐭𝐢𝐧𝐠, 𝐚𝐧𝐝 𝐏𝐨𝐭𝐞𝐧𝐭𝐢𝐚𝐥. 𝐲𝐨𝐮 𝐜𝐚𝐧 𝐞𝐯𝐞𝐧 𝐩𝐫𝐞𝐝𝐢𝐜𝐭 𝐰𝐡𝐨𝐥𝐞 𝐭𝐡𝐢𝐧𝐠𝐬 𝐰𝐢𝐭𝐡 𝐣𝐮𝐬𝐭 𝐭𝐡𝐞𝐬𝐞 𝐢𝐧𝐩𝐮𝐭𝐬. The impact of other features is literally minimal. now, i have to rebuild the model from scratch again with better feature engineering.
Even after this dumbest mistake, total wasted grinding, wasted time, wasted energy---for the first time in my learning journey, it feels like now i'm actually learning something.

Context


Day 62: Project @ Football Players Market Value Prediction - Finalizing Project and Deploying it

Finally, i’m able to go live with my first end-to-end ml project! 🎉 you can check it out here: https://paudelsamir.streamlit.app/

  • Today was all about polishing things i messed up, after some solid feature engineering, i was able to hit 85% accuracy, pretty solid start. then, i played around a hour with hyperparameter tuning, which bumped it up to 89%.

actual vs predicted

alt text

but the real magic happened when i tried ensemble learning techniques. after a bit of back and forth, gradient boosting took me all the way to 94% accuracy. and will be using the same algorithm for deployment too.

And with that, after 10 days of nonstop grinding, i’m officially closing this project. it’s been a fun ride, full of learning and surprises !!

𝐆𝐢𝐭𝐇𝐮𝐛 𝐑𝐞𝐩𝐨 For the Project: https://github.com/paudelsamir/ML-Based-Football-Players-Market-Value-Prediction

Day 63: Content-Based Movie Recommender System - Preprocessing

Today, I focused on the preprocessing phase of building a content-based movie recommender system. I created a tags feature by combining key keywords from columns like genres, descriptions, top 3 cast members, and crew, especially the director. This step was crucial to ensure that the recommendation engine has a rich set of features to work with. Notes:


Day 64: Content-Based Movie Recommender System - Building and Deployment

Today, I built the recommendation engine based on movie content similarity using vectorization (bag of words). I also deployed it with Streamlit, so now you can input a movie name and get the top 5 similar movies based on the similarity matrix. Additionally, I integrated an API to pull movie posters in real-time from the website TMDB!

𝐂𝐡𝐞𝐜𝐤 𝐨𝐮𝐭 𝐡𝐞 𝐥𝐢𝐯𝐞 𝐝𝐞𝐦𝐨 𝐡𝐞𝐫𝐞: https://lnkd.in/d7R3Wsnk


Day 65: Diving into Deep Learning

Explored deep learning concepts, including its significance, how it differs from machine learning, and whether it will replace ML. Covered key architectures like Feedforward Neural Networks (FNNs), Convolutional Neural Networks (CNNs) for image processing, Recurrent Neural Networks (RNNs) for sequential data, Autoencoders for feature learning, and Generative Adversarial Networks (GANs) for data generation. Notes from the day: Notes Notes

  • Types of Neural network: alt text

Day 66: Perceptrons

Today, I dived into the concept of perceptrons, which are the building blocks of neural networks. I explored the perceptron algorithm, its working mechanism, and how it can be used for binary classification tasks. A supervised learning algorithm used for binary classifiers.

  • Steps in Prceptron Algorithm: alt text

  • Perceptron from Scratch: alt text

Visuals:

  • Training Data: alt text
  • Perceptron Training: alt text

I explored the fundamentals of MLOps with this paper: Machine Learning Operations (MLOps): Overview, Definition, and Architecture : https://arxiv.org/pdf/2205.02302


Day 67: Perceptron Loss Function and Gradient Descent

Today, I explored the Perceptron Loss Function, which helps adjust weights when misclassification occurs, ensuring better decision boundaries. I learned how the perceptron updates its weights using the weight update rule and how Gradient Descent optimizes the loss function by iteratively moving in the direction of the negative gradient.

alt text alt text

Notes: alt text


Day 68: Multilayer Perceptron

The problem with Perceptrons lies in their limitation to learn complex patterns and functions, especially those that are not linearly separable. A Perceptron is a single-layer neural network with binary outputs, and it can only solve problems where the data points are linearly separable. If the data is not linearly separable, a Perceptron cannot converge and find a solution. alt text

So the solution is Multilayer Perceptron: GIF

There's a website named: https://playground.tensorflow.org/

I practiced different optimizations there some of them are,

  • Adding nodes to hidden layer
  • Adding nodes to input layer
  • Adding nodes to output layer (for multicalss)
  • Adding nuber of hidden layer

The conclusion is you can classify any type of problem within regression and classifcion by optimizing those nodes and others like activation and regularization.

Batch and Gradient Descent:

number of samples processed before updating model weights

alt text

  1. SGD (batch size = 1) → updates after each individual sample (one row at a time). noisy but good for escaping local minima.
  2. Mini-batch gradient descent (batch size = 16, 32, 64, etc.) → updates after a small group of samples (e.g., rows 1-100). balances speed and stability.
  3. Full-batch gradient descent (batch size = all samples) → updates after seeing the entire dataset. very stable but slow and memory-heavy.

each batch in mini-batch or full-batch contains multiple rows, and the loss is computed over those samples before updating weights.

so in SGD, you're updating weights after every single row (which makes it very random and noisy). in mini-batch, you take a chunk of rows, calculate gradients over that group, then update weights. in full-batch, you process all the rows at once and then update.

Notes: alt text


Day 69: MLP notation, Forward Propagation

today, i deepened my understanding of multi-layer perceptrons (MLPs), including their formal notation and the calculation of weights and biases for each layer. i also explored forward propagation and practiced matrix multiplication by manually constructing and multiplying matrices to intuitively follow the perceptron’s computations. alt text

additionally, i studied MLP training with pytorch from the book deep learning with python by françois chollet.

additionally, i explored Image processing with datacamp, here's image representation of what i learned today alt text


Day 70: Loss Functions for Deep Learning

  • Regression:

    • MSE : squares errors, punishes big mistakes more.
    • MAE : takes absolute difference, treats all errors equally.
    • Huber Loss: mix of mse & mae, good for outliers. alt text
  • Classificaiton:

    • Binary cross entropy : for yes/no classification (spam or not spam). alt text
    • Categorical Cross entropy : for multiple classes (yes/no/ maybe). alt text
    • Sparse Categorical cross entropy - same as categorical but works with integer labels.
    • Hinge Loss: used in SVMs, pushes correct class far from the wrong ones. alt text
  • Autoencoders / VAE loss:

    • KL divergence
  • GANs:

    • Discriminator loss : helps the discriminator tell real from fake.
    • Minmax Loss : generator tries to fool discriminator by minimizing its best-case performance.
  • Object Detection and segmentation loss:

    • interseciton over union loss
    • smoooth l1 loss
    • dice loss
  • Reinforcement loss:

    • policy gradient - rewards good actions.
    • Q-Learning loss : teaches agent to choose best long-term rewards.
    • proximal policy optimization
  • Custom losss function:

    • Perceptrual loss
    • combined loss functions: mix of different losses for better results (e.g., cross-entropy + dice loss).

Notes:

alt text alt text


Day 71: Deep Diving Backpropagation

I already explored backpropagation in andrew ng’s ml specialization course, but that was more of a surface level explanation just the mechanics of how it works.

Today, i’m diving deep. like, REALLY deep. i want an intuitive, mathematical understanding of backpropagation, not just the algorithmic steps. all my tracing and derivations are going into my handwritten notes. this is for intutive approach to understand Regression part.

Notes: alt text alt text

Tomorrow, i’ll probably implement backprop from scratch, test it on a proper dataset, and try to visualize what’s actually happening. the key question: how?

Also, i might challenge myself to explain why backprop works in my own words. maybe even turn it into an article.


Day 72: Implementing backpropagation for regression

Today was all about applying what i learned yesterday to code and visualizing backpropagation.

first, i created a toy dataset that looks like this:
alt text

Then, i wrote functions to implement backpropagation from scratch. after running the training loop, here’s what the final parameters looked like—no keras, no tensorflow, just raw python:
alt text

all the code is in my notebook:
Notebook: Backpropagation Regression

I also tried using keras' sequential api to train the same model. after around 700 epochs, the error dropped significantly.
Keras

final weights with keras:
alt text

PS: intentionally chose a confusing dataset to mess with my own head.


Day 73: Implementing Backpropagation for Classification

i already implemented backprop for regression, both handwritten and in code. today, i'm tweaking it for classification.

few things to change:

  • loss function is binary cross-entropy instead of MSE
  • activation function is sigmoid instead of linear

but the backprop algo stays the same. all derivatives are now based on the new log function. since i already get the intuition, i'm skipping the math and just coding it.

Notebook: implementation backprop classification

sample data looks like this:
sample data

final parameters after training:
final parameters

function to update parameters:
parameter update function

at last, i tried implementing the same using tensorflow to see how it compares to my scratch implementation:
tensorflow code


Day 74: Revising old days, Memoization

Today, i revised concept of gradient and derivatives, focusing on how subtracting the gradient term is helping minimize loss. gradient relies on derivatives to find the optimal weights, and the learning rate controls the step size too high can cause overshooting, while too low leads to slow convergence. another key takeaway was memoization, a technique to store previously computed values to optimize calculations. in neural networks, repeated derivative computations can slow down training, and memoization helps speed things up by avoiding redundant calculations. this approach is widely used in dynamic programming and can improve efficiency in deep learning models.

Notes: alt text alt text


Day 75: Vanishing Gradient, Exploding Gradient

Today i first revised Gradient descent in NN:

  • Batch : faster to complete epochs ( batch size = all)

  • Stochastic : faster to converge ( batch size = 1)

  • Mini- Batch : mostly suitable ( batch size around center)

  • vanishing gradient: when gradients become too small, causing early layers to learn very slowly or not at all. example: in deep networks using sigmoid activation, earlier layers stop updating because gradients shrink to near zero.

  • exploding gradient: when gradients become too large, leading to unstable updates and divergence. example: in rnn training, weights keep multiplying large gradients, causing values to explode to infinity.

problems

How to Handle VGD problem:

  • use better activation functions – replace sigmoid/tanh with relu, leaky relu, or elu. Using Sigmoid: alt text Using ReLU: alt text This is the final weights comparision: alt text
  • use proper weight initialization – xavier/glorot for sigmoid/tanh, he initialization for relu.
  • batch normalization – normalizes activations to maintain stable gradients.
  • residual connections (skip connections) – used in resnets to allow gradients to flow easily.
  • gradient clipping – caps gradients to prevent them from becoming too small.

Day 76: Implementing artificial neural networks (ann) for different datasets

  • Experimented with ann on two datasets: mnist for handwritten digit classification and a gre dataset for graduate admission prediction. the goal was to train models and analyze performance across different domains.

Notebook: GRE prediction
Notebook: MNIST Classification

  1. trained an ann on mnist to classify handwritten digits (0-9) using a sequential model with dense layers. training ran for 30 epochs with promising results.
    • training results (30 epochs):
      mnist training
    • model summary:
      mnist model summary
    • sample predictions:
      mnist prediction
  • applied ann to predict graduate school admission chances based on gre scores, gpa, and other factors. dataset required preprocessing before feeding into the model. trained for 100 epochs.
    • sample dataset:
      gre dataset sample
    • training results (100 epochs):
      gre training
    • model summary:
      gre model summary
    • model accuracy evaluation:
      gre model accuracy

next steps: hyperparameter tuning, dropout layers for regularization, and testing on additional datasets.

Day 77: Improving Neural Networks

Notes: notes

Fine-tuning Neural Network Hyperparameters

Fine-tuning neural network hyperparameters is about adjusting key settings to improve learning

  • Learning rate - controls how fast the model updates. Too high = unstable, too low = slow learning.
  • Batch size - number of samples processed before an update. Small = noisy but frequent updates, large = stable but slow.
  • Epochs - full passes through data. Too few = underfitting, too many = overfitting.
  • Layers & neurons - more can improve learning but make training harder.
  • Activation function - decides neuron output; common ones are ReLU, sigmoid, tanh. alt text
  • Dropout - turns off some neurons randomly to prevent overfitting.
  • Optimizer - algorithm that adjusts weights efficiently (Adam, SGD, etc.).

Problem Solving Strategies

  • Vanishing/exploding gradient – gradients shrink or blow up, stopping learning → use ReLU, batch norm, weight init, or residual connections.
  • Not enough data – small datasets cause poor generalization → apply data augmentation, transfer learning, or synthetic data generation.
  • Slow training – long training times due to large models or bad optimizers → use mini-batches, better optimizers, mixed precision, and GPUs.
  • Overfitting – model memorizes training data but fails on new data → apply dropout, regularization, early stopping, and data augmentation.
  • Underfitting – model is too simple and fails to learn patterns → increase complexity, train longer, improve features, and reduce regularization.
  • Imbalanced data – one class dominates, leading to biased predictions → use class weighting, oversampling, or synthetic data (SMOTE/GANs).
  • Poor generalization – model does well in training but fails on real-world data → ensure diverse data, reduce leakage, use domain adaptation, or adversarial training.
Transfer Learning

alt text alt text

Day 78: Sequence Modeling / RNNs - Just Overview

I just thought ki Before diving deeper into neural network improvement techniques, I should first gain a surface-level understanding of deep learning concepts I'll be tackling in the future as this learning technique is helping me alot. For this purpose, I found an excellent YouTube playlist: MIT 6.S191: Introduction to Deep Learning. There are approximately 10 to 15 videos that I plan to watch to build a foundational overview and i too will deep dive into these later

alt text alt text

Implementation Preview alt text alt text


Day 79: Transformers , Attention - Just Overview

Transformers replace RNNs by using self-attention, enabling parallel processing and handling long-range dependencies efficiently. Introduced in "Attention Is All You Need" (2017), they power models like BERT and GPT.

Self-attention computes Query (Q), Key (K), and Value (V) matrices to determine word relationships. Multi-head attention allows the model to capture different contextual meanings.

The transformer consists of encoder-decoder blocks with self-attention, feed-forward layers, and normalization. Encoders learn representations, while decoders generate sequences.

Transformers are used in chatbots, translation, search engines, and AI coding assistants. Key models include BERT (bi-directional understanding), GPT (text generation), and T5 (text-to-text tasks).

alt text alt text alt text


Day 80: CNNs - Just Overview Part 1

  • Computer vision enables machines to interpret visual data.
  • Used in self-driving cars, medical imaging, surveillance, AR.

What Computers "See"

  • Images are matrices of pixel values (grayscale: single matrix, RGB: 3 matrices).
    alt text

Feature Extraction & Convolution

  • CNNs learn features like edges, textures, and shapes automatically.
  • Convolution: uses filters (kernels) to extract patterns from images.
    alt text
  • Example Kernel (Edge Detection - traditional filter): alt text

CNN Architecture

  1. Convolution Layers: detect patterns.
  2. ReLU Activation: makes model non-linear.
  3. Pooling Layers: reduce size, keep key features.
  4. Fully Connected Layers: classify objects.
    alt text

Object Detection

use yolo when speed matters more than precision (e.g., real-time apps). use rcnn when accuracy is critical and speed isn’t a constraint. use faster rcnn for a balance between accuracy and speed. alt text

Self-Driving Cars

  • CNNs help detect lanes, pedestrians, and traffic signs.
  • End-to-end models predict steering angles using video input.
    alt text alt text

Day 81: CNNs - Just Overview Part 2

yesterday, i watched a video on cnn. the goal was just to explore it for a day, but i feel like this is an interesting topic. so today, i want to learn more—like, in-depth—about the feature extraction part, which i find the most interesting aspect of cnn.

i learned to use relu and understood convolution layers and how they work yesterday. but today, i learned about pooling and how it reduces the size. i also explored the classification process in more depth.

notes: alt text note: cnn by itself doesn't handle rotation and scaling well. for that, use data augmentation.


Day 82: Deep Generative modeling - Just overview

Watched this single video : https://www.youtube.com/watch?v=Dmm4UG-6jxA&t=3242s

today’s deep dive into generative models gave me a solid grasp of how ai can not only recognize patterns but also create new data from scratch. these models are the backbone of modern generative ai, and understanding them is key to keeping up with the field.

Generative models: what they do

  • they don’t just classify or analyze data; they generate entirely new data that resembles what they learned from.
  • examples include autoencoders, variational autoencoders (VAEs), GANs, and diffusion models—each with its own strengths. alt text

Latent variable models: finding the hidden structure

  • autoencoders learn to compress data into lower-dimensional representations and then reconstruct it.
  • vaes take this further by adding randomness, making them better at generating diverse outputs instead of just memorizing patterns.
  • plato’s cave analogy clicked here—observed data is like shadows on a wall, while latent variables represent the actual objects casting those shadows. alt text

Generative adversarial networks (GANs): competition makes better results

  • a generator creates fake data, while a discriminator tries to catch the fakes. alt text
  • through constant feedback, both improve, leading to highly realistic outputs.
  • training is tricky—imbalanced learning can cause mode collapse, where the generator keeps making similar outputs instead of diverse ones. alt text

Regularization & structure in vaes

  • forcing the latent space into a structured distribution (often gaussian) helps ensure smoothness and continuity, making vaes more useful for controlled generation. alt text

Applications beyond images

  • these models aren't just for ai art—they power speech synthesis, text generation, and even domain adaptation (like CycleGAN for translating images without paired data).

Diffusion models: a new generative powerhouse

  • instead of generating images all at once (like GANs), diffusion models gradually refine noise into meaningful images. alt text
  • they’re more stable and produce higher-quality results, making them the future of generative ai.

Day 83: Reinforcement Learning - Just Overview

with the help of this video: https://www.youtube.com/watch?v=8JVRbHAVCws&t=3242s

  • explored key concepts of reinforcement learning (RL), including Q-learning and policy learning algorithms. alt text

  • dived into Deep Q Networks (DQN) and their role in handling complex environments, like Atari games. alt text

  • understood the difference between discrete and continuous actions and how they impact RL models. alt text

  • learned about real-world applications of RL, from robotics to game AI, and cutting-edge technologies like AlphaGo and MuZero. alt text

  • explored training techniques like policy gradients and how they improve decision-making in RL agents. alt text alt text alt text alt text


Day 84: Deep learning: challenges & new frontiers - Just Overvview

Video Link: https://www.youtube.com/watch?v=N1fbskTpwZ0&t=3021s

alt text

  • neural network failure modes – ai can fail unpredictably, often overconfident in wrong answers. alt text
  • uncertainty in deep learning – models lack awareness of their own mistakes, leading to unreliable predictions. alt text
  • adversarial attacks – tiny changes in input can trick ai into making completely wrong decisions. alt text
  • algorithmic bias – ai inherits and amplifies biases from training data, leading to unfair outcomes. alt text

generative ai & diffusion models

  • the landscape – diffusion models are replacing GANs in high-quality image generation.
  • diffusion process – gradually add noise to data and train a model to reverse it.
  • noising & denoising – AI learns to reconstruct images by stepwise noise removal. alt text
  • text to image, beyond images – models like DALL·E generate images, but AI is expanding into text, music, and 3D.

large language models (LLMs)

  • using llms to generate text – models predict words to generate human-like text.
  • limitations – hallucinations, bias, and lack of real understanding. alt text
  • more parameters = better performance, but higher computational cost.
  • how they work – transformers + attention mechanism process and generate context-aware text

Day 85: Early Stopping, Normalizing Inputs, Dropout

These are the techniques i will cover in upcoming days:

  1. Vanishing Gradients

    • Activation Functions
    • Weight Initialization
  2. Overfitting

    • Reduce Complexity/Increase Data
    • Dropout Layers
    • Regularization (L1 & L2)
    • Early Stopping
  3. Normalization

    • Normalizing inputs
    • Batch Normalization
    • Normalizing Activations
  4. Gradient Checking and Clipping

  5. Optimizers

    • Momentum
    • Adagrad
    • RMSprop
  6. Learning rate scheduling

  7. Hyperparameter Tuning

    • No. of hidden layers
    • Nodes/layer
    • Batch size

Today, I explored Early Stopping, Normalizing Inputs, and Dropout techniques for improving neural network performance.

  • Early Stopping: Prevents overfitting by halting training when validation performance drops.
  • Normalizing Inputs: Scales features to a consistent range, aiding in faster and more stable learning.
  • Dropout: Reduces overfitting by randomly deactivating neurons during training, forcing the model to generalize better.

Notebook: Dropout on Regression alt text Notebook: Dropout on Classification alt text


Day 86: Regularization, Quantization

L1 and L2 regularization are typically used for smaller networks. For larger networks, it is better to use neural network-specific regularization which is dropout regularization. alt text

An evaluation procedure must be used when using a regularizer to monitor that regularization process. For this, we can plot model performance against the number of epochs during the training process.

Regularization Regularization

Notebook: without Regularization vs Applying Regularization alt text

Quantization:

quantization in deep learning reduces the precision of numbers in a model to save memory and speed up processing. it can be done in two ways: post-training quantization (ptq), which converts the model to lower precision after training for faster performance but may lose some accuracy, and quantization-aware training (qat), where quantization is simulated during training, resulting in better accuracy but requiring more time. frameworks like TensorFlow provide tools for both methods to help deploy lighter and faster models. alt text

Day 87 - Activation Functions - Revisited

why needed? introduce non-linearity to capture complex patterns.

alt text alt text ideal properties: non-linear, differentiable, computationally inexpensive, zero-centered, non-saturating.

  • use relu + he init, sigmoid/tanh + xavier init.
  • relu in early layers, tanh or sigmoid in later/output layers.
  • prefer gelu for transformers.
  • avoid dying relu with leaky relu or prelu.

relu: fast, simple but can die. leaky relu: small slope for negatives, avoids dead neurons. prelu: learnable slope, more flexible. elu: better generalization, but expensive. selu: self-normalizing, good for deep nets. alt text


Day 88: Weight Initialization

Zero Init: All weights as zero → no learning (same gradients).

weights = np.zeros((input_size, output_size))

alt text

One Init: All weights as one → same issue, no symmetry breaking.

weights = np.ones((input_size, output_size))

✅ Random Init: Small random values.

weights = np.random.randn(input_size, output_size) * 0.01

✅ Xavier Init (for tanh/sigmoid):

weights = np.random.randn(input_size, output_size) * np.sqrt(1 / input_size)

alt text

✅ He Init (for ReLU):

weights = np.random.randn(input_size, output_size) * np.sqrt(2 / input_si

alt text

Notebook: Weight Initialization

Notebook: Xavier and He initialization

Notes: alt text alt text


Day 89: Deeep Learning Optimizers

  • Gradient Descent: Fundamental optimizer that updates weights by moving in the opposite direction of the gradient of the loss function w.r.t. the weights. It's basic but effective for many problems.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01)

bgd

  • Stochastic Gradient Descent: Optimizes with each sample instead of the entire dataset, which is faster but more noisy.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01)
  • Momentum: Momentum helps accelerate SGD by moving along relevant directions and dampening oscillations. It adds a fraction of the previous update to the current one.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01, momentum=0.9)

momentum

  • NAG: An improvement over momentum. It looks ahead by calculating the gradient not just at the current position but slightly ahead, giving better convergence.
from tensorflow.keras.optimizers import SGD
optimizer = SGD(learning_rate=0.01, momentum=0.9, nesterov=True)

nag

  • AdaGrad: Adapts the learning rate to the parameters, performing larger updates for infrequent parameters and smaller updates for frequent ones.
from tensorflow.keras.optimizers import Adagrad
optimizer = Adagrad(learning_rate=0.01)

adagrad

  • RMSProp: Divides the learning rate by an exponentially decaying average of squared gradients. Useful for non-stationary objectives.
from tensorflow.keras.optimizers import RMSprop
optimizer = RMSprop(learning_rate=0.001)

rmsprop

  • ADAM (Adaptive Moment Estimation): Adaptive learning rate method that combines the advantages of RMSprop and momentum. Popular for its efficiency.
from tensorflow.keras.optimizers import Adam
optimizer = Adam(learning_rate=0.001)

Adam

Overall: overall Notes: notes notes notes

Interactive Visualization of Optimization Algorithms in Deep Learning: https://emiliendupont.github.io/2018/01/24/optimization-visualization/


Day 90: Keras Tuner

hyperparameters (like learning rate, batch size, optimizer) directly impact model performance. tuning helps optimize accuracy and generalization.

in my case, i worked with the diabetes dataset and focused on tuning key hyperparameters like learning rate, batch size, optimizer, number of neurons, and dropout rate. each of these parameters influences different aspects of training—for example, the learning rate affects how quickly the model converges, while dropout helps prevent overfitting. alt text

to streamline the tuning process, i used keras tuner’s RandomSearch. i defined a hypermodel where parameters like the number of layers, neurons, dropout rates, learning rates, and optimizer types were set as tunable. the objective was to maximize validation accuracy. i also configured settings like max_trials to control the search space and executions_per_trial to ensure consistent evaluation. alt text alt text after running the tuning process, the model achieved around 79% accuracy. the tuning helped balance model complexity and performance, reducing overfitting and improving generalization. alt text for further improvements, i could fine-tune hyperparameters like the learning rate and dropout in smaller increments, try advanced optimizers like adamw, or implement early stopping to avoid unnecessary training once the model stops improving.


Day 91: Deep Diving into CNNs:

Focused on planning a deep dive into cnns. explored why anns fall short for cnn tasks and uncovered some fascinating cnn applications. along the way, stumbled upon some surprisingly cool ideas for future projects. fueled by that curiosity, i tried something basic today—simple, but a solid starting point.

today, i explored opencv from scratch, debugged image loading issues, applied basic filters using custom convolution kernels in pure python as well as with Opencv, and created amazingly undefinable custom filters with the excitement. alt text Notebook:Trying OpenCV for the first time Watch out notebook what i did with this lovely image: alt text

Notes: alt text

architectures

  • early arch
    • lenet-5
    • alexnet
    • vgg net (16/19)
  • modern arch
    • resnet (skip connection)
    • inception net (google net)
    • mobilenet
    • efficientnet

Plan after Architecture

  1. data augmentation (rotation, scaling, flipping)
  2. transfer learning (using pretrained models like vgg, resnet)
  3. object detection (yolo, r-cnn, ssd, etc.)
  4. image segmentation (u-net, mask r-cnn)
  5. attention mechanisms (squeeze, spatial channel-wise)
  6. self-supervised learning (contrastive, ssl architectures)

Day 92: Understanding Paddings and Strides

How CNNs working with Grayscale and Rgb images?? alt text alt text

padding and strides are important in convolutional neural networks (cnns) because they affect feature extraction, output size, and computational efficiency.

Padding is used to prevent reduction in spatial dimensions and retain edge information. for example, a 5x5 image with a 3x3 filter produces a 3x3 feature map, which keeps shrinking with more layers. adding padding helps maintain the size. alt text there are two common types of padding:

  • valid: no padding, meaning the output size shrinks.
  • same: pads the input so the output size remains the same.

the output size with padding is calculated as:

(n + 2p - f + 1) × (n + 2p - f + 1), where n is the input size, f is the filter size, and p is the padding amount.

in keras, padding is applied like this Demo for MNIST:

# Importing necessary libraries
import tensorflow as tf
from tensorflow.keras.layers import Dense, Conv2D, Flatten
from tensorflow.keras.models import Sequential
from tensorflow.keras.datasets import mnist

# Loading MNIST dataset
(X_train, y_train), (X_test, y_test) = mnist.load_data()

# Creating a Sequential model
model = Sequential()

# Adding Convolutional layers with valid padding
model.add(Conv2D(32, kernel_size=(3,3), padding='valid', activation='relu', input_shape=(28,28,1)))
model.add(Conv2D(32, kernel_size=(3,3), padding='valid', activation='relu'))
model.add(Conv2D(32, kernel_size=(3,3), padding='valid', activation='relu'))

model.add(Flatten())

model.add(Dense(128, activation='relu'))
model.add(Dense(10, activation='softmax'))

# Printing model summary
model.summary()

strides control how far the filter moves across the image at each step. a stride of (1,1) moves one pixel at a time, while higher strides skip pixels, reducing spatial dimensions and computation time. stride = 2 output size with strides is calculated as:

((n + 2p - f) / s + 1) × ((n + 2p - f) / s + 1), where s is the stride value.

higher strides help capture larger patterns but reduce spatial resolution. for example, a stride of (2,2) makes the filter shift 2 pixels at a time:

# Importing necessary libraries
import tensorflow as tf
from tensorflow.keras.layers import Dense, Conv2D, Flatten
from tensorflow.keras.models import Sequential
from tensorflow.keras.datasets import mnist

# Loading MNIST dataset
(X_train, y_train), (X_test, y_test) = mnist.load_data()

# Creating a Sequential model with strides
model = Sequential()

model.add(Conv2D(32, kernel_size=(3,3), padding='same', strides=(2,2), activation='relu', input_shape=(28,28,1)))
model.add(Conv2D(32, kernel_size=(3,3), padding='same', strides=(2,2), activation='relu'))
model.add(Conv2D(32, kernel_size=(3,3), padding='same', strides=(2,2), activation='relu'))

model.add(Flatten())

model.add(Dense(128, activation='relu'))
model.add(Dense(10, activation='softmax'))

# Printing model summary
model.summary()

in summary, padding helps retain image details and control output size, while strides affect computational efficiency and feature abstraction. tuning these parameters is key to optimizing cnns.

Pooling:

pooling is used to downsample feature maps, reducing their size while retaining important information. it helps prevent overfitting and reduces computation.

common types of pooling: alt text max pooling: selects the maximum value in a region. average pooling: takes the average of values in a region. for example, a 2x2 max pooling layer with stride 2 reduces feature maps to half their original size.


# importing necessary libraries
import tensorflow as tf
from tensorflow.keras.layers import Dense, Conv2D, Flatten, MaxPooling2D
from tensorflow.keras.models import Sequential
from tensorflow.keras.datasets import mnist

# loading mnist dataset
(X_train, y_train), (X_test, y_test) = mnist.load_data()

# reshaping input data
X_train = X_train.reshape(-1, 28, 28, 1).astype('float32') / 255
X_test = X_test.reshape(-1, 28, 28, 1).astype('float32') / 255

# creating a sequential model
model = Sequential()

# adding convolutional layers with padding
model.add(Conv2D(32, kernel_size=(3,3), padding='same', activation='relu', input_shape=(28,28,1)))
model.add(MaxPooling2D(pool_size=(2,2), strides=(2,2)))

model.add(Conv2D(64, kernel_size=(3,3), padding='same', activation='relu'))
model.add(MaxPooling2D(pool_size=(2,2), strides=(2,2)))

model.add(Conv2D(128, kernel_size=(3,3), padding='same', activation='relu'))
model.add(MaxPooling2D(pool_size=(2,2), strides=(2,2)))

# flattening and adding dense layers
model.add(Flatten())
model.add(Dense(128, activation='relu'))
model.add(Dense(10, activation='softmax'))

# printing model summary
model.summary()


Day 93: Backpropagation in cnns: a quick breakdown

cnns learn by adjusting their parameters through backpropagation. here's how:

  1. forward propagation:
    • input passes through convolution → relu → max pooling → flatten → fully connected layers → output.
  2. backpropagation:
    • calculate loss and propagate errors backward.
    • use the chain rule to compute gradients for weights, biases, and activations.
  3. layer-wise backprop:
    • convolutional layers: errors backpropagate through filters.
    • max pooling: only max-selected neurons contribute gradients.
    • flattening: reshapes gradients before sending them back.

Notes: alt text alt text


Day 94: LeNet5, Cat Vs Dog Classification

LeNet-5, introduced by Yann LeCun in 1989, is one of the earliest convolutional neural networks (CNNs), designed for handwritten character recognition. It consists of seven layers:

alt text

  1. Input Layer: 32×32 grayscale image.
  2. Conv Layer 1 (C1): 6 filters (5×5), output size 28×28×6. alt text
  3. Pooling Layer 1 (S2): Average pooling (2×2), output 14×14×6. alt text
  4. Conv Layer 2 (C3): 16 filters (5×5), selectively connected, output 10×10×16. alt text
  5. Pooling Layer 2 (S4): Average pooling (2×2), output 5×5×16. alt text
  6. Fully Connected (C5): 120 neurons, connected to all 5×5×16 inputs. alt text
  7. Fully Connected (F6): 84 neurons. alt text
  8. Output Layer: Softmax with 10 classes (digits 0-9). alt text

Implementation:

def build_lenet(input_shape):
  # Define Sequential Model
  model = tf.keras.Sequential()
  
  # C1 Convolution Layer
  model.add(tf.keras.layers.Conv2D(filters=6, strides=(1,1), kernel_size=(5,5), activation='tanh', input_shape=input_shape))
  
  # S2 SubSampling Layer
  model.add(tf.keras.layers.AveragePooling2D(pool_size=(2,2), strides=(2,2)))

  # C3 Convolution Layer
  model.add(tf.keras.layers.Conv2D(filters=6, strides=(1,1), kernel_size=(5,5), activation='tanh'))

  # S4 SubSampling Layer
  model.add(tf.keras.layers.AveragePooling2D(pool_size=(2,2), strides=(2,2)))

  # C5 Fully Connected Layer
  model.add(tf.keras.layers.Dense(units=120, activation='tanh'))

  # Flatten the output so that we can connect it with the fully connected layers by converting it into a 1D Array
  model.add(tf.keras.layers.Flatten())

  # FC6 Fully Connected Layers
  model.add(tf.keras.layers.Dense(units=84, activation='tanh'))

  # Output Layer
  model.add(tf.keras.layers.Dense(units=10, activation='softmax'))

  # Compile the Model
  model.compile(loss='categorical_crossentropy', optimizer=tf.keras.optimizers.SGD(lr=0.1, momentum=0.0, decay=0.0), metrics=['accuracy'])

  return model


Day 95: GPU slow than CPU - well in my case?

Today, I trained an MNIST model on both CPU (Ryzen 7 6000) and GPU (RTX 3050 Ti), expecting a significant speedup with the GPU. Instead, the CPU performed slightly faster, and when I tried adding a small CNN, my GPU environment crashed, while the CPU handled it fine (but slower).

  • The GPU initially took longer due to kernel warm up and data transfer overhead.
  • Batch Size Impact – GPUs perform best with larger batch sizes (e.g., 512+), while I used a smaller batch.
  • Data Bottlenecks – My CPU handled data preloading better, while the GPU might have suffered from inefficient memory access.
  • GPU Utilization – The GPU wasn’t fully utilized, likely due to suboptimal parallelism in my setup.

And When I added even a small convo layer, my GPU environment crashed, but the CPU ran it (slowly). here’s what possible reason i estimated?

  • My RTX 3050 Ti has 4GB VRAM, and it might be running out.
  • Despite setting up a dedicated GPU environment (with proper CUDA, cuDNN, and TensorFlow versions), I couldn’t debug this issue today.
  • There might be a compatibility issue between TensorFlow, CUDA, and my system's architecture.

Could it be a CUDA/cuDNN issue or VRAM exhaustion?

I used two separate virtual environments:

  • CPU environment (normal setup)
  • GPU environment (proper versions of CUDA, TensorFlow, and cuDNN) Still, I couldn’t debug the crashes today. Any suggestions on debugging the bit architecture issue or possible TensorFlow config problems? and is it abnormal CPU performing better than GPU in my optimization??

Here's the comparision: GPU vs CPU GPU vs CPU

Some rough notebooks: Notebook: CPU Notebook: GPU


Day 96: Data Augmentation, Pretrained Models

Data augmentation is crucial in machine learning, especially for tasks like computer vision, to enhance model performance and prevent overfitting. It involves applying various transformations to existing data, such as rotation, translation, scaling, flipping, shearing, zooming, and adjusting brightness and contrast. These techniques help in creating a larger and more diverse training dataset, thereby improving model generalization. alt text

Why Use Data Augmentation?

  1. Increased Dataset Size: Enhances model training with more examples.
  2. Regularization: Adds noise to prevent overfitting.
  3. Improved Generalization: Helps models perform better on unseen data.

Notebook: data augmentation on cifar10 frog

Pretrained models in CNN:

Notes: Notes

before deep learning took over, imagenet models relied on classical ml methods like svm, decision trees, and hand-crafted features (think hog, sift, and lbp). this worked, but scaling to millions of images? a nightmare. then alexnet (2012) happened—deep cnns trained with relu activations and dropout on gpus. it crushed traditional methods, slashing classification error rates by half. alt text alt text

vgg (2014) pushed deeper with 3x3 convolutions, proving that simplicity + depth = power. same year, googlenet (inception v1) introduced inception modules—parallel conv layers reducing parameter overhead while boosting efficiency. resnet (2015) then solved the vanishing gradient problem with skip connections, making ultra-deep networks (152 layers!) trainable.

today, pretrained models built on imagenet—resnet, vgg, inception, efficientnet—are the backbone of modern deep learning. explored these today, and yeah, standing on the shoulders of giants makes life easier.

PaperLink: ImageNet Classification with Deep Convolutional Neural Networks

Medium Article on alexnet

Keras pretrained models: https://keras.io/api/applications/ alt text

I will test the examples from that website using their code there in following notebook: Notebook: Pretrained Model Testing

side syb side predictions


Day 97: Visualizing Convolutional Layers, Transfer learning

Today, i'll be using the article : https://machinelearningmastery.com/how-to-visualize-filters-and-feature-maps-in-convolutional-neural-networks/ with respect to following topics.

  • Visualizing Convolutional Layers
  • Pre-fit VGG Model
  • How to Visualize Filters
  • How to Visualize Feature Maps

Notebook: visualizing Layers with elephant image

notes: Notes

Transfer learning:

alt text transfer learning helps train deep learning models efficiently by leveraging pre-trained networks like vgg, resnet, or mobilenet. instead of starting from scratch, we use the convolutional base (which extracts features) and replace the fully connected layers with our own classifier.

two main approaches:

  1. feature extraction – freeze the convolutional layers and train only the new classifier. useful when the target dataset is similar to the original dataset.
  2. fine-tuning – unfreeze some deeper layers and retrain them along with the classifier. this helps when the target dataset is quite different.

Resource: https://www.tensorflow.org/tutorials/images/transfer_learning

Day 98: Keras functional API

How to use keras functional api for deep learning?

Today i went through an article: https://machinelearningmastery.com/keras-functional-api-deep-learning/

The Sequential model API is great for developing deep learning models in most situations, but it also has some limitations.

I started with the Sequential API to build familiarity:

  • Architecture: Simple CNN with Conv2D → MaxPooling → Flatten → Dense → Output.
  • Workflow:
    • Loaded data, normalized pixels (0-1), reshaped images (28x28x1), and one-hot encoded labels.

    • Built a linear stack of layers:

      model = Sequential([
          Conv2D(32, (3,3),
          MaxPooling2D(),
          Flatten(),
          Dense(128),
          Dense(10, activation='softmax')
      ])
      
    • Trained with model.fit(), achieving ~91% validation accuracy in 10 epochs.


2. Transition to Functional API

alt text

I rebuilt the same model using the Functional API to see the syntax shift:

  • Input Layer: Explicitly defined with Input(shape=(28,28,1)).

  • Layer Connections: Layers are chained like functions:

    x = Conv2D(32, (3,3)(input_layer)
    x = MaxPooling2D()(x)
    ...
    
  • Model Definition: Declared inputs/outputs explicitly:

    model_func = Model(inputs=input_layer, outputs=output
    

Truncated — view the full README on GitHub.

100daysofcode
100daysofml
365dayschallenge
365daysofdata
ai
dailylearning
data
datascience
deeplearning
machinelearning
ml

Languages

Jupyter Notebook

99.8%