Machine Learning From Scratch is a collection of machine learning algorithms implemented manually using Python and NumPy. The primary goal of this repository is to understand the mathematical foundations of machine learning by implementing algorithms from first principles instead of relying solely on high-level libraries.
Each algorithm is developed step by step, including data preprocessing, mathematical derivations, model training, evaluation, visualization, and comparison with the equivalent scikit-learn implementation.
- Learn Machine Learning fundamentals
- Understand the mathematics behind algorithms
- Implement machine learning algorithms from scratch
- Compare custom implementations with scikit-learn models
- Build practical machine learning projects
- Strengthen problem-solving and interview skills
MachineLearningFromScratch/
├── Linear-Regression/
├── Logistic-Regression/
├── KNN-Classifier/
├── KNN-Regression/
├── Naive-bayes/
├── Decision-Tree-Classifier/
├── Decision-Tree-Regressor/
├── SVM/
├── Practice/
├── ensemble-learning-scratch/
└── README.md
| Algorithm | Scratch Implementation | Scikit-learn Comparison | Status |
|---|---|---|---|
| Linear Regression | ✅ | ✅ | ✅ Completed |
| Logistic Regression | ✅ | ✅ | ✅ Completed |
| KNN Classification | ✅ | ✅ | ✅ Completed |
| KNN Regression | ✅ | ✅ | ✅ Completed |
| Decision Tree Classifier | ✅ | ✅ | ✅ Completed |
| Decision Tree Regressor | ✅ | ✅ | ✅ Completed |
| ensemble learning techniques | ⏳ | ⏳ | In Progress 📈 |
| SVM (Classifier) | ✅ | ✅ | ✅ Completed |
| Naive Bayes (Gaussian) | ✅ | ✅ | ✅ Completed |
├── ensemble-learning-scratch/
| Technique | Scratch Implementation | Scikit-learn Comparison | Status |
|---|---|---|---|
| Bagging Classifier | ✅ | ✅ | ✅ Completed |
| Bagging Regressor | ✅ | ✅ | ✅ Completed |
| Random Forest Classifier | ✅ | ✅ | ✅ Completed |
| Random Forest Regressor | ✅ | ✅ | ✅ Completed |
| AdaBoost Classifier | ✅ | ✅ | ✅ Completed |
| AdaBoost Regressor | ✅ | ✅ | ✅ Completed |
| Gradient Boosting | ⏳ | ⏳ | ⏳ Pending |
| Stacking | ⏳ | ⏳ | ⏳ Pending |
├── Practice/
The Practice/ directory contains practical experiments and end-to-end workflows focused on model evaluation, preprocessing pipelines, hyperparameter tuning, cross-validation, and ensemble learning using real-world datasets such as the Mobile Price Classification dataset.
Key concepts covered in this directory:
Worked with the basics of Scikit-learn Pipelines to combine preprocessing steps and machine learning models into a single workflow, making the training and prediction process more organized and helping avoid data leakage during cross-validation and hyperparameter tuning.
K-Fold cross-validation technique is used to evaluate model performance across different data splits and assess model stability and generalization.
Systematically searching for optimal model configurations using GridSearchCV and RandomizedSearchCV.
Evaluating performance improvements by comparing baseline models with hyperparameter-tuned estimators, including models such as Support Vector Machines and tree-based models.
Implemented using Scikit-learn's StackingClassifier with three base estimators:
- LogisticRegression
- DecisionTreeClassifier
- SVC
The final estimator is LogisticRegression.
Although Stacking did not improve the model's performance on the Mobile Price Classification dataset, implementing it was important for understanding how multiple different models can work together through a meta-model to make the final prediction. This is a valuable technique in real-world machine learning workflows.
Implemented using Scikit-learn's RandomForestClassifier to understand the concept of Bagging (Bootstrap Aggregating).
Bagging trains multiple models on different bootstrap samples of the training data and combines their predictions. In classification, the final prediction is generally determined through majority voting.
Random Forest extends this idea by combining multiple decision trees while also introducing randomness in feature selection, which helps improve model diversity, robustness, and generalization.
Explored Boosting as another major ensemble learning technique using Scikit-learn.
The practice includes three boosting algorithms:
- AdaBoost
- Gradient Boosting
- XGBoost
These algorithms were explored to understand how multiple weak or sequentially trained learners can be combined to build a stronger predictive model.
The goal was not only to compare their performance, but also to understand how boosting works, how to implement it using Scikit-learn, and when it can be useful in practical machine learning workflows.
Although boosting did not provide a significant performance improvement on the Mobile Price Classification dataset, exploring these algorithms was important for understanding their behavior and learning how ensemble methods can be applied to different datasets.
I also plan to implement these boosting algorithms from scratch in the future to understand their internal mechanisms and mathematical foundations more deeply.
- Python
- NumPy
- Pandas
- Matplotlib
- Seaborn
- Scikit-learn (for validation and performance comparison)
- Jupyter Notebook
Every algorithm in this repository includes:
- Mathematical intuition
- Step-by-step derivation
- From-scratch implementation
- Gradient-based optimization (where applicable)
- Data preprocessing
- Model evaluation
- Visualization
- Performance comparison with the equivalent scikit-learn implementation
- Well-documented Jupyter notebooks
Every custom implementation is evaluated and compared against the corresponding scikit-learn model using the same dataset, preprocessing pipeline, and train-test split.
The comparison includes standard evaluation metrics such as:
- Accuracy (Classification)
- Precision
- Recall
- F1 Score
- R² Score (Regression)
- Adjusted R²
- Mean Absolute Error (MAE)
- Mean Squared Error (MSE)
- Root Mean Squared Error (RMSE)
This comparison helps verify the correctness of each scratch implementation while demonstrating the performance differences between educational implementations and highly optimized production-grade machine learning libraries.
This repository is designed for students, beginners, and aspiring Machine Learning Engineers who want to build a strong understanding of how machine learning algorithms work internally.
Rather than only learning how to use machine learning libraries, the objective is to understand why the algorithms work by implementing them from scratch, validating them against industry-standard implementations, and applying them to real-world datasets.
Ayyan Ahmed
Machine Learning & AI Enthusiast