Jupyter Notebook
2
71 commits
updated Apr 27, 2024
In this work we have passed through different phases from fetching the data, to process it, shuffle the data for the next stage of preparation and run the model on. And how we come over these stages is decribed below.
First important, you need to install the requirements using snip code below, in case of missed libraries error try to pip3 install "name of library" like in snap code below.
pip3 install -r requirements.txt
# in case of missed libraries error
pip3 install nltk
# Then open Jupter notebook and run Deployment FastAPI Notebook to launch notebook "ipython notebook"
# request "http://0.0.0.0:8000/docs"
# Watch the video below to know how to use it.

Second important, We have used word2vec for word representation, we used some pretrained word2vec published AraVec by Eng. Abo Bakr and others, we used lightweight ones which use just unigrams, and as well as our pretrained one. So to come over this:
The dataset after fetching it from API is larger than 50 Miga, and we can not pushed to github so we have compress it, and you can extract it, but you can also run into the models directly by unzip the trained dataset inside "train" direction , and escape either fetching data or preprocess it, we save our work after each stage.
Helpful script to keep of some functions that we use in different files.
To see this file, check the configs.py script script, its fully documented.
We have the original dataset without the text column, which what we will use for feature engineering to predict which text belong to which dialect. So for that we have design our pipeline for fetching the data using the ids of the original dataset, once we get all of text related to all ids we save new csv file with the new text column.
To see this stage of fetching data, check the fetch_data.py script script, its fully documented, and to get overview of the result from this stage check Data Fetching.ipynb notebook.
Quick intuition about what we got from this stage

The first thing we started with after fetching the data is to process the text we got, and at this point as we dealing with Arabic text we start the cleaning process.
To see this stage of preprocessing data, check the data_preprocess.py script and to know how we shuffle the data, check the data_shuffling_split.py script its fully documented, and to get overview of the result from this stage check Data pre-processing.ipynb notebook.
Quick intuition about preprocessing and shuffling

Comparison the different splitting with the original data
|
After we have the data and splited with Stratified splitting, we have to get features (numbers) from that text as any ML models or DL models just fit to numbers.
So we use one of the modern approaches to get representation of each word "Word2Vec", instead of classical approaches like Vectorization or tf-idf and the the problems it runs from, either in encoding the context or the curse of dimension.
Once we got these word representation and run the required functions to convert from word representation into text representation, we run into the modeling phase and train different models as table below present the different results we got.
To see this stage of Data preparation & Modeling, check the ml_modeling.py script, and keras_models.py its fully documented, and to get overview of the result from this stage after run the notebooks check ML Models Train.ipynb notebook, and and DL Models Train.ipynb notebook
To see result of modeling, check the Compare ML Models.ipynb Notebook, or Compare DL Models.ipynb Notebook.
| Word2Vec Used | Model Used | F1-score |
|---|---|---|
| AraVec word Representation | AdaBoostClassifier | .32 |
| AraVec word Representation | Logistic Regression | .36 |
| Our word Representation | Logistic Regression | .41 |
| Our word Representation | GradientBoostingClassifier | .22 |
| Our word Representation | LinearSVC | .27 |
| Word2Vec Used | Model, Optimizer, Batching | F1-score |
|---|---|---|
| AraVec word Representation | LSTM, SGD, No Batch | .47 |
| AraVec word Representation | LSTM, SGD, Batch | .46 |
| AraVec word Representation | LSTM, Adam, No Batch | .47 |
| AraVec word Representation | LSTM, Adam, Batch | .47 |
| Our word Representation | LSTM, SGD, No Batch | .51 |
| Our word Representation | LSTM, SGD, Batch | .49 |
| Our word Representation | LSTM, Adam, No Batch | .51 |
Overview about the work till splitting data

Overview about Modeling using Tensorboard

Jupyter Notebook
78.0%
Python
22.0%
Jupyter Notebook
2
71 commits
updated Apr 27, 2024
In this work we have passed through different phases from fetching the data, to process it, shuffle the data for the next stage of preparation and run the model on. And how we come over these stages is decribed below.
First important, you need to install the requirements using snip code below, in case of missed libraries error try to pip3 install "name of library" like in snap code below.
pip3 install -r requirements.txt
# in case of missed libraries error
pip3 install nltk
# Then open Jupter notebook and run Deployment FastAPI Notebook to launch notebook "ipython notebook"
# request "http://0.0.0.0:8000/docs"
# Watch the video below to know how to use it.

Second important, We have used word2vec for word representation, we used some pretrained word2vec published AraVec by Eng. Abo Bakr and others, we used lightweight ones which use just unigrams, and as well as our pretrained one. So to come over this:
The dataset after fetching it from API is larger than 50 Miga, and we can not pushed to github so we have compress it, and you can extract it, but you can also run into the models directly by unzip the trained dataset inside "train" direction , and escape either fetching data or preprocess it, we save our work after each stage.
Helpful script to keep of some functions that we use in different files.
To see this file, check the configs.py script script, its fully documented.
We have the original dataset without the text column, which what we will use for feature engineering to predict which text belong to which dialect. So for that we have design our pipeline for fetching the data using the ids of the original dataset, once we get all of text related to all ids we save new csv file with the new text column.
To see this stage of fetching data, check the fetch_data.py script script, its fully documented, and to get overview of the result from this stage check Data Fetching.ipynb notebook.
Quick intuition about what we got from this stage

The first thing we started with after fetching the data is to process the text we got, and at this point as we dealing with Arabic text we start the cleaning process.
To see this stage of preprocessing data, check the data_preprocess.py script and to know how we shuffle the data, check the data_shuffling_split.py script its fully documented, and to get overview of the result from this stage check Data pre-processing.ipynb notebook.
Quick intuition about preprocessing and shuffling

Comparison the different splitting with the original data
|
After we have the data and splited with Stratified splitting, we have to get features (numbers) from that text as any ML models or DL models just fit to numbers.
So we use one of the modern approaches to get representation of each word "Word2Vec", instead of classical approaches like Vectorization or tf-idf and the the problems it runs from, either in encoding the context or the curse of dimension.
Once we got these word representation and run the required functions to convert from word representation into text representation, we run into the modeling phase and train different models as table below present the different results we got.
To see this stage of Data preparation & Modeling, check the ml_modeling.py script, and keras_models.py its fully documented, and to get overview of the result from this stage after run the notebooks check ML Models Train.ipynb notebook, and and DL Models Train.ipynb notebook
To see result of modeling, check the Compare ML Models.ipynb Notebook, or Compare DL Models.ipynb Notebook.
| Word2Vec Used | Model Used | F1-score |
|---|---|---|
| AraVec word Representation | AdaBoostClassifier | .32 |
| AraVec word Representation | Logistic Regression | .36 |
| Our word Representation | Logistic Regression | .41 |
| Our word Representation | GradientBoostingClassifier | .22 |
| Our word Representation | LinearSVC | .27 |
| Word2Vec Used | Model, Optimizer, Batching | F1-score |
|---|---|---|
| AraVec word Representation | LSTM, SGD, No Batch | .47 |
| AraVec word Representation | LSTM, SGD, Batch | .46 |
| AraVec word Representation | LSTM, Adam, No Batch | .47 |
| AraVec word Representation | LSTM, Adam, Batch | .47 |
| Our word Representation | LSTM, SGD, No Batch | .51 |
| Our word Representation | LSTM, SGD, Batch | .49 |
| Our word Representation | LSTM, Adam, No Batch | .51 |
Overview about the work till splitting data

Overview about Modeling using Tensorboard

Jupyter Notebook
78.0%
Python
22.0%