This repository contains the code and the data collected from a large database of Arabic journals (https://asjp.cerist.dz/en).
The folder src contains the script used to scrape the website in a file called scrape.py.
The folder notebooks contains the notebook explore_and_extract.ipynb that filters and collects the statistics of the collected papers. The following statistics are taken from this file:
The folder articles contains the scraped articles grouped by their journal, followed by their volume number and year, followed by their issue number and date. It contains boths, the json metadata file and the paper pdf.
Jupyter Notebook
100.0%
This repository contains the code and the data collected from a large database of Arabic journals (https://asjp.cerist.dz/en).
The folder src contains the script used to scrape the website in a file called scrape.py.
The folder notebooks contains the notebook explore_and_extract.ipynb that filters and collects the statistics of the collected papers. The following statistics are taken from this file:
The folder articles contains the scraped articles grouped by their journal, followed by their volume number and year, followed by their issue number and date. It contains boths, the json metadata file and the paper pdf.
Jupyter Notebook
100.0%