A Bibliometric and Scientometric Python Library Powered with Artificial Intelligence Tools
Python
222
299 commits
updated Jun 21, 2026
PEREIRA, V.; BASILIO, M.P.; SANTOS, C.H.T. (2025). PyBibX: A Python Library for Bibliometric and Scientometric Analysis Powered with Artificial Intelligence Tools. Data Technologies and Applications. Vol. 59, Iss. 2, pp. 302-337. doi: https://doi.org/10.1108/DTA-08-2023-0461
New to Python or prefer a graphical interface? The pybibx Web App lets you run your analysis in clicks, not lines of code.
import pybibx
# Start the web service using:
pybibx.web_app()
# Terminate the web service using:
pybibx.web_stop()
This Google Colab Demo is intended for quick demos only. For the best experience, run the Web UI locally or open it directly in a full browser.
A Bibliometric and Scientometric Python library that uses raw files generated by Scopus (.bib files or .csv files), WoS (Web of Science) (.bib files), PubMed (.txt files), and OpenAlex. For OpenAlex, the library supports JSON files obtained through the OpenAlex API and CSV (standard) files exported from the OpenAlex website. Also powered by Advanced AI technologies for analyzing bibliometric, scientometric, and textual data.
To export the correct file formats from Scopus, Web of Science, PubMed, and OpenAlex, follow these steps:
a) Scopus: Search, select articles, click Export, choose BibTeX or CSV, select all fields, and click Export again. When using the CSV format, the exported files include the References for the articles.
b) WoS: Search, select articles, click Export, choose Save to Other File Formats, select BibTeX, choose all fields, and click Send.
c) PubMed: Search, select articles, click Save, choose PubMed format, and click Save to download a .txt file. The exported files do not contain the References for the articles.
d) OpenAlex: API Option -> Retrieve records through the OpenAlex REST API and save the response as a .json file. OpenAlex's official programmatic interface is its API, designed for structured, machine-readable access (OpenAlex Developers). Website Option -> Search and filter records in the OpenAlex website, click Export, and choose CSV (standard). In particular, from the Website option, the exported files do not contain the References for the articles.
e) Compare: Users can compare records exported from Scopus, WoS, PubMed, and OpenAlex using pybibx.compare_sources(). This function identifies shared and unique documents across databases, evaluates source overlap, reports metadata completeness, indicates which documents were matched, and estimates the contribution of each database to the final corpus.
.bib files or .csv files), WoS (.bib files), PubMed (.txt files), and OpenAlex (API JSON files and website CSV (standard) files).pip install pybibx
This section indicates the libraries that inspired pybibx
BERT (https://smrzr.io/):
a) Github: https://github.com/dmmiller612/bert-extractive-summarizer
b) Paper: DEREK, M. (2019). Leveraging BERT for Extractive Text Summarization on Lectures. arXiv. doi: https://doi.org/10.48550/arXiv.1906.04165
SciBERT (https://huggingface.co/allenai/scibert_scivocab_uncased):
a) Github: https://github.com/allenai/scibert
b) Paper: BELTAGY, I.;, LO, K.; COHAN, A. (2019). SCIBERT: A Pretrained Language Model for Scientific Text. arXiv. doi: https://doi.org/10.48550/arXiv.1903.10676
BERTopic (https://maartengr.github.io/BERTopic/index.html):
a) Github: https://github.com/MaartenGr/BERTopic
b) Paper: GROOTENDORST, M. (2022). BERTopic: Neural Topic Modeling with a Class-based TF-IDF Procedure. arXiv. doi: https://doi.org/10.48550/arXiv.2203.05794
Bibliometrix (https://www.bibliometrix.org/home/):
a) Github: https://github.com/massimoaria/bibliometrix
b) Paper: ARIA, M.; CUCCURULLO, C. (2017). Bibliometrix: An R-tool for Comprehensive Science Mapping Analysis. Journal of Informetrics, 11(4), 959-975. doi: https://doi.org/10.1016/j.joi.2017.08.007
Gemini (https://gemini.google.com/app):
a) Github: https://github.com/google-gemini
b) Paper: Gemini Team Google (2024). Gemini: A Family of Highly Capable Multimodal Models. arXiv. doi: https://arxiv.org/abs/2312.11805
Gensim (https://radimrehurek.com/gensim/):
a) Github: https://github.com/piskvorky/gensim
b) Paper: REHUREK, R.; SOJKA, P. (2010). Software Framework for Topic Modelling with Large Corpora. LREC 2010. doi: https://doi.org/10.13140/2.1.2393.1847
chatGPT (https://chat.openai.com/chat):
a) Github: https://github.com/openai
b) Paper: OPENAI. (2023). GPT-4 Technical Report. arXiv. doi: https://doi.org/10.48550/arXiv.2303.08774
Metaknowledge (http://www.networkslab.org/metaknowledge):
a) Github: https://github.com/UWNETLAB/metaknowledge
b) Paper: McILROY-YOUNG, R.; McLEVEY, J.; ANDERSON, J. (2015). Metaknowledge: Open Source Software for Social Networks, Bibliometrics, and Sociology of Knowledge Research.
SentenceTransformers (https://www.sbert.net/):
a) Github: https://github.com/UKPLab/sentence-transformers
b) Paper: REIMERS, N.; GUREVYCH, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv. doi: https://arxiv.org/abs/1908.10084
PEGASUS (https://ai.googleblog.com/2020/06/pegasus-state-of-art-model-for.html?m=1):
a) Github: https://github.com/huggingface/transformers
b) Paper: ZHANG, J.; ZHAO, Y.; SALEH, M.; LIU, P.J. (2019). PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. arXiv. doi: https://doi.org/10.48550/arXiv.1912.08777
And to all the people who helped to improve or correct the code. Thank you very much!
A Bibliometric and Scientometric Python Library Powered with Artificial Intelligence Tools
Python
222
299 commits
updated Jun 21, 2026
PEREIRA, V.; BASILIO, M.P.; SANTOS, C.H.T. (2025). PyBibX: A Python Library for Bibliometric and Scientometric Analysis Powered with Artificial Intelligence Tools. Data Technologies and Applications. Vol. 59, Iss. 2, pp. 302-337. doi: https://doi.org/10.1108/DTA-08-2023-0461
New to Python or prefer a graphical interface? The pybibx Web App lets you run your analysis in clicks, not lines of code.
import pybibx
# Start the web service using:
pybibx.web_app()
# Terminate the web service using:
pybibx.web_stop()
This Google Colab Demo is intended for quick demos only. For the best experience, run the Web UI locally or open it directly in a full browser.
A Bibliometric and Scientometric Python library that uses raw files generated by Scopus (.bib files or .csv files), WoS (Web of Science) (.bib files), PubMed (.txt files), and OpenAlex. For OpenAlex, the library supports JSON files obtained through the OpenAlex API and CSV (standard) files exported from the OpenAlex website. Also powered by Advanced AI technologies for analyzing bibliometric, scientometric, and textual data.
To export the correct file formats from Scopus, Web of Science, PubMed, and OpenAlex, follow these steps:
a) Scopus: Search, select articles, click Export, choose BibTeX or CSV, select all fields, and click Export again. When using the CSV format, the exported files include the References for the articles.
b) WoS: Search, select articles, click Export, choose Save to Other File Formats, select BibTeX, choose all fields, and click Send.
c) PubMed: Search, select articles, click Save, choose PubMed format, and click Save to download a .txt file. The exported files do not contain the References for the articles.
d) OpenAlex: API Option -> Retrieve records through the OpenAlex REST API and save the response as a .json file. OpenAlex's official programmatic interface is its API, designed for structured, machine-readable access (OpenAlex Developers). Website Option -> Search and filter records in the OpenAlex website, click Export, and choose CSV (standard). In particular, from the Website option, the exported files do not contain the References for the articles.
e) Compare: Users can compare records exported from Scopus, WoS, PubMed, and OpenAlex using pybibx.compare_sources(). This function identifies shared and unique documents across databases, evaluates source overlap, reports metadata completeness, indicates which documents were matched, and estimates the contribution of each database to the final corpus.
.bib files or .csv files), WoS (.bib files), PubMed (.txt files), and OpenAlex (API JSON files and website CSV (standard) files).pip install pybibx
This section indicates the libraries that inspired pybibx
BERT (https://smrzr.io/):
a) Github: https://github.com/dmmiller612/bert-extractive-summarizer
b) Paper: DEREK, M. (2019). Leveraging BERT for Extractive Text Summarization on Lectures. arXiv. doi: https://doi.org/10.48550/arXiv.1906.04165
SciBERT (https://huggingface.co/allenai/scibert_scivocab_uncased):
a) Github: https://github.com/allenai/scibert
b) Paper: BELTAGY, I.;, LO, K.; COHAN, A. (2019). SCIBERT: A Pretrained Language Model for Scientific Text. arXiv. doi: https://doi.org/10.48550/arXiv.1903.10676
BERTopic (https://maartengr.github.io/BERTopic/index.html):
a) Github: https://github.com/MaartenGr/BERTopic
b) Paper: GROOTENDORST, M. (2022). BERTopic: Neural Topic Modeling with a Class-based TF-IDF Procedure. arXiv. doi: https://doi.org/10.48550/arXiv.2203.05794
Bibliometrix (https://www.bibliometrix.org/home/):
a) Github: https://github.com/massimoaria/bibliometrix
b) Paper: ARIA, M.; CUCCURULLO, C. (2017). Bibliometrix: An R-tool for Comprehensive Science Mapping Analysis. Journal of Informetrics, 11(4), 959-975. doi: https://doi.org/10.1016/j.joi.2017.08.007
Gemini (https://gemini.google.com/app):
a) Github: https://github.com/google-gemini
b) Paper: Gemini Team Google (2024). Gemini: A Family of Highly Capable Multimodal Models. arXiv. doi: https://arxiv.org/abs/2312.11805
Gensim (https://radimrehurek.com/gensim/):
a) Github: https://github.com/piskvorky/gensim
b) Paper: REHUREK, R.; SOJKA, P. (2010). Software Framework for Topic Modelling with Large Corpora. LREC 2010. doi: https://doi.org/10.13140/2.1.2393.1847
chatGPT (https://chat.openai.com/chat):
a) Github: https://github.com/openai
b) Paper: OPENAI. (2023). GPT-4 Technical Report. arXiv. doi: https://doi.org/10.48550/arXiv.2303.08774
Metaknowledge (http://www.networkslab.org/metaknowledge):
a) Github: https://github.com/UWNETLAB/metaknowledge
b) Paper: McILROY-YOUNG, R.; McLEVEY, J.; ANDERSON, J. (2015). Metaknowledge: Open Source Software for Social Networks, Bibliometrics, and Sociology of Knowledge Research.
SentenceTransformers (https://www.sbert.net/):
a) Github: https://github.com/UKPLab/sentence-transformers
b) Paper: REIMERS, N.; GUREVYCH, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv. doi: https://arxiv.org/abs/1908.10084
PEGASUS (https://ai.googleblog.com/2020/06/pegasus-state-of-art-model-for.html?m=1):
a) Github: https://github.com/huggingface/transformers
b) Paper: ZHANG, J.; ZHAO, Y.; SALEH, M.; LIU, P.J. (2019). PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. arXiv. doi: https://doi.org/10.48550/arXiv.1912.08777
And to all the people who helped to improve or correct the code. Thank you very much!