chandrabhuma/ChemVQA-2K

Dataset

🧪 ChemVQA-2K: A Visual Question Answering Dataset for Molecular Understanding

0

5 commits

1 linked in READMEs

updated Nov 4, 2025

See the code

README

icon4small

🧪 ChemVQA-2K: A Visual Question Answering Dataset for Molecular Understanding 📘 Overview ChemVQA-2K is a novel Visual Question Answering (VQA) dataset designed to bridge chemistry and multimodal AI. It contains approximately 2,000 high-resolution molecular images (512×512) generated from valid SMILES strings, accompanied by 10 structured Q&A pairs per molecule, resulting in ~20,000 image-question-answer triplets. Each image represents a 2D chemical structure rendered using RDKit, while each question tests the model’s ability to reason over molecular features such as formula, atom counts, bonds, functional groups, and polarity.


🧬 Dataset Structure Component Description ChemVQA_2K_images.zip 2,000 molecule renderings (mol_0.png, mol_1.png, …) ChemVQA_2K_full.csv Complete dataset with columns: id, image_name, question, answer

Each record follows: { "id": "mol_123", "image_name": "mol_123.png", "question": "What is the molecular formula of this molecule?", "answer": "C6H6O2" }


🔍 Example Questions Each molecule has multiple Q&A pairs, e.g.: Question Example Answer What is the molecular formula of this molecule? C₂H₅OH What is the molecular weight? 46.07 g/mol How many total atoms are present? 9 Which functional groups are present? Alcohol Is the molecule polar or non-polar? Polar


⚙️ Data Generation Process • Molecules generated by concatenating random organic fragments and validated using RDKit. • Each molecule’s image created with Draw.MolToFile() at 512×512 px resolution. • Functional groups detected via SMARTS pattern matching. • Q&A pairs auto-generated from chemical descriptors (MolWt, CalcMolFormula, substructure matches).


🚀 Intended Use ChemVQA-2K is ideal for: • Fine-tuning Vision-Language Models (VLMs) for scientific visual reasoning. • Developing chemistry-aware question answering systems. • Training vision encoders on molecular visual patterns. • Exploring RL-based visual understanding of chemical structures.


📊 Dataset Statistics Property Value Images 1924 Image resolution 512×512 px Q&A pairs 19240 Functional groups detected 16 File size (approx.) ~25 MB (images + CSVs)


🧠 Potential Research Directions • Multimodal Chemistry Understanding — connecting visual structure with symbolic reasoning. • Scientific Vision-Language Pretraining — use as domain-specific VQA benchmark. • Explainable Chemistry AI — models that describe functional features and molecular properties.

✅ ChemVQA-2K Dataset Benefits the Chemistry Community

  1. Bridges Chemistry and AI Literacy • Helps chemistry students and researchers learn to interact with AI systems using domain-specific visual queries. • Encourages adoption of AI tools in chemical education and research.

  2. Enables Vision-Language Model (VLM) Development for Chemistry • Provides a benchmark for training/fine-tuning multimodal models (e.g., BLIP-2, LLaVA, Qwen-VL) on chemical structure understanding. • Supports the creation of chemistry-aware AI assistants that can "see" molecules and answer questions.

  3. Supports Automated Molecular Analysis • Models trained on this data can automatically extract properties (e.g., atom count, functional groups) from molecular diagrams—useful in digitizing legacy chemical literature or lab notebooks.

  4. Enhances Chemistry Education Tools • Can power interactive learning apps where students upload a molecule image and get instant Q&A feedback (e.g., “How many oxygen atoms?” → “2”). • Useful for self-assessment and virtual tutoring systems.

  5. Facilitates Accessibility in Chemistry • Assists visually impaired researchers/students via multimodal AI that describes molecular structures verbally or in text. • Converts visual chemical information into accessible natural language.

  6. Promotes Reproducible & Scalable Data Curation • Demonstrates a programmatic, open-source pipeline to generate large-scale VQA datasets from SMILES—inspiring similar efforts for reactions, spectra, or crystal structures.

  7. Encourages Domain-Specific AI Benchmarking • Offers a standardized testbed to evaluate how well general VLMs understand scientific imagery vs. models fine-tuned on chemistry data.

  8. Supports Low-Resource Learning • The structured Q&A format is ideal for few-shot or instruction-tuning scenarios, reducing the need for massive labeled datasets.

  9. Integrates with Cheminformatics Workflows • Can be combined with tools like RDKit, PubChem, or ChemSpider to build smart search or annotation systems that answer questions about retrieved molecules.

  10. Fosters Interdisciplinary Collaboration • Creates a common ground for chemists, computer scientists, and educators to collaborate on AI-driven scientific discovery.

📂 File Organization ChemVQA_2K_images.zip / ├── images/ │ ├── mol_0.png │ ├── mol_1.png │ └── ... ├── ChemVQA_2K_full.csv


📜 License 🆓 Open for academic and research use under the CC-BY-4.0 License. Please cite this dataset if you use it in your work.


🧩 Suggested Tags chemistry, VQA, molecule, SMILES, vision-language, AI4Science, RDKit, functional-groups, scientific-vision, chemoinformatics


✨ Visual Summary 🧪 Molecule → 🖼️ Image → ❓ Question → 💡 Answer A structured dataset for visual reasoning over molecular representations.

👨‍🔬 Dataset Authors


🧑‍🏫 Author🏛️ Department🧾 ORCID ID
Dr. B. Chandra Mohan**Dept. of ECE, Bapatla Engineering College, Bapatla–522102🔗 0000-0002-7566-4739
Dr. P. Sumanth Kumar**Dept. of ECE, Bapatla Engineering College, Bapatla–522102🔗 0000-0003-3010-6052
Dr. V. Madhavarao**Dept. of Chemistry, Bapatla Engineering College, Bapatla–522102🔗 0000-0002-0031-5200
Dr. N. Srinivasarao**Dept. of Chemistry, Bapatla Engineering College, Bapatla–522102🔗 0000-0001-7965-6587

Samples from the Dataset image image

chemistry

Contributors

chandrabhuma

5 commits

chandrabhuma/ChemVQA-2K

Dataset

🧪 ChemVQA-2K: A Visual Question Answering Dataset for Molecular Understanding

0

5 commits

1 linked in READMEs

updated Nov 4, 2025

See the code

README

icon4small

🧪 ChemVQA-2K: A Visual Question Answering Dataset for Molecular Understanding 📘 Overview ChemVQA-2K is a novel Visual Question Answering (VQA) dataset designed to bridge chemistry and multimodal AI. It contains approximately 2,000 high-resolution molecular images (512×512) generated from valid SMILES strings, accompanied by 10 structured Q&A pairs per molecule, resulting in ~20,000 image-question-answer triplets. Each image represents a 2D chemical structure rendered using RDKit, while each question tests the model’s ability to reason over molecular features such as formula, atom counts, bonds, functional groups, and polarity.


🧬 Dataset Structure Component Description ChemVQA_2K_images.zip 2,000 molecule renderings (mol_0.png, mol_1.png, …) ChemVQA_2K_full.csv Complete dataset with columns: id, image_name, question, answer

Each record follows: { "id": "mol_123", "image_name": "mol_123.png", "question": "What is the molecular formula of this molecule?", "answer": "C6H6O2" }


🔍 Example Questions Each molecule has multiple Q&A pairs, e.g.: Question Example Answer What is the molecular formula of this molecule? C₂H₅OH What is the molecular weight? 46.07 g/mol How many total atoms are present? 9 Which functional groups are present? Alcohol Is the molecule polar or non-polar? Polar


⚙️ Data Generation Process • Molecules generated by concatenating random organic fragments and validated using RDKit. • Each molecule’s image created with Draw.MolToFile() at 512×512 px resolution. • Functional groups detected via SMARTS pattern matching. • Q&A pairs auto-generated from chemical descriptors (MolWt, CalcMolFormula, substructure matches).


🚀 Intended Use ChemVQA-2K is ideal for: • Fine-tuning Vision-Language Models (VLMs) for scientific visual reasoning. • Developing chemistry-aware question answering systems. • Training vision encoders on molecular visual patterns. • Exploring RL-based visual understanding of chemical structures.


📊 Dataset Statistics Property Value Images 1924 Image resolution 512×512 px Q&A pairs 19240 Functional groups detected 16 File size (approx.) ~25 MB (images + CSVs)


🧠 Potential Research Directions • Multimodal Chemistry Understanding — connecting visual structure with symbolic reasoning. • Scientific Vision-Language Pretraining — use as domain-specific VQA benchmark. • Explainable Chemistry AI — models that describe functional features and molecular properties.

✅ ChemVQA-2K Dataset Benefits the Chemistry Community

  1. Bridges Chemistry and AI Literacy • Helps chemistry students and researchers learn to interact with AI systems using domain-specific visual queries. • Encourages adoption of AI tools in chemical education and research.

  2. Enables Vision-Language Model (VLM) Development for Chemistry • Provides a benchmark for training/fine-tuning multimodal models (e.g., BLIP-2, LLaVA, Qwen-VL) on chemical structure understanding. • Supports the creation of chemistry-aware AI assistants that can "see" molecules and answer questions.

  3. Supports Automated Molecular Analysis • Models trained on this data can automatically extract properties (e.g., atom count, functional groups) from molecular diagrams—useful in digitizing legacy chemical literature or lab notebooks.

  4. Enhances Chemistry Education Tools • Can power interactive learning apps where students upload a molecule image and get instant Q&A feedback (e.g., “How many oxygen atoms?” → “2”). • Useful for self-assessment and virtual tutoring systems.

  5. Facilitates Accessibility in Chemistry • Assists visually impaired researchers/students via multimodal AI that describes molecular structures verbally or in text. • Converts visual chemical information into accessible natural language.

  6. Promotes Reproducible & Scalable Data Curation • Demonstrates a programmatic, open-source pipeline to generate large-scale VQA datasets from SMILES—inspiring similar efforts for reactions, spectra, or crystal structures.

  7. Encourages Domain-Specific AI Benchmarking • Offers a standardized testbed to evaluate how well general VLMs understand scientific imagery vs. models fine-tuned on chemistry data.

  8. Supports Low-Resource Learning • The structured Q&A format is ideal for few-shot or instruction-tuning scenarios, reducing the need for massive labeled datasets.

  9. Integrates with Cheminformatics Workflows • Can be combined with tools like RDKit, PubChem, or ChemSpider to build smart search or annotation systems that answer questions about retrieved molecules.

  10. Fosters Interdisciplinary Collaboration • Creates a common ground for chemists, computer scientists, and educators to collaborate on AI-driven scientific discovery.

📂 File Organization ChemVQA_2K_images.zip / ├── images/ │ ├── mol_0.png │ ├── mol_1.png │ └── ... ├── ChemVQA_2K_full.csv


📜 License 🆓 Open for academic and research use under the CC-BY-4.0 License. Please cite this dataset if you use it in your work.


🧩 Suggested Tags chemistry, VQA, molecule, SMILES, vision-language, AI4Science, RDKit, functional-groups, scientific-vision, chemoinformatics


✨ Visual Summary 🧪 Molecule → 🖼️ Image → ❓ Question → 💡 Answer A structured dataset for visual reasoning over molecular representations.

👨‍🔬 Dataset Authors


🧑‍🏫 Author🏛️ Department🧾 ORCID ID
Dr. B. Chandra Mohan**Dept. of ECE, Bapatla Engineering College, Bapatla–522102🔗 0000-0002-7566-4739
Dr. P. Sumanth Kumar**Dept. of ECE, Bapatla Engineering College, Bapatla–522102🔗 0000-0003-3010-6052
Dr. V. Madhavarao**Dept. of Chemistry, Bapatla Engineering College, Bapatla–522102🔗 0000-0002-0031-5200
Dr. N. Srinivasarao**Dept. of Chemistry, Bapatla Engineering College, Bapatla–522102🔗 0000-0001-7965-6587

Samples from the Dataset image image

chemistry

Contributors

chandrabhuma

5 commits