TheLuoFengLab/MolLangData

A tool for genetating molecular structure-language description and the resulting large-scale dataset.

Python

2

17 commits

updated Sep 25, 2026

See the code

README

MolLangData: A Large-Scale Dataset for Molecular Structure–Language Description

GitHub stars GitHub forks GitHub issues Discord arXiv Hugging Face License: MIT

MolLangData is a large-scale dataset of molecular structures paired with natural-language descriptions, generated via a rule-regularized method. It supports training and evaluating models for molecular structure–language alignment.


Table of contents


Project Resources

ResourceLinksDescription
MolLangDataGitHub · Hugging FaceMain dataset on Hugging Face (~163k samples). We are actively expanding beyond this release.
MolLangBench (ICLR 2026)GitHub · Hugging FaceHuman-curated benchmark for molecular structure recognition, editing, and generation. The generation task aligns with structural description in this work and serves as a standard, validated evaluation.
LangMolDiodeGitHubLanguage-conditional molecule generator trained with SFT and reinforcement learning on MolLangData.

(back to top)


OPSIN (IUPAC → XML / SMILES)

We use a customized OPSIN fork that adds complete XML structure metadata for building prompts:

  • Fork: feiyang-cai/opsin_mollangdata
  • The repository is for reference; a compiled JAR is provided for the single-molecule workflow below.
  • The OPSIN fork README and usage will be updated later (see TODO).

(back to top)


Requirements

  • Python 3.8+
  • Java (JRE or JDK on PATH)
  • Python dependencies: pip install -r requirements.txt (installs openai, rdkit, and related packages)

(back to top)


Single molecule: get prompt and description from IUPAC

The script get_prompt_description_from_iupac.py takes one IUPAC name, runs OPSIN to obtain XML and SMILES, computes a difficulty level (easy / medium / hard), builds the prompt, and optionally calls an LLM to generate a structure description.

Files involved

PathPurpose
get_prompt_description_from_iupac.pyMain script
jar/Place the OPSIN JAR here: jar/opsin-core-*-jar-with-dependencies.jar
prompts/smiles_iupac_metadata_v13/Prompt template and semantic sections (default, bridged/fused/spiro rings)
config/llm_config.jsonLLM backend (Azure/OpenAI), model, reasoning effort per difficulty, timeouts (used only with --get-description)

Run (from repo root)

  • Prompt only (no LLM):

    python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone"
    
  • With LLM description (model and reasoning chosen by difficulty; backend from config):

    python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone" --get-description
    
  • Write output to a folder (prompt.md, descriptions.txt, generation_info.json):

    python3 get_prompt_description_from_iupac.py "ethane" --get-description -o out/
    
Options
OptionDescription
--opsin-jar PATHPath to OPSIN JAR (default: jar/opsin-core-*-jar-with-dependencies.jar)
--prompts-dir DIRPrompt templates directory (default: prompts/smiles_iupac_metadata_v13)
-o DIR / --output DIROutput folder: writes prompt.md; with --get-description also descriptions.txt and generation_info.json
--get-descriptionCall LLM to generate structure description (uses config/llm_config.json unless overridden)
--llm-config FILELLM config JSON (default: config/llm_config.json)
--model MODELOverride LLM model
--reasoning-effort EFFORTOverride LLM reasoning effort
Output
  • Without -o: Prints IUPAC, OPSIN status, SMILES, difficulty level, and (if --get-description) the description and prompt to stdout.
  • With -o DIR: Writes DIR/prompt.md; with --get-description also writes DIR/descriptions.txt and DIR/generation_info.json (backend, model, reasoning effort, duration, etc.).

(back to top)


Dataset generation pipeline

This section describes how to obtain or regenerate the MolLangData dataset. Pre-sampled and pre-processed data are available on Box (recommended); you can also run the full pipeline from raw PubChem data. Note: LLM outputs are non-deterministic—re-running the pipeline will not reproduce the same descriptions.

Download from Box to start from Step 4 (or Step 3 if you start from the sampled TSV only). Step 3 parsing output and Response API JSONL job files are also on Box, so you can skip to Step 4 or Step 5 as needed.

Box resourceContents
MolLangData PubChem sampled TSV• Sampled TSV: 8 rounds, 200k samples per round (round_0/sampled.tsv, round_1/sampled.tsv, …).
• Parsing output (Step 3): e.g. round_1/parsing_out_mollangdata_0.1.3.
• Response API JSONL (Step 4): e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs for starting at Step 5.

The published MolLangData dataset on Hugging Face corresponds to round_0 data.

Pipeline overview

StepDescription
1SDF → TSV — convert PubChem SDF to TSV (batch_prompt_generation/miscellaneous/1_sdf_to_tsv)
2Sampling — deterministic sampling from TSV into chunks (e.g. 200k per round) (batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm)
3OPSIN + MolLangData parsing — full XML structure data per sampled.tsv
4Create batch prompt JSONL — build LLM job files from parsing output
5Run LLM jobs — submit to obtain structure descriptions (OpenAI Batch or one-by-one)

Steps 1 and 2 can be time-consuming; we strongly recommend using the pre-sampled Box data above. Scripts: batch_prompt_generation/miscellaneous/1_sdf_to_tsv and batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm.


Step 3 — OPSIN + MolLangData parsing

Run the custom OPSIN tool on each sampled chunk to obtain full XML structure data. Parsing output for all rounds is on Box (see above); you can skip this step when using that data.

  • Input: A folder containing sampled.tsv (e.g. a round folder from Box).
  • Output: Under that folder (or --out-dir) as parsing_out_mollangdata_<version>/, containing opsin_xml/, mollangdata_xml/, and parsing_results.tsv.

The script auto-detects the OPSIN JAR in the repo (opsin/opsin-core/target/ or jar/); override with --opsin-jar.

Run (from repo root; <sample_folder> is the folder containing sampled.tsv):

python3 batch_prompt_generation/3_run_opsin_mollangdata_on_sampled.py <sample_folder>
Step 3 — Arguments
ArgumentDescription
sample_folderFolder containing sampled.tsv (required). TSV must have PUBCHEM_COMPOUND_CID, SMILES, canonical_smiles, and an IUPAC column (see --iupac-column).
--opsin-jar PATHPath to OPSIN MolLangData JAR. Default: auto-detect under repo opsin/opsin-core/target/ or jar/.
--out-dir DIROutput directory. Default: <sample_folder>/parsing_out; output is under parsing_out_mollangdata_<version>/.
--iupac-column NAMEIUPAC column name. Default: PUBCHEM_IUPAC_SYSTEMATIC_NAME.
--timeout-s NTimeout in seconds per molecule (default: 30).
--max-rows NProcess only first N rows (for testing).
--verboseVerbose logging.

Step 4 — Create batch prompt JSONL

Build LLM job files from Step 3 output. The script assigns difficulty (easy/medium/hard), builds the prompt with the same dynamic template as the single-molecule script (including fused/spiro/bridged sections when applicable), and routes model and reasoning effort from config/llm_config.json. Output is JSONL (and optional per-prompt TXT) for one-by-one or batch LLM runs. Supports Azure and OpenAI, and both Chat Completions and Responses APIs. Pre-built prompts for all rounds are on Box (e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs).

Step 4 — Backend / API recommendations
  • Azure: GPT-5.2 does not support batch on Azure as of 2/13/2026. Use one-by-one with Responses API (--api-format responses). This is what we used for data generation.
  • OpenAI: Both Responses and Chat Completions work. The Batch API with the Responses API has been reported to cause tasks to run repeatedly and incur extra cost (Batch API task runs repeatedly with gpt-5.2-pro). We recommend Chat Completions (--api-format chat_completions) when using OpenAI.

Run (from repo root). <input_folder> must be Step 3 output containing parsing_results.tsv (e.g. parsing_out_mollangdata_0.1.3). Example: on Box.

python3 batch_prompt_generation/4_create_batch_prompt_jsonl.py <input_folder> <output_folder> [prompts_folder] --api-format responses
Step 4 — Arguments
ArgumentDescription
input_folderFolder containing parsing_results.tsv (and mollangdata_xml/ or opsin_xml/). Typically Step 3 output (e.g. parsing_out_mollangdata_0.1.3).
output_folderDirectory for JSONL output (and optional TXT). Files are split by model and reasoning effort.
prompts_folderOptional; default: prompts/smiles_iupac_metadata_v13.
--llm-config FILELLM config (default: config/llm_config.json). Must have easy, medium, hard with model and reasoning_effort.
--api-formatresponses or chat_completions. Default: chat_completions. Use responses for Responses API.
--chat-completions-url URLURL for chat-completions (default: /v1/chat/completions).
--responses-url URLURL for responses (default: /v1/responses).
--xml-typemollangdata_xml or opsin_xml (default: mollangdata_xml).
--exclude-iupacOmit IUPAC Name: block from prompts.
--exclude-xmlOmit XML Metadata: block from prompts.
--allow-dotsInclude rows whose SMILES contain a dot (disconnected component).
--sample-size NRandomly sample N rows (for testing).
--random-seed NSeed for sampling (default: 533).
--custom-prefix STRPrefix for compound IDs in custom_id.
--no-txtDo not write per-prompt .txt; JSONL only.

Step 5 — Run LLM jobs (structure descriptions)

⚠️ Cost warning: Step 5 calls the LLM API for every prompt and can incur large costs (e.g. hundreds of thousands of requests). Check usage and billing before running. We recommend using the published MolLangData dataset unless you need to regenerate or extend descriptions.

Two options:

OptionScriptUse case
5a — OpenAI Batch5a_submit_openai_batch_jobs.pyUpload JSONL to OpenAI Batch API, wait, and retrieve. OpenAI only. Batch pricing is half of on-demand. Requires OPENAI_API_KEY.
5b — One-by-one5b_run_requests_one_by_one.py, 5b_run_tmux_jobs.shSequential or parallel (e.g. via tmux). Works with Azure and OpenAI. We use 5b because batch is not supported for GPT-5.2 on Azure (as of 2/13/2026). Responses API is recommended for background mode (submit → poll).
Step 5 — Recommendations
  • OpenAI: Prefer 5a (Batch) — half the price. Use 5b for testing or when batch is unavailable.
  • Azure: Use 5b (one-by-one); batch is not supported for GPT-5.2 on Azure.
  • API format: For 5b, prefer Responses API (background mode is more stable). Generate job JSONL with --api-format responses in Step 4 when using 5b.

Run (from repo root)

  • 5a — OpenAI Batch (submit, wait, retrieve). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.

    python3 batch_prompt_generation/5a_submit_openai_batch_jobs.py <input> --output-dir ./batch_results
    
  • 5b — One-by-one, single process (Azure, resume with pool). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.

    python3 batch_prompt_generation/5b_run_requests_one_by_one.py <input> --backend azure --output-dir ./runs/run1 --resume --pool-folder ./runs/pool
    
  • 5b — One-by-one, multiple processes (e.g. tmux; split one JSONL into 4 workers). Replace <input.jsonl> with a Step 4 .jsonl file:

    bash ./batch_prompt_generation/5b_run_tmux_jobs.sh -f <input.jsonl> -s myrun -n 4 -o ./runs_out
    
Step 5a — Arguments (OpenAI Batch)
ArgumentDescription
inputSingle .jsonl file or directory of .jsonl (from Step 4).
--output-dir DIRWhere to save retrieved output/error files (default: batch_results).
--no-waitSubmit only; do not wait or download (default: wait and retrieve).
--poll-interval NSeconds between status polls (default: 60).
--max-requests NMax requests per batch file when splitting (default: 50000).
--completion-window DURBatch completion window (default: 24h).
--dry-runList files and endpoints only; no upload or batch creation.
Step 5b — Arguments

One-by-one runner (5b_run_requests_one_by_one.py)

ArgumentDescription
input(Required.) A single .jsonl file or a directory of .jsonl files from Step 4. Each line is one LLM request.
--output-dir DIR(Required.) Directory for outputs: results_<backend>.jsonl, stats_<backend>.jsonl, and request_outputs/<custom_id>/ per request.
--backendazure or openai. Selects which LLM API to call.
--llm-config FILEPath to LLM config JSON (default: config/llm_config.json). Used for endpoint and auth.
--resumeIf set, skip requests that already have an entry in --output-dir and continue from the rest. Use after an interrupted run.
--pool-folder DIRWhen using Responses API background mode: folder to store in-flight request IDs for polling. Required for --resume to work correctly.
--model MODELOverride the model from config (e.g. gpt-5.2).
--reasoning-effort LEVELOverride reasoning effort from config (e.g. high, xhigh).
--max-requests NProcess at most N requests (for testing). Omit to process all.
--timeout NRequest timeout in seconds.
--poll-interval NSeconds between polls when using Responses API background mode.
--max-retries N, --retry-sleep SECRetry failed requests up to N times, waiting SEC seconds between retries.
--sleep SECOptional delay in seconds between requests (rate limiting).
--no-tqdmDisable progress bar.

Tmux multi-process (5b_run_tmux_jobs.sh)

Splits one JSONL into N chunks and runs N workers in tmux windows, each calling 5b_run_requests_one_by_one.py on one chunk. For Azure, each worker can use a different deployment (e.g. gpt-5.2, gpt-5.2-2) via -i START_IDX.

ArgumentDescription
-f FILE(Required.) Input JSONL file from Step 4.
-s SESSION_NAME(Required.) Tmux session name for the worker windows.
-n N(Required.) Number of splits = number of parallel workers (e.g. 4 → 4 tmux windows).
-o OUTPUT_BASEBase output directory; each worker gets a subdir (e.g. ./runs_out).
-p POOL_FOLDERPool folder for Responses API background mode (passed to each worker).
-i START_IDX(Azure.) Deployment start index so workers use different deployments (e.g. 0 → gpt-5.2, 1 → gpt-5.2-2).
--llm-config, --backend, --model, --reasoning-effortPassed through to each worker.
-t TIMEOUTRequest timeout in seconds.
--no-tqdmDisable progress bar in workers.

(back to top)


Dataset on Hugging Face: structure and validation

The MolLangData dataset on Hugging Face is provided in two configurations with the following structure and validation results.

Dataset structure

  1. validated_data

    • All validated data (2k samples), including descriptions that passed validation and those that did not.
    • See validation precision in the Dataset statistics table below.
  2. generated_data

    • All generated data from round 0, excluding the validated subset.
    • These samples are not validated.

Dataset statistics

DifficultyModelReasoning effortGenerated samplesValidated samplesValidation precision
EasyGPT-5.2high105,085 (65.2%)1,317 (65.8%)1,300 (98.7%)
MediumGPT-5.2xhigh40,916 (25.4%)496 (24.8%)492 (99.2%)
HardGPT-5.2xhigh15,110 (9.4%)187 (9.4%)180 (96.3%)
Overall——161,1112,0001,972 (98.6%)

(back to top)


TODO & roadmap

  • Full pipeline
    • Single-molecule generation (IUPAC/SMILES + metadata → one description).
    • Difficulty-based routing (easy/medium/hard) and model/reasoning settings.
    • Batch workflow: prepare prompts, call batch API, optional validation and export.
  • OPSIN: README and usage for feiyang-cai/opsin_mollangdata; this repo will link to it and document JAR usage in the pipeline.

(back to top)


Citation

If you use MolLangData in your research, please cite:

@article{MolLangData,
  title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
  author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  year={2026},
  journal={arXiv preprint arXiv:2602.02320},
}

For the related benchmark MolLangBench (ICLR 2026), you may also cite our previous work:

@inproceedings{MolLangBench,
  title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
  author={Cai, Feiyang and Bai, Jiahui and Tang, Tao and He, Guijuan and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
}

(back to top)


License

This project is licensed under the MIT License.

(back to top)


Contact

Maintainer: Feiyang Cai — feiyang@clemson.edu

(back to top)

Contributors

feiyang-cai

17 commits

TheLuoFengLab/MolLangData

A tool for genetating molecular structure-language description and the resulting large-scale dataset.

Python

2

17 commits

updated Sep 25, 2026

See the code

README

MolLangData: A Large-Scale Dataset for Molecular Structure–Language Description

GitHub stars GitHub forks GitHub issues Discord arXiv Hugging Face License: MIT

MolLangData is a large-scale dataset of molecular structures paired with natural-language descriptions, generated via a rule-regularized method. It supports training and evaluating models for molecular structure–language alignment.


Table of contents


Project Resources

ResourceLinksDescription
MolLangDataGitHub · Hugging FaceMain dataset on Hugging Face (~163k samples). We are actively expanding beyond this release.
MolLangBench (ICLR 2026)GitHub · Hugging FaceHuman-curated benchmark for molecular structure recognition, editing, and generation. The generation task aligns with structural description in this work and serves as a standard, validated evaluation.
LangMolDiodeGitHubLanguage-conditional molecule generator trained with SFT and reinforcement learning on MolLangData.

(back to top)


OPSIN (IUPAC → XML / SMILES)

We use a customized OPSIN fork that adds complete XML structure metadata for building prompts:

  • Fork: feiyang-cai/opsin_mollangdata
  • The repository is for reference; a compiled JAR is provided for the single-molecule workflow below.
  • The OPSIN fork README and usage will be updated later (see TODO).

(back to top)


Requirements

  • Python 3.8+
  • Java (JRE or JDK on PATH)
  • Python dependencies: pip install -r requirements.txt (installs openai, rdkit, and related packages)

(back to top)


Single molecule: get prompt and description from IUPAC

The script get_prompt_description_from_iupac.py takes one IUPAC name, runs OPSIN to obtain XML and SMILES, computes a difficulty level (easy / medium / hard), builds the prompt, and optionally calls an LLM to generate a structure description.

Files involved

PathPurpose
get_prompt_description_from_iupac.pyMain script
jar/Place the OPSIN JAR here: jar/opsin-core-*-jar-with-dependencies.jar
prompts/smiles_iupac_metadata_v13/Prompt template and semantic sections (default, bridged/fused/spiro rings)
config/llm_config.jsonLLM backend (Azure/OpenAI), model, reasoning effort per difficulty, timeouts (used only with --get-description)

Run (from repo root)

  • Prompt only (no LLM):

    python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone"
    
  • With LLM description (model and reasoning chosen by difficulty; backend from config):

    python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone" --get-description
    
  • Write output to a folder (prompt.md, descriptions.txt, generation_info.json):

    python3 get_prompt_description_from_iupac.py "ethane" --get-description -o out/
    
Options
OptionDescription
--opsin-jar PATHPath to OPSIN JAR (default: jar/opsin-core-*-jar-with-dependencies.jar)
--prompts-dir DIRPrompt templates directory (default: prompts/smiles_iupac_metadata_v13)
-o DIR / --output DIROutput folder: writes prompt.md; with --get-description also descriptions.txt and generation_info.json
--get-descriptionCall LLM to generate structure description (uses config/llm_config.json unless overridden)
--llm-config FILELLM config JSON (default: config/llm_config.json)
--model MODELOverride LLM model
--reasoning-effort EFFORTOverride LLM reasoning effort
Output
  • Without -o: Prints IUPAC, OPSIN status, SMILES, difficulty level, and (if --get-description) the description and prompt to stdout.
  • With -o DIR: Writes DIR/prompt.md; with --get-description also writes DIR/descriptions.txt and DIR/generation_info.json (backend, model, reasoning effort, duration, etc.).

(back to top)


Dataset generation pipeline

This section describes how to obtain or regenerate the MolLangData dataset. Pre-sampled and pre-processed data are available on Box (recommended); you can also run the full pipeline from raw PubChem data. Note: LLM outputs are non-deterministic—re-running the pipeline will not reproduce the same descriptions.

Download from Box to start from Step 4 (or Step 3 if you start from the sampled TSV only). Step 3 parsing output and Response API JSONL job files are also on Box, so you can skip to Step 4 or Step 5 as needed.

Box resourceContents
MolLangData PubChem sampled TSV• Sampled TSV: 8 rounds, 200k samples per round (round_0/sampled.tsv, round_1/sampled.tsv, …).
• Parsing output (Step 3): e.g. round_1/parsing_out_mollangdata_0.1.3.
• Response API JSONL (Step 4): e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs for starting at Step 5.

The published MolLangData dataset on Hugging Face corresponds to round_0 data.

Pipeline overview

StepDescription
1SDF → TSV — convert PubChem SDF to TSV (batch_prompt_generation/miscellaneous/1_sdf_to_tsv)
2Sampling — deterministic sampling from TSV into chunks (e.g. 200k per round) (batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm)
3OPSIN + MolLangData parsing — full XML structure data per sampled.tsv
4Create batch prompt JSONL — build LLM job files from parsing output
5Run LLM jobs — submit to obtain structure descriptions (OpenAI Batch or one-by-one)

Steps 1 and 2 can be time-consuming; we strongly recommend using the pre-sampled Box data above. Scripts: batch_prompt_generation/miscellaneous/1_sdf_to_tsv and batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm.


Step 3 — OPSIN + MolLangData parsing

Run the custom OPSIN tool on each sampled chunk to obtain full XML structure data. Parsing output for all rounds is on Box (see above); you can skip this step when using that data.

  • Input: A folder containing sampled.tsv (e.g. a round folder from Box).
  • Output: Under that folder (or --out-dir) as parsing_out_mollangdata_<version>/, containing opsin_xml/, mollangdata_xml/, and parsing_results.tsv.

The script auto-detects the OPSIN JAR in the repo (opsin/opsin-core/target/ or jar/); override with --opsin-jar.

Run (from repo root; <sample_folder> is the folder containing sampled.tsv):

python3 batch_prompt_generation/3_run_opsin_mollangdata_on_sampled.py <sample_folder>
Step 3 — Arguments
ArgumentDescription
sample_folderFolder containing sampled.tsv (required). TSV must have PUBCHEM_COMPOUND_CID, SMILES, canonical_smiles, and an IUPAC column (see --iupac-column).
--opsin-jar PATHPath to OPSIN MolLangData JAR. Default: auto-detect under repo opsin/opsin-core/target/ or jar/.
--out-dir DIROutput directory. Default: <sample_folder>/parsing_out; output is under parsing_out_mollangdata_<version>/.
--iupac-column NAMEIUPAC column name. Default: PUBCHEM_IUPAC_SYSTEMATIC_NAME.
--timeout-s NTimeout in seconds per molecule (default: 30).
--max-rows NProcess only first N rows (for testing).
--verboseVerbose logging.

Step 4 — Create batch prompt JSONL

Build LLM job files from Step 3 output. The script assigns difficulty (easy/medium/hard), builds the prompt with the same dynamic template as the single-molecule script (including fused/spiro/bridged sections when applicable), and routes model and reasoning effort from config/llm_config.json. Output is JSONL (and optional per-prompt TXT) for one-by-one or batch LLM runs. Supports Azure and OpenAI, and both Chat Completions and Responses APIs. Pre-built prompts for all rounds are on Box (e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs).

Step 4 — Backend / API recommendations
  • Azure: GPT-5.2 does not support batch on Azure as of 2/13/2026. Use one-by-one with Responses API (--api-format responses). This is what we used for data generation.
  • OpenAI: Both Responses and Chat Completions work. The Batch API with the Responses API has been reported to cause tasks to run repeatedly and incur extra cost (Batch API task runs repeatedly with gpt-5.2-pro). We recommend Chat Completions (--api-format chat_completions) when using OpenAI.

Run (from repo root). <input_folder> must be Step 3 output containing parsing_results.tsv (e.g. parsing_out_mollangdata_0.1.3). Example: on Box.

python3 batch_prompt_generation/4_create_batch_prompt_jsonl.py <input_folder> <output_folder> [prompts_folder] --api-format responses
Step 4 — Arguments
ArgumentDescription
input_folderFolder containing parsing_results.tsv (and mollangdata_xml/ or opsin_xml/). Typically Step 3 output (e.g. parsing_out_mollangdata_0.1.3).
output_folderDirectory for JSONL output (and optional TXT). Files are split by model and reasoning effort.
prompts_folderOptional; default: prompts/smiles_iupac_metadata_v13.
--llm-config FILELLM config (default: config/llm_config.json). Must have easy, medium, hard with model and reasoning_effort.
--api-formatresponses or chat_completions. Default: chat_completions. Use responses for Responses API.
--chat-completions-url URLURL for chat-completions (default: /v1/chat/completions).
--responses-url URLURL for responses (default: /v1/responses).
--xml-typemollangdata_xml or opsin_xml (default: mollangdata_xml).
--exclude-iupacOmit IUPAC Name: block from prompts.
--exclude-xmlOmit XML Metadata: block from prompts.
--allow-dotsInclude rows whose SMILES contain a dot (disconnected component).
--sample-size NRandomly sample N rows (for testing).
--random-seed NSeed for sampling (default: 533).
--custom-prefix STRPrefix for compound IDs in custom_id.
--no-txtDo not write per-prompt .txt; JSONL only.

Step 5 — Run LLM jobs (structure descriptions)

⚠️ Cost warning: Step 5 calls the LLM API for every prompt and can incur large costs (e.g. hundreds of thousands of requests). Check usage and billing before running. We recommend using the published MolLangData dataset unless you need to regenerate or extend descriptions.

Two options:

OptionScriptUse case
5a — OpenAI Batch5a_submit_openai_batch_jobs.pyUpload JSONL to OpenAI Batch API, wait, and retrieve. OpenAI only. Batch pricing is half of on-demand. Requires OPENAI_API_KEY.
5b — One-by-one5b_run_requests_one_by_one.py, 5b_run_tmux_jobs.shSequential or parallel (e.g. via tmux). Works with Azure and OpenAI. We use 5b because batch is not supported for GPT-5.2 on Azure (as of 2/13/2026). Responses API is recommended for background mode (submit → poll).
Step 5 — Recommendations
  • OpenAI: Prefer 5a (Batch) — half the price. Use 5b for testing or when batch is unavailable.
  • Azure: Use 5b (one-by-one); batch is not supported for GPT-5.2 on Azure.
  • API format: For 5b, prefer Responses API (background mode is more stable). Generate job JSONL with --api-format responses in Step 4 when using 5b.

Run (from repo root)

  • 5a — OpenAI Batch (submit, wait, retrieve). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.

    python3 batch_prompt_generation/5a_submit_openai_batch_jobs.py <input> --output-dir ./batch_results
    
  • 5b — One-by-one, single process (Azure, resume with pool). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.

    python3 batch_prompt_generation/5b_run_requests_one_by_one.py <input> --backend azure --output-dir ./runs/run1 --resume --pool-folder ./runs/pool
    
  • 5b — One-by-one, multiple processes (e.g. tmux; split one JSONL into 4 workers). Replace <input.jsonl> with a Step 4 .jsonl file:

    bash ./batch_prompt_generation/5b_run_tmux_jobs.sh -f <input.jsonl> -s myrun -n 4 -o ./runs_out
    
Step 5a — Arguments (OpenAI Batch)
ArgumentDescription
inputSingle .jsonl file or directory of .jsonl (from Step 4).
--output-dir DIRWhere to save retrieved output/error files (default: batch_results).
--no-waitSubmit only; do not wait or download (default: wait and retrieve).
--poll-interval NSeconds between status polls (default: 60).
--max-requests NMax requests per batch file when splitting (default: 50000).
--completion-window DURBatch completion window (default: 24h).
--dry-runList files and endpoints only; no upload or batch creation.
Step 5b — Arguments

One-by-one runner (5b_run_requests_one_by_one.py)

ArgumentDescription
input(Required.) A single .jsonl file or a directory of .jsonl files from Step 4. Each line is one LLM request.
--output-dir DIR(Required.) Directory for outputs: results_<backend>.jsonl, stats_<backend>.jsonl, and request_outputs/<custom_id>/ per request.
--backendazure or openai. Selects which LLM API to call.
--llm-config FILEPath to LLM config JSON (default: config/llm_config.json). Used for endpoint and auth.
--resumeIf set, skip requests that already have an entry in --output-dir and continue from the rest. Use after an interrupted run.
--pool-folder DIRWhen using Responses API background mode: folder to store in-flight request IDs for polling. Required for --resume to work correctly.
--model MODELOverride the model from config (e.g. gpt-5.2).
--reasoning-effort LEVELOverride reasoning effort from config (e.g. high, xhigh).
--max-requests NProcess at most N requests (for testing). Omit to process all.
--timeout NRequest timeout in seconds.
--poll-interval NSeconds between polls when using Responses API background mode.
--max-retries N, --retry-sleep SECRetry failed requests up to N times, waiting SEC seconds between retries.
--sleep SECOptional delay in seconds between requests (rate limiting).
--no-tqdmDisable progress bar.

Tmux multi-process (5b_run_tmux_jobs.sh)

Splits one JSONL into N chunks and runs N workers in tmux windows, each calling 5b_run_requests_one_by_one.py on one chunk. For Azure, each worker can use a different deployment (e.g. gpt-5.2, gpt-5.2-2) via -i START_IDX.

ArgumentDescription
-f FILE(Required.) Input JSONL file from Step 4.
-s SESSION_NAME(Required.) Tmux session name for the worker windows.
-n N(Required.) Number of splits = number of parallel workers (e.g. 4 → 4 tmux windows).
-o OUTPUT_BASEBase output directory; each worker gets a subdir (e.g. ./runs_out).
-p POOL_FOLDERPool folder for Responses API background mode (passed to each worker).
-i START_IDX(Azure.) Deployment start index so workers use different deployments (e.g. 0 → gpt-5.2, 1 → gpt-5.2-2).
--llm-config, --backend, --model, --reasoning-effortPassed through to each worker.
-t TIMEOUTRequest timeout in seconds.
--no-tqdmDisable progress bar in workers.

(back to top)


Dataset on Hugging Face: structure and validation

The MolLangData dataset on Hugging Face is provided in two configurations with the following structure and validation results.

Dataset structure

  1. validated_data

    • All validated data (2k samples), including descriptions that passed validation and those that did not.
    • See validation precision in the Dataset statistics table below.
  2. generated_data

    • All generated data from round 0, excluding the validated subset.
    • These samples are not validated.

Dataset statistics

DifficultyModelReasoning effortGenerated samplesValidated samplesValidation precision
EasyGPT-5.2high105,085 (65.2%)1,317 (65.8%)1,300 (98.7%)
MediumGPT-5.2xhigh40,916 (25.4%)496 (24.8%)492 (99.2%)
HardGPT-5.2xhigh15,110 (9.4%)187 (9.4%)180 (96.3%)
Overall——161,1112,0001,972 (98.6%)

(back to top)


TODO & roadmap

  • Full pipeline
    • Single-molecule generation (IUPAC/SMILES + metadata → one description).
    • Difficulty-based routing (easy/medium/hard) and model/reasoning settings.
    • Batch workflow: prepare prompts, call batch API, optional validation and export.
  • OPSIN: README and usage for feiyang-cai/opsin_mollangdata; this repo will link to it and document JAR usage in the pipeline.

(back to top)


Citation

If you use MolLangData in your research, please cite:

@article{MolLangData,
  title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
  author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  year={2026},
  journal={arXiv preprint arXiv:2602.02320},
}

For the related benchmark MolLangBench (ICLR 2026), you may also cite our previous work:

@inproceedings{MolLangBench,
  title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
  author={Cai, Feiyang and Bai, Jiahui and Tang, Tao and He, Guijuan and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
}

(back to top)


License

This project is licensed under the MIT License.

(back to top)


Contact

Maintainer: Feiyang Cai — feiyang@clemson.edu

(back to top)

Contributors

feiyang-cai

17 commits

Languages

Python

91.3%

Shell

8.7%