A tool for genetating molecular structure-language description and the resulting large-scale dataset.
Python
2
17 commits
updated Sep 25, 2026
MolLangData is a large-scale dataset of molecular structures paired with natural-language descriptions, generated via a rule-regularized method. It supports training and evaluating models for molecular structure–language alignment.
| Resource | Links | Description |
|---|---|---|
| MolLangData | GitHub · Hugging Face | Main dataset on Hugging Face (~163k samples). We are actively expanding beyond this release. |
| MolLangBench (ICLR 2026) | GitHub · Hugging Face | Human-curated benchmark for molecular structure recognition, editing, and generation. The generation task aligns with structural description in this work and serves as a standard, validated evaluation. |
| LangMolDiode | GitHub | Language-conditional molecule generator trained with SFT and reinforcement learning on MolLangData. |
We use a customized OPSIN fork that adds complete XML structure metadata for building prompts:
PATH)pip install -r requirements.txt (installs openai, rdkit, and related packages)The script get_prompt_description_from_iupac.py takes one IUPAC name, runs OPSIN to obtain XML and SMILES, computes a difficulty level (easy / medium / hard), builds the prompt, and optionally calls an LLM to generate a structure description.
Files involved
| Path | Purpose |
|---|---|
get_prompt_description_from_iupac.py | Main script |
jar/ | Place the OPSIN JAR here: jar/opsin-core-*-jar-with-dependencies.jar |
prompts/smiles_iupac_metadata_v13/ | Prompt template and semantic sections (default, bridged/fused/spiro rings) |
config/llm_config.json | LLM backend (Azure/OpenAI), model, reasoning effort per difficulty, timeouts (used only with --get-description) |
Run (from repo root)
Prompt only (no LLM):
python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone"
With LLM description (model and reasoning chosen by difficulty; backend from config):
python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone" --get-description
Write output to a folder (prompt.md, descriptions.txt, generation_info.json):
python3 get_prompt_description_from_iupac.py "ethane" --get-description -o out/
| Option | Description |
|---|---|
--opsin-jar PATH | Path to OPSIN JAR (default: jar/opsin-core-*-jar-with-dependencies.jar) |
--prompts-dir DIR | Prompt templates directory (default: prompts/smiles_iupac_metadata_v13) |
-o DIR / --output DIR | Output folder: writes prompt.md; with --get-description also descriptions.txt and generation_info.json |
--get-description | Call LLM to generate structure description (uses config/llm_config.json unless overridden) |
--llm-config FILE | LLM config JSON (default: config/llm_config.json) |
--model MODEL | Override LLM model |
--reasoning-effort EFFORT | Override LLM reasoning effort |
-o: Prints IUPAC, OPSIN status, SMILES, difficulty level, and (if --get-description) the description and prompt to stdout.-o DIR: Writes DIR/prompt.md; with --get-description also writes DIR/descriptions.txt and DIR/generation_info.json (backend, model, reasoning effort, duration, etc.).This section describes how to obtain or regenerate the MolLangData dataset. Pre-sampled and pre-processed data are available on Box (recommended); you can also run the full pipeline from raw PubChem data. Note: LLM outputs are non-deterministic—re-running the pipeline will not reproduce the same descriptions.
Download from Box to start from Step 4 (or Step 3 if you start from the sampled TSV only). Step 3 parsing output and Response API JSONL job files are also on Box, so you can skip to Step 4 or Step 5 as needed.
| Box resource | Contents |
|---|---|
| MolLangData PubChem sampled TSV | • Sampled TSV: 8 rounds, 200k samples per round (round_0/sampled.tsv, round_1/sampled.tsv, …).• Parsing output (Step 3): e.g. round_1/parsing_out_mollangdata_0.1.3.• Response API JSONL (Step 4): e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs for starting at Step 5. |
The published MolLangData dataset on Hugging Face corresponds to round_0 data.
| Step | Description |
|---|---|
| 1 | SDF → TSV — convert PubChem SDF to TSV (batch_prompt_generation/miscellaneous/1_sdf_to_tsv) |
| 2 | Sampling — deterministic sampling from TSV into chunks (e.g. 200k per round) (batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm) |
| 3 | OPSIN + MolLangData parsing — full XML structure data per sampled.tsv |
| 4 | Create batch prompt JSONL — build LLM job files from parsing output |
| 5 | Run LLM jobs — submit to obtain structure descriptions (OpenAI Batch or one-by-one) |
Steps 1 and 2 can be time-consuming; we strongly recommend using the pre-sampled Box data above. Scripts: batch_prompt_generation/miscellaneous/1_sdf_to_tsv and batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm.
Run the custom OPSIN tool on each sampled chunk to obtain full XML structure data. Parsing output for all rounds is on Box (see above); you can skip this step when using that data.
sampled.tsv (e.g. a round folder from Box).--out-dir) as parsing_out_mollangdata_<version>/, containing opsin_xml/, mollangdata_xml/, and parsing_results.tsv.The script auto-detects the OPSIN JAR in the repo (opsin/opsin-core/target/ or jar/); override with --opsin-jar.
Run (from repo root; <sample_folder> is the folder containing sampled.tsv):
python3 batch_prompt_generation/3_run_opsin_mollangdata_on_sampled.py <sample_folder>
| Argument | Description |
|---|---|
sample_folder | Folder containing sampled.tsv (required). TSV must have PUBCHEM_COMPOUND_CID, SMILES, canonical_smiles, and an IUPAC column (see --iupac-column). |
--opsin-jar PATH | Path to OPSIN MolLangData JAR. Default: auto-detect under repo opsin/opsin-core/target/ or jar/. |
--out-dir DIR | Output directory. Default: <sample_folder>/parsing_out; output is under parsing_out_mollangdata_<version>/. |
--iupac-column NAME | IUPAC column name. Default: PUBCHEM_IUPAC_SYSTEMATIC_NAME. |
--timeout-s N | Timeout in seconds per molecule (default: 30). |
--max-rows N | Process only first N rows (for testing). |
--verbose | Verbose logging. |
Build LLM job files from Step 3 output. The script assigns difficulty (easy/medium/hard), builds the prompt with the same dynamic template as the single-molecule script (including fused/spiro/bridged sections when applicable), and routes model and reasoning effort from config/llm_config.json. Output is JSONL (and optional per-prompt TXT) for one-by-one or batch LLM runs. Supports Azure and OpenAI, and both Chat Completions and Responses APIs. Pre-built prompts for all rounds are on Box (e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs).
--api-format responses). This is what we used for data generation.--api-format chat_completions) when using OpenAI.Run (from repo root). <input_folder> must be Step 3 output containing parsing_results.tsv (e.g. parsing_out_mollangdata_0.1.3). Example: on Box.
python3 batch_prompt_generation/4_create_batch_prompt_jsonl.py <input_folder> <output_folder> [prompts_folder] --api-format responses
| Argument | Description |
|---|---|
input_folder | Folder containing parsing_results.tsv (and mollangdata_xml/ or opsin_xml/). Typically Step 3 output (e.g. parsing_out_mollangdata_0.1.3). |
output_folder | Directory for JSONL output (and optional TXT). Files are split by model and reasoning effort. |
prompts_folder | Optional; default: prompts/smiles_iupac_metadata_v13. |
--llm-config FILE | LLM config (default: config/llm_config.json). Must have easy, medium, hard with model and reasoning_effort. |
--api-format | responses or chat_completions. Default: chat_completions. Use responses for Responses API. |
--chat-completions-url URL | URL for chat-completions (default: /v1/chat/completions). |
--responses-url URL | URL for responses (default: /v1/responses). |
--xml-type | mollangdata_xml or opsin_xml (default: mollangdata_xml). |
--exclude-iupac | Omit IUPAC Name: block from prompts. |
--exclude-xml | Omit XML Metadata: block from prompts. |
--allow-dots | Include rows whose SMILES contain a dot (disconnected component). |
--sample-size N | Randomly sample N rows (for testing). |
--random-seed N | Seed for sampling (default: 533). |
--custom-prefix STR | Prefix for compound IDs in custom_id. |
--no-txt | Do not write per-prompt .txt; JSONL only. |
⚠️ Cost warning: Step 5 calls the LLM API for every prompt and can incur large costs (e.g. hundreds of thousands of requests). Check usage and billing before running. We recommend using the published MolLangData dataset unless you need to regenerate or extend descriptions.
Two options:
| Option | Script | Use case |
|---|---|---|
| 5a — OpenAI Batch | 5a_submit_openai_batch_jobs.py | Upload JSONL to OpenAI Batch API, wait, and retrieve. OpenAI only. Batch pricing is half of on-demand. Requires OPENAI_API_KEY. |
| 5b — One-by-one | 5b_run_requests_one_by_one.py, 5b_run_tmux_jobs.sh | Sequential or parallel (e.g. via tmux). Works with Azure and OpenAI. We use 5b because batch is not supported for GPT-5.2 on Azure (as of 2/13/2026). Responses API is recommended for background mode (submit → poll). |
--api-format responses in Step 4 when using 5b.Run (from repo root)
5a — OpenAI Batch (submit, wait, retrieve). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.
python3 batch_prompt_generation/5a_submit_openai_batch_jobs.py <input> --output-dir ./batch_results
5b — One-by-one, single process (Azure, resume with pool). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.
python3 batch_prompt_generation/5b_run_requests_one_by_one.py <input> --backend azure --output-dir ./runs/run1 --resume --pool-folder ./runs/pool
5b — One-by-one, multiple processes (e.g. tmux; split one JSONL into 4 workers). Replace <input.jsonl> with a Step 4 .jsonl file:
bash ./batch_prompt_generation/5b_run_tmux_jobs.sh -f <input.jsonl> -s myrun -n 4 -o ./runs_out
| Argument | Description |
|---|---|
input | Single .jsonl file or directory of .jsonl (from Step 4). |
--output-dir DIR | Where to save retrieved output/error files (default: batch_results). |
--no-wait | Submit only; do not wait or download (default: wait and retrieve). |
--poll-interval N | Seconds between status polls (default: 60). |
--max-requests N | Max requests per batch file when splitting (default: 50000). |
--completion-window DUR | Batch completion window (default: 24h). |
--dry-run | List files and endpoints only; no upload or batch creation. |
One-by-one runner (5b_run_requests_one_by_one.py)
| Argument | Description |
|---|---|
input | (Required.) A single .jsonl file or a directory of .jsonl files from Step 4. Each line is one LLM request. |
--output-dir DIR | (Required.) Directory for outputs: results_<backend>.jsonl, stats_<backend>.jsonl, and request_outputs/<custom_id>/ per request. |
--backend | azure or openai. Selects which LLM API to call. |
--llm-config FILE | Path to LLM config JSON (default: config/llm_config.json). Used for endpoint and auth. |
--resume | If set, skip requests that already have an entry in --output-dir and continue from the rest. Use after an interrupted run. |
--pool-folder DIR | When using Responses API background mode: folder to store in-flight request IDs for polling. Required for --resume to work correctly. |
--model MODEL | Override the model from config (e.g. gpt-5.2). |
--reasoning-effort LEVEL | Override reasoning effort from config (e.g. high, xhigh). |
--max-requests N | Process at most N requests (for testing). Omit to process all. |
--timeout N | Request timeout in seconds. |
--poll-interval N | Seconds between polls when using Responses API background mode. |
--max-retries N, --retry-sleep SEC | Retry failed requests up to N times, waiting SEC seconds between retries. |
--sleep SEC | Optional delay in seconds between requests (rate limiting). |
--no-tqdm | Disable progress bar. |
Tmux multi-process (5b_run_tmux_jobs.sh)
Splits one JSONL into N chunks and runs N workers in tmux windows, each calling 5b_run_requests_one_by_one.py on one chunk. For Azure, each worker can use a different deployment (e.g. gpt-5.2, gpt-5.2-2) via -i START_IDX.
| Argument | Description |
|---|---|
-f FILE | (Required.) Input JSONL file from Step 4. |
-s SESSION_NAME | (Required.) Tmux session name for the worker windows. |
-n N | (Required.) Number of splits = number of parallel workers (e.g. 4 → 4 tmux windows). |
-o OUTPUT_BASE | Base output directory; each worker gets a subdir (e.g. ./runs_out). |
-p POOL_FOLDER | Pool folder for Responses API background mode (passed to each worker). |
-i START_IDX | (Azure.) Deployment start index so workers use different deployments (e.g. 0 → gpt-5.2, 1 → gpt-5.2-2). |
--llm-config, --backend, --model, --reasoning-effort | Passed through to each worker. |
-t TIMEOUT | Request timeout in seconds. |
--no-tqdm | Disable progress bar in workers. |
The MolLangData dataset on Hugging Face is provided in two configurations with the following structure and validation results.
validated_data
generated_data
| Difficulty | Model | Reasoning effort | Generated samples | Validated samples | Validation precision |
|---|---|---|---|---|---|
| Easy | GPT-5.2 | high | 105,085 (65.2%) | 1,317 (65.8%) | 1,300 (98.7%) |
| Medium | GPT-5.2 | xhigh | 40,916 (25.4%) | 496 (24.8%) | 492 (99.2%) |
| Hard | GPT-5.2 | xhigh | 15,110 (9.4%) | 187 (9.4%) | 180 (96.3%) |
| Overall | — | — | 161,111 | 2,000 | 1,972 (98.6%) |
If you use MolLangData in your research, please cite:
@article{MolLangData,
title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
year={2026},
journal={arXiv preprint arXiv:2602.02320},
}
For the related benchmark MolLangBench (ICLR 2026), you may also cite our previous work:
@inproceedings{MolLangBench,
title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
author={Cai, Feiyang and Bai, Jiahui and Tang, Tao and He, Guijuan and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
}
This project is licensed under the MIT License.
Maintainer: Feiyang Cai — feiyang@clemson.edu
17 commits
Python
91.3%
Shell
8.7%
A tool for genetating molecular structure-language description and the resulting large-scale dataset.
Python
2
17 commits
updated Sep 25, 2026
MolLangData is a large-scale dataset of molecular structures paired with natural-language descriptions, generated via a rule-regularized method. It supports training and evaluating models for molecular structure–language alignment.
| Resource | Links | Description |
|---|---|---|
| MolLangData | GitHub · Hugging Face | Main dataset on Hugging Face (~163k samples). We are actively expanding beyond this release. |
| MolLangBench (ICLR 2026) | GitHub · Hugging Face | Human-curated benchmark for molecular structure recognition, editing, and generation. The generation task aligns with structural description in this work and serves as a standard, validated evaluation. |
| LangMolDiode | GitHub | Language-conditional molecule generator trained with SFT and reinforcement learning on MolLangData. |
We use a customized OPSIN fork that adds complete XML structure metadata for building prompts:
PATH)pip install -r requirements.txt (installs openai, rdkit, and related packages)The script get_prompt_description_from_iupac.py takes one IUPAC name, runs OPSIN to obtain XML and SMILES, computes a difficulty level (easy / medium / hard), builds the prompt, and optionally calls an LLM to generate a structure description.
Files involved
| Path | Purpose |
|---|---|
get_prompt_description_from_iupac.py | Main script |
jar/ | Place the OPSIN JAR here: jar/opsin-core-*-jar-with-dependencies.jar |
prompts/smiles_iupac_metadata_v13/ | Prompt template and semantic sections (default, bridged/fused/spiro rings) |
config/llm_config.json | LLM backend (Azure/OpenAI), model, reasoning effort per difficulty, timeouts (used only with --get-description) |
Run (from repo root)
Prompt only (no LLM):
python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone"
With LLM description (model and reasoning chosen by difficulty; backend from config):
python3 get_prompt_description_from_iupac.py "3,4-dihydro-2H-1,5-benzodioxepin-7-yl-(2-fluorophenyl)methanone" --get-description
Write output to a folder (prompt.md, descriptions.txt, generation_info.json):
python3 get_prompt_description_from_iupac.py "ethane" --get-description -o out/
| Option | Description |
|---|---|
--opsin-jar PATH | Path to OPSIN JAR (default: jar/opsin-core-*-jar-with-dependencies.jar) |
--prompts-dir DIR | Prompt templates directory (default: prompts/smiles_iupac_metadata_v13) |
-o DIR / --output DIR | Output folder: writes prompt.md; with --get-description also descriptions.txt and generation_info.json |
--get-description | Call LLM to generate structure description (uses config/llm_config.json unless overridden) |
--llm-config FILE | LLM config JSON (default: config/llm_config.json) |
--model MODEL | Override LLM model |
--reasoning-effort EFFORT | Override LLM reasoning effort |
-o: Prints IUPAC, OPSIN status, SMILES, difficulty level, and (if --get-description) the description and prompt to stdout.-o DIR: Writes DIR/prompt.md; with --get-description also writes DIR/descriptions.txt and DIR/generation_info.json (backend, model, reasoning effort, duration, etc.).This section describes how to obtain or regenerate the MolLangData dataset. Pre-sampled and pre-processed data are available on Box (recommended); you can also run the full pipeline from raw PubChem data. Note: LLM outputs are non-deterministic—re-running the pipeline will not reproduce the same descriptions.
Download from Box to start from Step 4 (or Step 3 if you start from the sampled TSV only). Step 3 parsing output and Response API JSONL job files are also on Box, so you can skip to Step 4 or Step 5 as needed.
| Box resource | Contents |
|---|---|
| MolLangData PubChem sampled TSV | • Sampled TSV: 8 rounds, 200k samples per round (round_0/sampled.tsv, round_1/sampled.tsv, …).• Parsing output (Step 3): e.g. round_1/parsing_out_mollangdata_0.1.3.• Response API JSONL (Step 4): e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs for starting at Step 5. |
The published MolLangData dataset on Hugging Face corresponds to round_0 data.
| Step | Description |
|---|---|
| 1 | SDF → TSV — convert PubChem SDF to TSV (batch_prompt_generation/miscellaneous/1_sdf_to_tsv) |
| 2 | Sampling — deterministic sampling from TSV into chunks (e.g. 200k per round) (batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm) |
| 3 | OPSIN + MolLangData parsing — full XML structure data per sampled.tsv |
| 4 | Create batch prompt JSONL — build LLM job files from parsing output |
| 5 | Run LLM jobs — submit to obtain structure descriptions (OpenAI Batch or one-by-one) |
Steps 1 and 2 can be time-consuming; we strongly recommend using the pre-sampled Box data above. Scripts: batch_prompt_generation/miscellaneous/1_sdf_to_tsv and batch_prompt_generation/miscellaneous/2_sampling_from_pubchecm.
Run the custom OPSIN tool on each sampled chunk to obtain full XML structure data. Parsing output for all rounds is on Box (see above); you can skip this step when using that data.
sampled.tsv (e.g. a round folder from Box).--out-dir) as parsing_out_mollangdata_<version>/, containing opsin_xml/, mollangdata_xml/, and parsing_results.tsv.The script auto-detects the OPSIN JAR in the repo (opsin/opsin-core/target/ or jar/); override with --opsin-jar.
Run (from repo root; <sample_folder> is the folder containing sampled.tsv):
python3 batch_prompt_generation/3_run_opsin_mollangdata_on_sampled.py <sample_folder>
| Argument | Description |
|---|---|
sample_folder | Folder containing sampled.tsv (required). TSV must have PUBCHEM_COMPOUND_CID, SMILES, canonical_smiles, and an IUPAC column (see --iupac-column). |
--opsin-jar PATH | Path to OPSIN MolLangData JAR. Default: auto-detect under repo opsin/opsin-core/target/ or jar/. |
--out-dir DIR | Output directory. Default: <sample_folder>/parsing_out; output is under parsing_out_mollangdata_<version>/. |
--iupac-column NAME | IUPAC column name. Default: PUBCHEM_IUPAC_SYSTEMATIC_NAME. |
--timeout-s N | Timeout in seconds per molecule (default: 30). |
--max-rows N | Process only first N rows (for testing). |
--verbose | Verbose logging. |
Build LLM job files from Step 3 output. The script assigns difficulty (easy/medium/hard), builds the prompt with the same dynamic template as the single-molecule script (including fused/spiro/bridged sections when applicable), and routes model and reasoning effort from config/llm_config.json. Output is JSONL (and optional per-prompt TXT) for one-by-one or batch LLM runs. Supports Azure and OpenAI, and both Chat Completions and Responses APIs. Pre-built prompts for all rounds are on Box (e.g. round_1/parsing_out_mollangdata_0.1.3_prompts_jobs).
--api-format responses). This is what we used for data generation.--api-format chat_completions) when using OpenAI.Run (from repo root). <input_folder> must be Step 3 output containing parsing_results.tsv (e.g. parsing_out_mollangdata_0.1.3). Example: on Box.
python3 batch_prompt_generation/4_create_batch_prompt_jsonl.py <input_folder> <output_folder> [prompts_folder] --api-format responses
| Argument | Description |
|---|---|
input_folder | Folder containing parsing_results.tsv (and mollangdata_xml/ or opsin_xml/). Typically Step 3 output (e.g. parsing_out_mollangdata_0.1.3). |
output_folder | Directory for JSONL output (and optional TXT). Files are split by model and reasoning effort. |
prompts_folder | Optional; default: prompts/smiles_iupac_metadata_v13. |
--llm-config FILE | LLM config (default: config/llm_config.json). Must have easy, medium, hard with model and reasoning_effort. |
--api-format | responses or chat_completions. Default: chat_completions. Use responses for Responses API. |
--chat-completions-url URL | URL for chat-completions (default: /v1/chat/completions). |
--responses-url URL | URL for responses (default: /v1/responses). |
--xml-type | mollangdata_xml or opsin_xml (default: mollangdata_xml). |
--exclude-iupac | Omit IUPAC Name: block from prompts. |
--exclude-xml | Omit XML Metadata: block from prompts. |
--allow-dots | Include rows whose SMILES contain a dot (disconnected component). |
--sample-size N | Randomly sample N rows (for testing). |
--random-seed N | Seed for sampling (default: 533). |
--custom-prefix STR | Prefix for compound IDs in custom_id. |
--no-txt | Do not write per-prompt .txt; JSONL only. |
⚠️ Cost warning: Step 5 calls the LLM API for every prompt and can incur large costs (e.g. hundreds of thousands of requests). Check usage and billing before running. We recommend using the published MolLangData dataset unless you need to regenerate or extend descriptions.
Two options:
| Option | Script | Use case |
|---|---|---|
| 5a — OpenAI Batch | 5a_submit_openai_batch_jobs.py | Upload JSONL to OpenAI Batch API, wait, and retrieve. OpenAI only. Batch pricing is half of on-demand. Requires OPENAI_API_KEY. |
| 5b — One-by-one | 5b_run_requests_one_by_one.py, 5b_run_tmux_jobs.sh | Sequential or parallel (e.g. via tmux). Works with Azure and OpenAI. We use 5b because batch is not supported for GPT-5.2 on Azure (as of 2/13/2026). Responses API is recommended for background mode (submit → poll). |
--api-format responses in Step 4 when using 5b.Run (from repo root)
5a — OpenAI Batch (submit, wait, retrieve). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.
python3 batch_prompt_generation/5a_submit_openai_batch_jobs.py <input> --output-dir ./batch_results
5b — One-by-one, single process (Azure, resume with pool). Replace <input> with a Step 4 .jsonl file or a directory of .jsonl files. Example JSONL: on Box.
python3 batch_prompt_generation/5b_run_requests_one_by_one.py <input> --backend azure --output-dir ./runs/run1 --resume --pool-folder ./runs/pool
5b — One-by-one, multiple processes (e.g. tmux; split one JSONL into 4 workers). Replace <input.jsonl> with a Step 4 .jsonl file:
bash ./batch_prompt_generation/5b_run_tmux_jobs.sh -f <input.jsonl> -s myrun -n 4 -o ./runs_out
| Argument | Description |
|---|---|
input | Single .jsonl file or directory of .jsonl (from Step 4). |
--output-dir DIR | Where to save retrieved output/error files (default: batch_results). |
--no-wait | Submit only; do not wait or download (default: wait and retrieve). |
--poll-interval N | Seconds between status polls (default: 60). |
--max-requests N | Max requests per batch file when splitting (default: 50000). |
--completion-window DUR | Batch completion window (default: 24h). |
--dry-run | List files and endpoints only; no upload or batch creation. |
One-by-one runner (5b_run_requests_one_by_one.py)
| Argument | Description |
|---|---|
input | (Required.) A single .jsonl file or a directory of .jsonl files from Step 4. Each line is one LLM request. |
--output-dir DIR | (Required.) Directory for outputs: results_<backend>.jsonl, stats_<backend>.jsonl, and request_outputs/<custom_id>/ per request. |
--backend | azure or openai. Selects which LLM API to call. |
--llm-config FILE | Path to LLM config JSON (default: config/llm_config.json). Used for endpoint and auth. |
--resume | If set, skip requests that already have an entry in --output-dir and continue from the rest. Use after an interrupted run. |
--pool-folder DIR | When using Responses API background mode: folder to store in-flight request IDs for polling. Required for --resume to work correctly. |
--model MODEL | Override the model from config (e.g. gpt-5.2). |
--reasoning-effort LEVEL | Override reasoning effort from config (e.g. high, xhigh). |
--max-requests N | Process at most N requests (for testing). Omit to process all. |
--timeout N | Request timeout in seconds. |
--poll-interval N | Seconds between polls when using Responses API background mode. |
--max-retries N, --retry-sleep SEC | Retry failed requests up to N times, waiting SEC seconds between retries. |
--sleep SEC | Optional delay in seconds between requests (rate limiting). |
--no-tqdm | Disable progress bar. |
Tmux multi-process (5b_run_tmux_jobs.sh)
Splits one JSONL into N chunks and runs N workers in tmux windows, each calling 5b_run_requests_one_by_one.py on one chunk. For Azure, each worker can use a different deployment (e.g. gpt-5.2, gpt-5.2-2) via -i START_IDX.
| Argument | Description |
|---|---|
-f FILE | (Required.) Input JSONL file from Step 4. |
-s SESSION_NAME | (Required.) Tmux session name for the worker windows. |
-n N | (Required.) Number of splits = number of parallel workers (e.g. 4 → 4 tmux windows). |
-o OUTPUT_BASE | Base output directory; each worker gets a subdir (e.g. ./runs_out). |
-p POOL_FOLDER | Pool folder for Responses API background mode (passed to each worker). |
-i START_IDX | (Azure.) Deployment start index so workers use different deployments (e.g. 0 → gpt-5.2, 1 → gpt-5.2-2). |
--llm-config, --backend, --model, --reasoning-effort | Passed through to each worker. |
-t TIMEOUT | Request timeout in seconds. |
--no-tqdm | Disable progress bar in workers. |
The MolLangData dataset on Hugging Face is provided in two configurations with the following structure and validation results.
validated_data
generated_data
| Difficulty | Model | Reasoning effort | Generated samples | Validated samples | Validation precision |
|---|---|---|---|---|---|
| Easy | GPT-5.2 | high | 105,085 (65.2%) | 1,317 (65.8%) | 1,300 (98.7%) |
| Medium | GPT-5.2 | xhigh | 40,916 (25.4%) | 496 (24.8%) | 492 (99.2%) |
| Hard | GPT-5.2 | xhigh | 15,110 (9.4%) | 187 (9.4%) | 180 (96.3%) |
| Overall | — | — | 161,111 | 2,000 | 1,972 (98.6%) |
If you use MolLangData in your research, please cite:
@article{MolLangData,
title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
year={2026},
journal={arXiv preprint arXiv:2602.02320},
}
For the related benchmark MolLangBench (ICLR 2026), you may also cite our previous work:
@inproceedings{MolLangBench,
title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
author={Cai, Feiyang and Bai, Jiahui and Tang, Tao and He, Guijuan and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
}
This project is licensed under the MIT License.
Maintainer: Feiyang Cai — feiyang@clemson.edu
17 commits
Python
91.3%
Shell
8.7%