NarrativeWeaver is a controllable multi-image generation framework that unifies narrative planning, visual generation, and identity-consistent synthesis in a single model. Given a reference image and a high-level instruction, NarrativeWeaver first auto-regressively plans per-image descriptions, then renders each image with both semantic coherence and visual consistency across the sequence. A fine-grained alignment stage further preserves subject identity across frames, and an optional Memory Bank extension enables long-sequence generation with stronger cross-image consistency.
This codebase builds upon UniWorld.
conda create -n narrative_weaver python=3.10
conda activate narrative_weaver
pip install -r requirements.txt
Requirements:
flex-attention support (used by the FLUX denoiser and the vision encoder)requirements.txt for the full dependency listNarrativeWeaver depends on the following pretrained models. Download them from HuggingFace before training or inference:
| Model | HuggingFace ID |
|---|---|
| Qwen2.5-VL-3B-Instruct | Qwen/Qwen2.5-VL-3B-Instruct |
| FLUX.1-dev | black-forest-labs/FLUX.1-dev |
| SigLIP (Stage 3 only) | google/siglip2-so400m-patch16-512 |
| SigLIP MLP weights (Stage 3 only) | flux-redux-siglipv2-512.bin (from the UniWorld repo) |
Note: We found that SigLIP can play a similar role to a VAE for consistency control while offering a higher compression ratio, so we adopt SigLIP in this project.
After downloading, merge the Qwen2.5-VL and FLUX weights into a single NarrativeWeaver initialization checkpoint:
python scripts/make_NarrativeWeaver_weight.py \
--origin_flux_ckpt_path /path/to/FLUX.1-dev \
--origin_qwenvl_ckpt_path /path/to/Qwen2.5-VL-3B-Instruct \
--save_path /path/to/output/NarrativeWeaver-init
This produces the NarrativeWeaver-init checkpoint that serves as the starting point for Stage 0 (T2I Pretrain).
.txt file)Each line of the training data txt file contains three comma-separated fields:
/path/to/images_dir,/path/to/annotations.json,false
| Field | Description |
|---|---|
/path/to/images_dir | Root directory containing the images referenced in the JSON |
/path/to/annotations.json | JSON file with annotation entries (see below) |
false | Default flag (keep as false) |
Each entry in the annotation JSON follows this structure:
{
"id": "sample_id",
"image": [
"sample_id/condition.jpg",
"sample_id/output_img_1.png",
"sample_id/output_img_2.png"
],
"conversations": [
{
"from": "human",
"value": "Your instruction prompt here. <image>"
},
{
"from": "gpt",
"value": "Description for image 1. <gen_image>Description for image 2. <gen_image>"
}
]
}
Key points:
image list is the condition / reference image (e.g., a product photo or a character reference).<image> in the human turn refers to the condition / reference image.<gen_image> in the gpt turn acts as a delimiter between descriptions of consecutive generated images. Each text segment before a <gen_image> token describes the corresponding output image.<gen_image> tokens should equal the number of target images (i.e., len(image) - 1).NarrativeWeaver follows a 4-stage training pipeline. Each stage loads from the checkpoint produced by the previous stage.
Before training, set your WandB API key and update every path in the YAML config files to point to your own data and checkpoints:
export WANDB_API_KEY="your_wandb_api_key"
Training uses DeepSpeed ZeRO via HuggingFace Accelerate. Multi-node configs are provided under scripts/accelerate_configs/.
Initializes the learnable MetaQuery and MLP projector by training on text-to-image data, with the LLM frozen.
Config: scripts/denoiser/T2I_pretrain/T2I_flux_qwen2p5vl_3b_vlm_pretrain.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/NarrativeWeaver-init
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
dataset_config:
data_txt: /path/to/data/data_t2i.txt
training_config:
output_dir: /path/to/output/T2I_pretrain
Run:
bash scripts/denoiser/T2I_pretrain/T2I_flux_qwen2p5vl_3b_vlm_pretrain.sh
Fine-tunes the LLM to perform narrative planning — producing per-image textual descriptions from a reference image and an instruction. The denoiser is not updated in this stage.
Config: scripts/denoiser/E-Commerce/Interleaved_e-commerce_step1.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/T2I_pretrain/checkpoint-32000/univa
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
dataset_config:
data_txt: /path/to/data/data_e-commerce_train.txt
validation_json_path: /path/to/data/val.json
validation_image_dir: /path/to/data/images
training_config:
output_dir: /path/to/output/step1_narrative_planning
Run:
bash scripts/denoiser/E-Commerce/Interleaved_e-commerce_step1.sh
Trains the FLUX denoiser end-to-end with the LLM frozen, enabling the model to generate images conditioned on the planned narrative descriptions.
Config: scripts/denoiser/E-Commerce/Interleaved_e-commerce_step2.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/step1/checkpoint/univa
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
dataset_config:
data_txt: /path/to/data/data_e-commerce_train.txt
validation_json_path: /path/to/data/val.json
validation_image_dir: /path/to/data/images
training_config:
output_dir: /path/to/output/step2_visual_generation
Run:
bash scripts/denoiser/E-Commerce/Interleaved_e-commerce_step2.sh
Introduces SigLIP-based identity alignment by injecting fine-grained visual features from the reference image into the denoiser, improving subject consistency across generated images.
Config: scripts/denoiser/E-Commerce/Interleaved_e-commerce_step3.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/step2/checkpoint/univa
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
pretrained_siglip_name_or_path: /path/to/siglip2-so400m-patch16-512
pretrained_siglip_mlp_path: /path/to/flux-redux-siglipv2-512.bin
dataset_config:
data_txt: /path/to/data/data_e-commerce_train.txt
validation_json_path: /path/to/data/val.json
validation_image_dir: /path/to/data/images
training_config:
output_dir: /path/to/output/step3_fine_grained_alignment
Run:
bash scripts/denoiser/E-Commerce/Interleaved_e-commerce_step3.sh
A Memory Bank extension of Stage 3 for long-sequence generation. It maintains a cross-image memory buffer that improves consistency across many frames.
Config: scripts/denoiser/MemoryBank/Interleaved_e-commerce_step3.yaml
Run:
bash scripts/denoiser/MemoryBank/Interleaved_e-commerce_step3.sh
All inference scripts live under univa/serve/. Output images are written to --save_dir.
Runs the complete NarrativeWeaver pipeline — the model first generates per-image descriptions from the instruction and reference image, then synthesizes all images.
python -m univa.serve.test_interleaved_flex_visualize \
--model_path /path/to/checkpoint/univa \
--flux_path /path/to/FLUX.1-dev \
--siglip_path /path/to/siglip2-so400m-patch16-512 \
--test_path /path/to/test.json \
--test_image_path /path/to/images \
--height 480 \
--width 832 \
--save_dir /path/to/output \
--stage 2
Skips narrative planning and generates images directly from existing per-image descriptions in the test JSON.
python -m univa.serve.test_interleaved_imageOnly_flex \
--model_path /path/to/checkpoint/univa \
--flux_path /path/to/FLUX.1-dev \
--siglip_path /path/to/siglip2-so400m-patch16-512 \
--test_path /path/to/test.json \
--test_image_path /path/to/images \
--height 480 \
--width 832 \
--save_dir /path/to/output \
--stage 2
Uses a Memory Bank checkpoint for long-sequence generation with enhanced cross-image consistency.
python -m univa.serve.test_interleaved_imageOnly_flex_memoryBank \
--model_path /path/to/memory_bank_checkpoint/univa \
--flux_path /path/to/FLUX.1-dev \
--siglip_path /path/to/siglip2-so400m-patch16-512 \
--test_path /path/to/test.json \
--test_image_path /path/to/images \
--height 480 \
--width 832 \
--save_dir /path/to/output \
--stage 2
If you find this work useful, please cite:
@article{yao2026narrative,
title = {Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning},
author = {Yao, Zhengjian and Li, Yongzhi and Gao, Xinyuan and Chen, Quan and Jiang, Peng and Lu, Yanye},
journal = {arXiv preprint arXiv:2603.06688},
year = {2026}
}
This codebase builds upon UniWorld. We thank the authors for their excellent work and for releasing their code.
1 commits
Python
99.5%
NarrativeWeaver is a controllable multi-image generation framework that unifies narrative planning, visual generation, and identity-consistent synthesis in a single model. Given a reference image and a high-level instruction, NarrativeWeaver first auto-regressively plans per-image descriptions, then renders each image with both semantic coherence and visual consistency across the sequence. A fine-grained alignment stage further preserves subject identity across frames, and an optional Memory Bank extension enables long-sequence generation with stronger cross-image consistency.
This codebase builds upon UniWorld.
conda create -n narrative_weaver python=3.10
conda activate narrative_weaver
pip install -r requirements.txt
Requirements:
flex-attention support (used by the FLUX denoiser and the vision encoder)requirements.txt for the full dependency listNarrativeWeaver depends on the following pretrained models. Download them from HuggingFace before training or inference:
| Model | HuggingFace ID |
|---|---|
| Qwen2.5-VL-3B-Instruct | Qwen/Qwen2.5-VL-3B-Instruct |
| FLUX.1-dev | black-forest-labs/FLUX.1-dev |
| SigLIP (Stage 3 only) | google/siglip2-so400m-patch16-512 |
| SigLIP MLP weights (Stage 3 only) | flux-redux-siglipv2-512.bin (from the UniWorld repo) |
Note: We found that SigLIP can play a similar role to a VAE for consistency control while offering a higher compression ratio, so we adopt SigLIP in this project.
After downloading, merge the Qwen2.5-VL and FLUX weights into a single NarrativeWeaver initialization checkpoint:
python scripts/make_NarrativeWeaver_weight.py \
--origin_flux_ckpt_path /path/to/FLUX.1-dev \
--origin_qwenvl_ckpt_path /path/to/Qwen2.5-VL-3B-Instruct \
--save_path /path/to/output/NarrativeWeaver-init
This produces the NarrativeWeaver-init checkpoint that serves as the starting point for Stage 0 (T2I Pretrain).
.txt file)Each line of the training data txt file contains three comma-separated fields:
/path/to/images_dir,/path/to/annotations.json,false
| Field | Description |
|---|---|
/path/to/images_dir | Root directory containing the images referenced in the JSON |
/path/to/annotations.json | JSON file with annotation entries (see below) |
false | Default flag (keep as false) |
Each entry in the annotation JSON follows this structure:
{
"id": "sample_id",
"image": [
"sample_id/condition.jpg",
"sample_id/output_img_1.png",
"sample_id/output_img_2.png"
],
"conversations": [
{
"from": "human",
"value": "Your instruction prompt here. <image>"
},
{
"from": "gpt",
"value": "Description for image 1. <gen_image>Description for image 2. <gen_image>"
}
]
}
Key points:
image list is the condition / reference image (e.g., a product photo or a character reference).<image> in the human turn refers to the condition / reference image.<gen_image> in the gpt turn acts as a delimiter between descriptions of consecutive generated images. Each text segment before a <gen_image> token describes the corresponding output image.<gen_image> tokens should equal the number of target images (i.e., len(image) - 1).NarrativeWeaver follows a 4-stage training pipeline. Each stage loads from the checkpoint produced by the previous stage.
Before training, set your WandB API key and update every path in the YAML config files to point to your own data and checkpoints:
export WANDB_API_KEY="your_wandb_api_key"
Training uses DeepSpeed ZeRO via HuggingFace Accelerate. Multi-node configs are provided under scripts/accelerate_configs/.
Initializes the learnable MetaQuery and MLP projector by training on text-to-image data, with the LLM frozen.
Config: scripts/denoiser/T2I_pretrain/T2I_flux_qwen2p5vl_3b_vlm_pretrain.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/NarrativeWeaver-init
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
dataset_config:
data_txt: /path/to/data/data_t2i.txt
training_config:
output_dir: /path/to/output/T2I_pretrain
Run:
bash scripts/denoiser/T2I_pretrain/T2I_flux_qwen2p5vl_3b_vlm_pretrain.sh
Fine-tunes the LLM to perform narrative planning — producing per-image textual descriptions from a reference image and an instruction. The denoiser is not updated in this stage.
Config: scripts/denoiser/E-Commerce/Interleaved_e-commerce_step1.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/T2I_pretrain/checkpoint-32000/univa
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
dataset_config:
data_txt: /path/to/data/data_e-commerce_train.txt
validation_json_path: /path/to/data/val.json
validation_image_dir: /path/to/data/images
training_config:
output_dir: /path/to/output/step1_narrative_planning
Run:
bash scripts/denoiser/E-Commerce/Interleaved_e-commerce_step1.sh
Trains the FLUX denoiser end-to-end with the LLM frozen, enabling the model to generate images conditioned on the planned narrative descriptions.
Config: scripts/denoiser/E-Commerce/Interleaved_e-commerce_step2.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/step1/checkpoint/univa
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
dataset_config:
data_txt: /path/to/data/data_e-commerce_train.txt
validation_json_path: /path/to/data/val.json
validation_image_dir: /path/to/data/images
training_config:
output_dir: /path/to/output/step2_visual_generation
Run:
bash scripts/denoiser/E-Commerce/Interleaved_e-commerce_step2.sh
Introduces SigLIP-based identity alignment by injecting fine-grained visual features from the reference image into the denoiser, improving subject consistency across generated images.
Config: scripts/denoiser/E-Commerce/Interleaved_e-commerce_step3.yaml
Key paths to set:
model_config:
pretrained_lvlm_name_or_path: /path/to/step2/checkpoint/univa
pretrained_denoiser_name_or_path: /path/to/FLUX.1-dev
pretrained_siglip_name_or_path: /path/to/siglip2-so400m-patch16-512
pretrained_siglip_mlp_path: /path/to/flux-redux-siglipv2-512.bin
dataset_config:
data_txt: /path/to/data/data_e-commerce_train.txt
validation_json_path: /path/to/data/val.json
validation_image_dir: /path/to/data/images
training_config:
output_dir: /path/to/output/step3_fine_grained_alignment
Run:
bash scripts/denoiser/E-Commerce/Interleaved_e-commerce_step3.sh
A Memory Bank extension of Stage 3 for long-sequence generation. It maintains a cross-image memory buffer that improves consistency across many frames.
Config: scripts/denoiser/MemoryBank/Interleaved_e-commerce_step3.yaml
Run:
bash scripts/denoiser/MemoryBank/Interleaved_e-commerce_step3.sh
All inference scripts live under univa/serve/. Output images are written to --save_dir.
Runs the complete NarrativeWeaver pipeline — the model first generates per-image descriptions from the instruction and reference image, then synthesizes all images.
python -m univa.serve.test_interleaved_flex_visualize \
--model_path /path/to/checkpoint/univa \
--flux_path /path/to/FLUX.1-dev \
--siglip_path /path/to/siglip2-so400m-patch16-512 \
--test_path /path/to/test.json \
--test_image_path /path/to/images \
--height 480 \
--width 832 \
--save_dir /path/to/output \
--stage 2
Skips narrative planning and generates images directly from existing per-image descriptions in the test JSON.
python -m univa.serve.test_interleaved_imageOnly_flex \
--model_path /path/to/checkpoint/univa \
--flux_path /path/to/FLUX.1-dev \
--siglip_path /path/to/siglip2-so400m-patch16-512 \
--test_path /path/to/test.json \
--test_image_path /path/to/images \
--height 480 \
--width 832 \
--save_dir /path/to/output \
--stage 2
Uses a Memory Bank checkpoint for long-sequence generation with enhanced cross-image consistency.
python -m univa.serve.test_interleaved_imageOnly_flex_memoryBank \
--model_path /path/to/memory_bank_checkpoint/univa \
--flux_path /path/to/FLUX.1-dev \
--siglip_path /path/to/siglip2-so400m-patch16-512 \
--test_path /path/to/test.json \
--test_image_path /path/to/images \
--height 480 \
--width 832 \
--save_dir /path/to/output \
--stage 2
If you find this work useful, please cite:
@article{yao2026narrative,
title = {Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning},
author = {Yao, Zhengjian and Li, Yongzhi and Gao, Xinyuan and Chen, Quan and Jiang, Peng and Lu, Yanye},
journal = {arXiv preprint arXiv:2603.06688},
year = {2026}
}
This codebase builds upon UniWorld. We thank the authors for their excellent work and for releasing their code.
1 commits
Python
99.5%