"DeepInnovator: AI Research Assistant - Idea Spark & Scientific Discovery"
See the code
| π‘ Generate Research Ideas and Hypotheses | π Discovers Cross-Disciplinary Connections |
| π Research Gap & Trend Analysis | π οΈ AI-Powered Scientific Problem Solving |
π¬ DeepInnovator is an AI research copilot powered by our built scientific foundation model trained specifically.
π‘ DeepInnovator transforms how researchers discover and develop breakthrough ideas for research discovery.
β’ DeepInnovator-14B significantly outperforms Qwen-14B-Instruct across all evaluation dimensions.
β’ It achieves impressive win rates of 80.53%-93.81% against the base model in automated evaluations.
β’ Despite smaller parameter size, DeepInnovator matches performance of GPT-4o and Gemini-2.5-pro.
β’ DeepInnovator even surpasses GPT-4o in well-justified rationale evaluation, scoring 82.3% vs 77.9%
β’ The model shows strong zero-shot transfer capabilities to completely unseen research domains.
β’ It generates high-quality research ideas in law, education, and biotechnology despite being trained on STEM.
| Domain | Metric | Much Better | Better | Worse | Much Worse | Both Bad | Avg. Winrate (vs Qwen-14B-IT / vs GPT-4o) |
|---|---|---|---|---|---|---|---|
| Law | Novelty | 0 / 0 | 9 / 7 | 3 / 6 | 1 / 0 | 0 / 0 | 69.2% / 53.8% |
| Feasibility | 1 / 0 | 5 / 4 | 3 / 7 | 1 / 1 | 3 / 1 | 60.0% / 33.3% | |
| Effectiveness | 2 / 1 | 5 / 4 | 2 / 3 | 1 / 3 | 3 / 2 | 70.0% / 45.5% | |
| Detailedness | 7 / 1 | 3 / 4 | 2 / 5 | 1 / 0 | 0 / 3 | 76.9% / 50.0% | |
| Education | Novelty | 3 / 0 | 9 / 9 | 2 / 5 | 1 / 1 | 0 / 0 | 80.0% / 60.0% |
| Feasibility | 3 / 0 | 5 / 0 | 5 / 7 | 1 / 3 | 1 / 5 | 57.1% / 0.0% | |
| Effectiveness | 2 / 0 | 6 / 0 | 4 / 5 | 3 / 4 | 0 / 6 | 53.3% / 0.0% | |
| Detailedness | 1 / 0 | 8 / 4 | 3 / 5 | 0 / 2 | 3 / 4 | 75.0% / 36.4% | |
| Biotech | Novelty | 3 / 2 | 8 / 6 | 1 / 5 | 0 / 0 | 2 / 2 | 91.7% / 61.5% |
| Feasibility | 2 / 0 | 6 / 5 | 0 / 3 | 0 / 5 | 6 / 2 | 100.0% / 38.5% | |
| Effectiveness | 1 / 2 | 6 / 5 | 4 / 2 | 1 / 4 | 2 / 2 | 58.3% / 53.8% | |
| Detailedness | 6 / 3 | 4 / 2 | 1 / 4 | 0 / 5 | 3 / 1 | 90.9% / 35.7% |
β’ DeepInnovator-14B achieves 100% win rate in Biotech Feasibility against Qwen2.5-14B-IT
β’ Demonstrates 91.7% win rate in Biotech Novelty versus the baseline model
β’ Secures 90.9% win rate in Biotech Detailedness compared to Qwen2.5-14B-IT
β’ Maintains 61.5% win rate in Biotech Novelty when benchmarked against GPT-4o
β’ Shows 60.0% win rate in Education Novelty against the advanced GPT-4o model

DeepInnovator/
βββ recipe/
β βββ DeepInnovator/
β βββ data_preparation/ # Data preparation pipeline
β β βββ config/ # Agent and model configurations
β β βββ data_prepare/ # Pipeline scripts
β β βββ run.sh # Quick run script
β β βββ README.md # Data preparation documentation
β βββ config/ # Training configurations
β β βββ agent.yaml # Agent loop configuration
β β βββ reward_config.yaml # Reward function configuration
β β βββ ResearchGAN_interaction_config.yaml # Interaction configuration
β βββ metrics/ # Reward metrics
β β βββ basic_reward.py
β β βββ delta_reward.py
β β βββ token_amount.py
β βββ preprocess.py # Dataset preprocessing script
β βββ preprocess.sh # Preprocessing script runner
β βββ reward_function.py # Main reward function
β βββ DeepInnovator_interation.py # Interaction logic
β βββ DeepInnovator_agent_loop.py # Agent loop implementation
β βββ train_rl.sh # Training script
β βββ utils.py # Utility functions
βββ verl/ # VERL framework (for RL training)
# Core dependencies
pip install openai omegaconf python-dotenv feedparser requests PyPDF2 tqdm python-dateutil
pip install datasets numpy torch transformers
Create a .env file in the project root:
# API Configuration for data preparation
OPENAI_API_BASE=your_api_base_url
OPENAI_API_KEY=your_api_key
# For training (if needed)
WANDB_API_KEY=your_wandb_api_key
WANDB_BASE_URL=your_wandb_base_url
Edit recipe/DeepInnovator/data_preparation/config/models/providers.yaml to set your API endpoints:
openai:
base_url: ${env:OPENAI_API_BASE}
api_key: ${env:OPENAI_API_KEY}
The data preparation pipeline processes academic papers through multiple stages to generate training data.
Run the complete pipeline:
cd recipe/DeepInnovator/data_preparation
bash run.sh [total_papers] [datapath]
Example:
cd recipe/DeepInnovator/data_preparation
bash run.sh 100 ./data/arxiv_data
Download papers from arXiv across predefined categories (cs, stat, q-fin, math):
cd recipe/DeepInnovator/data_preparation
python data_prepare/pull_papers.py --total_papers 100 --datapath ./data/arxiv_data
Parameters:
--total_papers: Total number of papers to download--datapath: Data save pathOutput: Papers saved to {datapath}/raw_paper/ directory
Extract ideas from target papers:
cd recipe/DeepInnovator/data_preparation
python data_prepare/get_target_paper_idea.py --datapath ./data/arxiv_data
Output: {datapath}/{paper_id}/target_paper/raw_paper/paper_idea.json
Process papers through the full pipeline (step1-step4) to generate training data:
cd recipe/DeepInnovator/data_preparation
python data_prepare/get_training_data.py
Output Structure:
layer0/: Paper analysis resultslayer1/: Paper groups and memorieslayer2/: Connections, serendipity, and trendsinsights/: Generated research ideasAfter generating training data, preprocess it for RL training:
cd recipe/DeepInnovator
python preprocess.py \
--input_dir ./data/arxiv_data \
--output_dir ./data/train \
--task_desc "refine a research idea" \
--validation_size 0.1 \
--seed 42 \
--num_proc 1 \
--dataset_type "rl" \
--test False \
--layer0 False \
--layer1 False \
--layer2 True
Parameters:
--input_dir: Input directory containing processed papers--output_dir: Output directory for preprocessed data--task_desc: Task description for the dataset--validation_size: Validation split ratio (default: 0.1)--seed: Random seed (default: 42)--num_proc: Number of parallel workers (default: 1)--dataset_type: Dataset type - "rl" or "sft" (default: "rl")--test: Test mode - sample fixed number of examples (default: True)--layer0/1/2: Include layer data in prompts (default: False)Output:
rl_train.parquet: Training datasetrl_validation.parquet: Validation datasetOr use the convenience script:
cd recipe/DeepInnovator
bash preprocess.sh
Before training, configure the following files:
recipe/DeepInnovator/config/reward_config.yaml: Configure reward function parameters
config:
metric_weights:
delta_reward: 5
token_amount: 0.1
default_reward_kwargs:
model: "your model"
api_key: "your api key"
api_base: "your api base"
recipe/DeepInnovator/config/ResearchGAN_interaction_config.yaml: Configure interaction settings
interaction:
- name: "DeepInnovator"
discriminator_kwargs:
discriminator_model: "your model"
api_key: "your api key"
api_base: "your api base"
recipe/DeepInnovator/train_rl.sh: Update training parameters
MODEL_DIR: Path to base modelDATASET_DIR: Path to preprocessed datasetWANDB_PROJECT_NAME: Weights & Biases project nameWANDB_EXPERIMENT_NAME: Experiment namecd recipe/DeepInnovator
bash train_rl.sh [resume_path]
Parameters:
resume_path (optional): Path to checkpoint to resume fromTraining Configuration:
The training process involves:
recipe/DeepInnovator/reward_function.py)Computes conversation-level rewards by combining multiple metrics:
delta_reward: Measures improvement between iterationstoken_amount: Length-based rewardreward_config.yamlrecipe/DeepInnovator/DeepInnovator_interation.py)Implements the interaction logic:
recipe/DeepInnovator/DeepInnovator_agent_loop.py)Manages the agent's decision-making process:
After data preparation:
data/
βββ {paper_id}/
βββ target_paper/ # Target paper and references
β βββ raw_paper/
β βββ paper_md/ # Markdown files
β βββ paper_idea.json # Extracted ideas
βββ raw_paper/ # Raw downloaded papers
βββ layer0/ # Paper analysis
β βββ paper_memory/ # Structured paper data
βββ layer1/ # Paper grouping
β βββ inner_paper_memory.json
β βββ inter_paper_group.json
βββ layer2/ # Connections and insights
β βββ connections.json
β βββ serendipity.json
β βββ research_trending.json
βββ insights/ # Generated ideas
βββ idea_spark.json
After preprocessing:
data/train/
βββ rl_train.parquet # Training dataset
βββ rl_validation.parquet # Validation dataset
This project is licensed under the MIT License - see the LICENSE file for details.
@article{fan2026deepinnovator,
title={DeepInnovator: Triggering the Innovative Capabilities of LLMs},
author={Fan, Tianyu and Zhang, Fengji and Zheng, Yuxiang and Chen, Bei and Niu, Xinyao and Huang, Chengen and Lin, Junyang and Huang, Chao},
journal={arXiv preprint arXiv:2602.18920},
year={2026}
}
Thanks for visiting β¨ DeepInnovator!
Python
96.2%
Shell
3.8%
"DeepInnovator: AI Research Assistant - Idea Spark & Scientific Discovery"
See the code
| π‘ Generate Research Ideas and Hypotheses | π Discovers Cross-Disciplinary Connections |
| π Research Gap & Trend Analysis | π οΈ AI-Powered Scientific Problem Solving |
π¬ DeepInnovator is an AI research copilot powered by our built scientific foundation model trained specifically.
π‘ DeepInnovator transforms how researchers discover and develop breakthrough ideas for research discovery.
β’ DeepInnovator-14B significantly outperforms Qwen-14B-Instruct across all evaluation dimensions.
β’ It achieves impressive win rates of 80.53%-93.81% against the base model in automated evaluations.
β’ Despite smaller parameter size, DeepInnovator matches performance of GPT-4o and Gemini-2.5-pro.
β’ DeepInnovator even surpasses GPT-4o in well-justified rationale evaluation, scoring 82.3% vs 77.9%
β’ The model shows strong zero-shot transfer capabilities to completely unseen research domains.
β’ It generates high-quality research ideas in law, education, and biotechnology despite being trained on STEM.
| Domain | Metric | Much Better | Better | Worse | Much Worse | Both Bad | Avg. Winrate (vs Qwen-14B-IT / vs GPT-4o) |
|---|---|---|---|---|---|---|---|
| Law | Novelty | 0 / 0 | 9 / 7 | 3 / 6 | 1 / 0 | 0 / 0 | 69.2% / 53.8% |
| Feasibility | 1 / 0 | 5 / 4 | 3 / 7 | 1 / 1 | 3 / 1 | 60.0% / 33.3% | |
| Effectiveness | 2 / 1 | 5 / 4 | 2 / 3 | 1 / 3 | 3 / 2 | 70.0% / 45.5% | |
| Detailedness | 7 / 1 | 3 / 4 | 2 / 5 | 1 / 0 | 0 / 3 | 76.9% / 50.0% | |
| Education | Novelty | 3 / 0 | 9 / 9 | 2 / 5 | 1 / 1 | 0 / 0 | 80.0% / 60.0% |
| Feasibility | 3 / 0 | 5 / 0 | 5 / 7 | 1 / 3 | 1 / 5 | 57.1% / 0.0% | |
| Effectiveness | 2 / 0 | 6 / 0 | 4 / 5 | 3 / 4 | 0 / 6 | 53.3% / 0.0% | |
| Detailedness | 1 / 0 | 8 / 4 | 3 / 5 | 0 / 2 | 3 / 4 | 75.0% / 36.4% | |
| Biotech | Novelty | 3 / 2 | 8 / 6 | 1 / 5 | 0 / 0 | 2 / 2 | 91.7% / 61.5% |
| Feasibility | 2 / 0 | 6 / 5 | 0 / 3 | 0 / 5 | 6 / 2 | 100.0% / 38.5% | |
| Effectiveness | 1 / 2 | 6 / 5 | 4 / 2 | 1 / 4 | 2 / 2 | 58.3% / 53.8% | |
| Detailedness | 6 / 3 | 4 / 2 | 1 / 4 | 0 / 5 | 3 / 1 | 90.9% / 35.7% |
β’ DeepInnovator-14B achieves 100% win rate in Biotech Feasibility against Qwen2.5-14B-IT
β’ Demonstrates 91.7% win rate in Biotech Novelty versus the baseline model
β’ Secures 90.9% win rate in Biotech Detailedness compared to Qwen2.5-14B-IT
β’ Maintains 61.5% win rate in Biotech Novelty when benchmarked against GPT-4o
β’ Shows 60.0% win rate in Education Novelty against the advanced GPT-4o model

DeepInnovator/
βββ recipe/
β βββ DeepInnovator/
β βββ data_preparation/ # Data preparation pipeline
β β βββ config/ # Agent and model configurations
β β βββ data_prepare/ # Pipeline scripts
β β βββ run.sh # Quick run script
β β βββ README.md # Data preparation documentation
β βββ config/ # Training configurations
β β βββ agent.yaml # Agent loop configuration
β β βββ reward_config.yaml # Reward function configuration
β β βββ ResearchGAN_interaction_config.yaml # Interaction configuration
β βββ metrics/ # Reward metrics
β β βββ basic_reward.py
β β βββ delta_reward.py
β β βββ token_amount.py
β βββ preprocess.py # Dataset preprocessing script
β βββ preprocess.sh # Preprocessing script runner
β βββ reward_function.py # Main reward function
β βββ DeepInnovator_interation.py # Interaction logic
β βββ DeepInnovator_agent_loop.py # Agent loop implementation
β βββ train_rl.sh # Training script
β βββ utils.py # Utility functions
βββ verl/ # VERL framework (for RL training)
# Core dependencies
pip install openai omegaconf python-dotenv feedparser requests PyPDF2 tqdm python-dateutil
pip install datasets numpy torch transformers
Create a .env file in the project root:
# API Configuration for data preparation
OPENAI_API_BASE=your_api_base_url
OPENAI_API_KEY=your_api_key
# For training (if needed)
WANDB_API_KEY=your_wandb_api_key
WANDB_BASE_URL=your_wandb_base_url
Edit recipe/DeepInnovator/data_preparation/config/models/providers.yaml to set your API endpoints:
openai:
base_url: ${env:OPENAI_API_BASE}
api_key: ${env:OPENAI_API_KEY}
The data preparation pipeline processes academic papers through multiple stages to generate training data.
Run the complete pipeline:
cd recipe/DeepInnovator/data_preparation
bash run.sh [total_papers] [datapath]
Example:
cd recipe/DeepInnovator/data_preparation
bash run.sh 100 ./data/arxiv_data
Download papers from arXiv across predefined categories (cs, stat, q-fin, math):
cd recipe/DeepInnovator/data_preparation
python data_prepare/pull_papers.py --total_papers 100 --datapath ./data/arxiv_data
Parameters:
--total_papers: Total number of papers to download--datapath: Data save pathOutput: Papers saved to {datapath}/raw_paper/ directory
Extract ideas from target papers:
cd recipe/DeepInnovator/data_preparation
python data_prepare/get_target_paper_idea.py --datapath ./data/arxiv_data
Output: {datapath}/{paper_id}/target_paper/raw_paper/paper_idea.json
Process papers through the full pipeline (step1-step4) to generate training data:
cd recipe/DeepInnovator/data_preparation
python data_prepare/get_training_data.py
Output Structure:
layer0/: Paper analysis resultslayer1/: Paper groups and memorieslayer2/: Connections, serendipity, and trendsinsights/: Generated research ideasAfter generating training data, preprocess it for RL training:
cd recipe/DeepInnovator
python preprocess.py \
--input_dir ./data/arxiv_data \
--output_dir ./data/train \
--task_desc "refine a research idea" \
--validation_size 0.1 \
--seed 42 \
--num_proc 1 \
--dataset_type "rl" \
--test False \
--layer0 False \
--layer1 False \
--layer2 True
Parameters:
--input_dir: Input directory containing processed papers--output_dir: Output directory for preprocessed data--task_desc: Task description for the dataset--validation_size: Validation split ratio (default: 0.1)--seed: Random seed (default: 42)--num_proc: Number of parallel workers (default: 1)--dataset_type: Dataset type - "rl" or "sft" (default: "rl")--test: Test mode - sample fixed number of examples (default: True)--layer0/1/2: Include layer data in prompts (default: False)Output:
rl_train.parquet: Training datasetrl_validation.parquet: Validation datasetOr use the convenience script:
cd recipe/DeepInnovator
bash preprocess.sh
Before training, configure the following files:
recipe/DeepInnovator/config/reward_config.yaml: Configure reward function parameters
config:
metric_weights:
delta_reward: 5
token_amount: 0.1
default_reward_kwargs:
model: "your model"
api_key: "your api key"
api_base: "your api base"
recipe/DeepInnovator/config/ResearchGAN_interaction_config.yaml: Configure interaction settings
interaction:
- name: "DeepInnovator"
discriminator_kwargs:
discriminator_model: "your model"
api_key: "your api key"
api_base: "your api base"
recipe/DeepInnovator/train_rl.sh: Update training parameters
MODEL_DIR: Path to base modelDATASET_DIR: Path to preprocessed datasetWANDB_PROJECT_NAME: Weights & Biases project nameWANDB_EXPERIMENT_NAME: Experiment namecd recipe/DeepInnovator
bash train_rl.sh [resume_path]
Parameters:
resume_path (optional): Path to checkpoint to resume fromTraining Configuration:
The training process involves:
recipe/DeepInnovator/reward_function.py)Computes conversation-level rewards by combining multiple metrics:
delta_reward: Measures improvement between iterationstoken_amount: Length-based rewardreward_config.yamlrecipe/DeepInnovator/DeepInnovator_interation.py)Implements the interaction logic:
recipe/DeepInnovator/DeepInnovator_agent_loop.py)Manages the agent's decision-making process:
After data preparation:
data/
βββ {paper_id}/
βββ target_paper/ # Target paper and references
β βββ raw_paper/
β βββ paper_md/ # Markdown files
β βββ paper_idea.json # Extracted ideas
βββ raw_paper/ # Raw downloaded papers
βββ layer0/ # Paper analysis
β βββ paper_memory/ # Structured paper data
βββ layer1/ # Paper grouping
β βββ inner_paper_memory.json
β βββ inter_paper_group.json
βββ layer2/ # Connections and insights
β βββ connections.json
β βββ serendipity.json
β βββ research_trending.json
βββ insights/ # Generated ideas
βββ idea_spark.json
After preprocessing:
data/train/
βββ rl_train.parquet # Training dataset
βββ rl_validation.parquet # Validation dataset
This project is licensed under the MIT License - see the LICENSE file for details.
@article{fan2026deepinnovator,
title={DeepInnovator: Triggering the Innovative Capabilities of LLMs},
author={Fan, Tianyu and Zhang, Fengji and Zheng, Yuxiang and Chen, Bei and Niu, Xinyao and Huang, Chengen and Lin, Junyang and Huang, Chao},
journal={arXiv preprint arXiv:2602.18920},
year={2026}
}
Thanks for visiting β¨ DeepInnovator!
Python
96.2%
Shell
3.8%