solis-team/Hydra

[FSE 2026] Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond

Python

138

28 commits

updated Jun 4, 2026

See the code

README

Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond

Repository-Level Code Generation • Structure-Aware Indexing • Dependency-Aware Retrieval

Python version License: Apache 2.0 arXiv HF Model HF Data

Quick Start • Documentation • Paper

Introduction

Large language models for code (CodeLLMs) have demonstrated remarkable success in standalone code completion and generation, yet their effectiveness diminishes in repository-level settings where cross-file dependencies and structural context are essential. Existing Retrieval-Augmented Generation (RAG) approaches often borrow strategies from NLP, relying on chunking-based indexing and similarity-based retrieval that overlook structural relationships and miss functionally relevant dependencies.

We present Hydra, a repository-level code generation framework that treats code as structured code rather than natural language. Our approach introduces: (i) structure-aware indexing that preserves code structure and dependencies, (ii) a lightweight dependency-aware retriever (DAR) that identifies true dependencies, and (iii) hybrid retrieval combining dependency-aware and similarity-based methods.

Extensive experiments on DevEval and RepoExec benchmarks show that Hydra achieves state-of-the-art performance, surpassing the strongest baseline by over 5% in Pass@1 and enabling smaller models to match larger ones.

Quick Start

Prerequisites

Create a new conda environment and install dependencies:

# Create conda environment
conda create -n hydra python=3.10.0
conda activate hydra

# Install required packages
pip install -r requirements.txt

Setup and Installation

Important: You must complete the following setup steps before running any experiments.

  1. Extract benchmark data:

    cd data
    unzip temp.zip
    
    # Extract RepoExec benchmark
    cd ../benchmark/RepoExec
    unzip test-apps.zip
    
    # Extract DevEval benchmark
    cd ../DevEval
    tar -xzf data.tar.gz
    wget https://huggingface.co/datasets/LJ0815/DevEval/resolve/main/Source_Code.tar.gz
    tar -xvzf Source_Code.tar.gz
    
  2. Prepare structured context (required for experiments):

    # For RepoExec benchmark
    bash src/context_formulation/structured_indexer/run.sh --dataset RepoExec
    
    # For DevEval benchmark  
    bash src/context_formulation/structured_indexer/run.sh --dataset DevEval
    

Main Results

Comparison of Hydra with prior retrieval-based approaches and no-context baselines. Results are reported in Pass@1/3/5.

GPT-4.1 mini

MethodRepoExec Pass@1RepoExec Pass@3RepoExec Pass@5DevEval Pass@1DevEval Pass@3DevEval Pass@5
No Context21.5824.4225.6319.7223.1924.71
RepoCoder22.2026.0827.8917.4823.1525.70
RepoFormer39.1542.4243.9430.8934.2135.40
RLCoder38.1442.1743.3829.4632.7634.14
Hydra43.5545.7246.4831.9135.5636.99

QwenCoder-1.5B-Instruct

MethodRepoExec Pass@1RepoExec Pass@3RepoExec Pass@5DevEval Pass@1DevEval Pass@3DevEval Pass@5
No Context5.758.319.303.535.205.97
RepoCoder7.1511.7214.374.548.089.81
RepoFormer11.1516.4218.875.587.948.99
RLCoder14.8721.0423.949.3412.9014.47
Hydra15.7221.3023.3810.7114.5016.05

QwenCoder-7B-Instruct

MethodRepoExec Pass@1RepoExec Pass@3RepoExec Pass@5DevEval Pass@1DevEval Pass@3DevEval Pass@5
No Context13.3017.0418.037.109.1610.03
RepoCoder14.8221.9925.076.3910.6312.82
RepoFormer17.6925.0428.4510.4113.6814.90
RLCoder20.1723.6927.6113.0017.6719.61
Hydra23.3231.3234.3617.2722.4424.44

Documentation

Important Before reproducing experiments, you must first train the DAR (Dependency-Aware Retriever) model or using our pretrained model at huggingface DAR model.

For detailed instructions and comprehensive guides, please refer to:

  • Training.md - DAR (Dependency-Aware Retriever) training guide including:

    • Dataset construction methodology
    • Model architecture and training procedures
  • Reproduce.md - Complete experimental reproduction guide including:

    • Benchmark setup and data preparation
    • Research questions reproduction (RQ1-RQ4)
    • Code generation pipeline
    • Evaluation and metrics calculation

Citation

If you found this repository to be useful, please cite:

@misc{leanh2026treatcodenaturallanguage,
      title={Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond}, 
      author={Minh Le-Anh and Huyen Nguyen and Khanh An Tran and Nam Le Hai and Linh Ngo Van and Nghi D. Q. Bui and Bach Le},
      year={2026},
      eprint={2602.11671},
      archivePrefix={arXiv},
      primaryClass={cs.SE},
      url={https://arxiv.org/abs/2602.11671}, 
}
ai4se
code-generation

solis-team/Hydra

[FSE 2026] Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond

Python

138

28 commits

updated Jun 4, 2026

See the code

README

Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond

Repository-Level Code Generation • Structure-Aware Indexing • Dependency-Aware Retrieval

Python version License: Apache 2.0 arXiv HF Model HF Data

Quick Start • Documentation • Paper

Introduction

Large language models for code (CodeLLMs) have demonstrated remarkable success in standalone code completion and generation, yet their effectiveness diminishes in repository-level settings where cross-file dependencies and structural context are essential. Existing Retrieval-Augmented Generation (RAG) approaches often borrow strategies from NLP, relying on chunking-based indexing and similarity-based retrieval that overlook structural relationships and miss functionally relevant dependencies.

We present Hydra, a repository-level code generation framework that treats code as structured code rather than natural language. Our approach introduces: (i) structure-aware indexing that preserves code structure and dependencies, (ii) a lightweight dependency-aware retriever (DAR) that identifies true dependencies, and (iii) hybrid retrieval combining dependency-aware and similarity-based methods.

Extensive experiments on DevEval and RepoExec benchmarks show that Hydra achieves state-of-the-art performance, surpassing the strongest baseline by over 5% in Pass@1 and enabling smaller models to match larger ones.

Quick Start

Prerequisites

Create a new conda environment and install dependencies:

# Create conda environment
conda create -n hydra python=3.10.0
conda activate hydra

# Install required packages
pip install -r requirements.txt

Setup and Installation

Important: You must complete the following setup steps before running any experiments.

  1. Extract benchmark data:

    cd data
    unzip temp.zip
    
    # Extract RepoExec benchmark
    cd ../benchmark/RepoExec
    unzip test-apps.zip
    
    # Extract DevEval benchmark
    cd ../DevEval
    tar -xzf data.tar.gz
    wget https://huggingface.co/datasets/LJ0815/DevEval/resolve/main/Source_Code.tar.gz
    tar -xvzf Source_Code.tar.gz
    
  2. Prepare structured context (required for experiments):

    # For RepoExec benchmark
    bash src/context_formulation/structured_indexer/run.sh --dataset RepoExec
    
    # For DevEval benchmark  
    bash src/context_formulation/structured_indexer/run.sh --dataset DevEval
    

Main Results

Comparison of Hydra with prior retrieval-based approaches and no-context baselines. Results are reported in Pass@1/3/5.

GPT-4.1 mini

MethodRepoExec Pass@1RepoExec Pass@3RepoExec Pass@5DevEval Pass@1DevEval Pass@3DevEval Pass@5
No Context21.5824.4225.6319.7223.1924.71
RepoCoder22.2026.0827.8917.4823.1525.70
RepoFormer39.1542.4243.9430.8934.2135.40
RLCoder38.1442.1743.3829.4632.7634.14
Hydra43.5545.7246.4831.9135.5636.99

QwenCoder-1.5B-Instruct

MethodRepoExec Pass@1RepoExec Pass@3RepoExec Pass@5DevEval Pass@1DevEval Pass@3DevEval Pass@5
No Context5.758.319.303.535.205.97
RepoCoder7.1511.7214.374.548.089.81
RepoFormer11.1516.4218.875.587.948.99
RLCoder14.8721.0423.949.3412.9014.47
Hydra15.7221.3023.3810.7114.5016.05

QwenCoder-7B-Instruct

MethodRepoExec Pass@1RepoExec Pass@3RepoExec Pass@5DevEval Pass@1DevEval Pass@3DevEval Pass@5
No Context13.3017.0418.037.109.1610.03
RepoCoder14.8221.9925.076.3910.6312.82
RepoFormer17.6925.0428.4510.4113.6814.90
RLCoder20.1723.6927.6113.0017.6719.61
Hydra23.3231.3234.3617.2722.4424.44

Documentation

Important Before reproducing experiments, you must first train the DAR (Dependency-Aware Retriever) model or using our pretrained model at huggingface DAR model.

For detailed instructions and comprehensive guides, please refer to:

  • Training.md - DAR (Dependency-Aware Retriever) training guide including:

    • Dataset construction methodology
    • Model architecture and training procedures
  • Reproduce.md - Complete experimental reproduction guide including:

    • Benchmark setup and data preparation
    • Research questions reproduction (RQ1-RQ4)
    • Code generation pipeline
    • Evaluation and metrics calculation

Citation

If you found this repository to be useful, please cite:

@misc{leanh2026treatcodenaturallanguage,
      title={Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond}, 
      author={Minh Le-Anh and Huyen Nguyen and Khanh An Tran and Nam Le Hai and Linh Ngo Van and Nghi D. Q. Bui and Bach Le},
      year={2026},
      eprint={2602.11671},
      archivePrefix={arXiv},
      primaryClass={cs.SE},
      url={https://arxiv.org/abs/2602.11671}, 
}
ai4se
code-generation

Languages

Python

97.9%

Shell

1.4%