DPO dataset meant to enhance python coding abilities.
This dataset uses the excellent https://huggingface.co/datasets/Vezora/Tested-22k-Python-Alpaca dataset as the "chosen" responses, given this dataset was already tested and validated.
The "rejected" values were generated with a mix of airoboros-l2-13b-3.1 and bagel-7b-v0.1.
The rejected values may actually be perfectly fine, but the assumption here is that the values are generally a lower quality than the chosen counterpart. Items with duplicate code blocks were removed.
If you're interested in new functionality/datasets, take a look at bagel repo and airoboros and either make a PR or open an issue with details.
To help me with the fine-tuning costs, dataset generation, etc., please use one of the following:
6 commits
1 commits
DPO dataset meant to enhance python coding abilities.
This dataset uses the excellent https://huggingface.co/datasets/Vezora/Tested-22k-Python-Alpaca dataset as the "chosen" responses, given this dataset was already tested and validated.
The "rejected" values were generated with a mix of airoboros-l2-13b-3.1 and bagel-7b-v0.1.
The rejected values may actually be perfectly fine, but the assumption here is that the values are generally a lower quality than the chosen counterpart. Items with duplicate code blocks were removed.
If you're interested in new functionality/datasets, take a look at bagel repo and airoboros and either make a PR or open an issue with details.
To help me with the fine-tuning costs, dataset generation, etc., please use one of the following:
6 commits
1 commits