Following the principles highlighted in arXiv:2305.11206 by Zhou et al. and replicated in some aspects by Kaiokendev with SuperHOT, the archive in this repository contains about 2000 manually selected and curated 1-on-1 human-human roleplaying conversations and associated LLM-generated persona and scenario data. The RP conversations all feature only two human participants, although occasionally the participants may play the role of more than one character.
The conversation data is in the form of source files in .yaml format + basic Python script for building the dataset, intended to be finetuned in "completion" format (similar to unsupervised finetuning).
Having reached the minimum number of examples suggested in the LIMA paper and after putting overall probably more than 500 hours of work on manually gathering and curating the data, LimaRP can be considered a finished project at this point in time. Future work (cleaning, trimming, expansion) would require more resources and community help.
data-long) were designed for an 8192
tokens context size. Note that while the 8k samples can be reduced to 4k size, this can confuse the model to
some extent, as scenario and persona data may end up referring to events removed from the context.LimaRPLimaRP has a few notable issues, here in subjective decreasing order of severity.
data-short). The data needs to be carefully
checked to make sure that no issue in this regard exists.gpt-4-isms and can be repetitive, lack depth and miss certain character traits; manual
editing will be needed to make them more human-like and respond to more specialized personality
traits and keywords—as a result, LimaRP-generated text may appear to ignore certain character traits.
A more powerful personality summarizer capable of being both accurate while generating sufficiently
long descriptions could be conceived for solving this issue._MULTI
or _GROUP in the filename) have been assigned a name with the format Char1&Char2. Testing didn't reveal issues
with this, but it's something to keep in mind if more severe impersonation problems occur compared to the initial
release of LimaRP. Furthermore, in a few conversations additional characters (roleplayed by either of the two users)
may also temporarily participate to the story. These have often (but not always) been assigned a _BAD tag in the filename.content field and in most cases they
contain keywords like shemale, futa, futanari, trans, transgender when relevant to assist filtering.Only one format has been used: forum/novel-style. This includes:
Other RP styles have been excluded, and messages showing them have been fixed when possible and feasible.
Jessica looked at Mark with disdain."I say this."*thud*_What is he doing?_''The Jungle Book''... with a trailing space when a word follows) and em-dashes always converted
to three consecutive dashes (---) without any surrounding space.
—).<FIRST> is always
assumed to be the bot/model, and <SECOND> always assumed to be the human/user. All conversations terminate with
a message by <FIRST>.
Weights are naively calculated in terms of bytes for the entire conversation files as of 2023-11-10.
| Source | Notes | Weight |
|---|---|---|
| All The Fallen | Registration required | 5.1% |
| Black Dahlia Roleplaying | Registration required, 18+ characters only | 0.9% |
| Blue Moon Roleplaying | Mostly open-access, Lolisho forbidden | 18.4% |
| Darknest Fantasy | Registration required, 18+ characters only | 0.2% |
| Eka's Portal | Open-access | 1.6% |
| Elliquiy | Approval required, Lolisho forbidden | 50.8% |
| Lolicit | Registration required, Defunct website | 10.5% |
| Redlight Ponyville | Approval required | 0.6% |
| The Inner Sanctum | Registration required, 18+ characters only | 11.8% |
Note that users are required to be 18+ to write in the listed ERP forums or forum subsections.
Usernames, OOC and other personal information have not been included in the training data, only the names of the roleplayed characters as used in the conversations (or sometimes with minor changes).
Ideas in random order that could be applied for improving the dataset. Some have been already mentioned earlier.
tokens/10
40 commits
1 commits
Following the principles highlighted in arXiv:2305.11206 by Zhou et al. and replicated in some aspects by Kaiokendev with SuperHOT, the archive in this repository contains about 2000 manually selected and curated 1-on-1 human-human roleplaying conversations and associated LLM-generated persona and scenario data. The RP conversations all feature only two human participants, although occasionally the participants may play the role of more than one character.
The conversation data is in the form of source files in .yaml format + basic Python script for building the dataset, intended to be finetuned in "completion" format (similar to unsupervised finetuning).
Having reached the minimum number of examples suggested in the LIMA paper and after putting overall probably more than 500 hours of work on manually gathering and curating the data, LimaRP can be considered a finished project at this point in time. Future work (cleaning, trimming, expansion) would require more resources and community help.
data-long) were designed for an 8192
tokens context size. Note that while the 8k samples can be reduced to 4k size, this can confuse the model to
some extent, as scenario and persona data may end up referring to events removed from the context.LimaRPLimaRP has a few notable issues, here in subjective decreasing order of severity.
data-short). The data needs to be carefully
checked to make sure that no issue in this regard exists.gpt-4-isms and can be repetitive, lack depth and miss certain character traits; manual
editing will be needed to make them more human-like and respond to more specialized personality
traits and keywords—as a result, LimaRP-generated text may appear to ignore certain character traits.
A more powerful personality summarizer capable of being both accurate while generating sufficiently
long descriptions could be conceived for solving this issue._MULTI
or _GROUP in the filename) have been assigned a name with the format Char1&Char2. Testing didn't reveal issues
with this, but it's something to keep in mind if more severe impersonation problems occur compared to the initial
release of LimaRP. Furthermore, in a few conversations additional characters (roleplayed by either of the two users)
may also temporarily participate to the story. These have often (but not always) been assigned a _BAD tag in the filename.content field and in most cases they
contain keywords like shemale, futa, futanari, trans, transgender when relevant to assist filtering.Only one format has been used: forum/novel-style. This includes:
Other RP styles have been excluded, and messages showing them have been fixed when possible and feasible.
Jessica looked at Mark with disdain."I say this."*thud*_What is he doing?_''The Jungle Book''... with a trailing space when a word follows) and em-dashes always converted
to three consecutive dashes (---) without any surrounding space.
—).<FIRST> is always
assumed to be the bot/model, and <SECOND> always assumed to be the human/user. All conversations terminate with
a message by <FIRST>.
Weights are naively calculated in terms of bytes for the entire conversation files as of 2023-11-10.
| Source | Notes | Weight |
|---|---|---|
| All The Fallen | Registration required | 5.1% |
| Black Dahlia Roleplaying | Registration required, 18+ characters only | 0.9% |
| Blue Moon Roleplaying | Mostly open-access, Lolisho forbidden | 18.4% |
| Darknest Fantasy | Registration required, 18+ characters only | 0.2% |
| Eka's Portal | Open-access | 1.6% |
| Elliquiy | Approval required, Lolisho forbidden | 50.8% |
| Lolicit | Registration required, Defunct website | 10.5% |
| Redlight Ponyville | Approval required | 0.6% |
| The Inner Sanctum | Registration required, 18+ characters only | 11.8% |
Note that users are required to be 18+ to write in the listed ERP forums or forum subsections.
Usernames, OOC and other personal information have not been included in the training data, only the names of the roleplayed characters as used in the conversations (or sometimes with minor changes).
Ideas in random order that could be applied for improving the dataset. Some have been already mentioned earlier.
tokens/10
40 commits
1 commits