andersonbcdefg/synthetic_retrieval_tasks

Dataset

79

stars

9

commits

1

linked in READMEs

Feb 3, 2024

updated

synthetic

README

Synthetic data designed as prompts for generating embeddings training data for retrieval. The "iteration" column refers to how the data was generated.

Iteration 1: Use the following pool of seed tasks, prompt GPT-3.5-Turbo to generate additional tasks.

RETRIEVAL_EXAMPLES = [
  'Provide a scientific claim as query, retrieve documents that help verify or refute the claim.',
  'Search for documents that answers a FAQ-style query on children\'s nutrition.',
  "Retrieve company's financial reports for a given stock ticker symbol.",
  "Given a book name as a query, retrieve reviews, ratings and summaries of that book.",
  "Search for scientific research papers supporting a medical diagnosis for a specified disease.",
  "Given a question, retrieve Wikipedia passages that answer the question.",
  "Provided a user question, retrieve the highest voted answers on Reddit ELI5 forum.",
  "Given a web search engine query, retrieve relevant passages that answer the query.",
  "Find Amazon reviews similar to the input review.",
  "Find the song lyrics most related to the user's search.",
  "Given a multi-hop question, retrieve documents that can help answer the question.",
  "Retrieve tweets that are semantically similar to the given tweet",
  "Given a news summary, retrieve other semantically similar summaries",
  "Given a question, retrieve relevant answers from Stackexchange",
  "Given a scientific paper title, retrieve paper abstracts that are cited by the given paper."
]

Iteration 2: Use the ~40,000 tasks generated in Iteration 1 as seed tasks, prompt GPT-3.5-Turbo to generate additional tasks.

Iteration 3: Use the ~80,000 tasks generated in Iterations 1-2 as seed tasks, prompt GPT-4-Turbo to generate additional tasks.

Contributors

davanstrien

1 commits

andersonbcdefg/synthetic_retrieval_tasks

Dataset

79

stars

9

commits

1

linked in READMEs

Feb 3, 2024

updated

synthetic

README

Synthetic data designed as prompts for generating embeddings training data for retrieval. The "iteration" column refers to how the data was generated.

Iteration 1: Use the following pool of seed tasks, prompt GPT-3.5-Turbo to generate additional tasks.

RETRIEVAL_EXAMPLES = [
  'Provide a scientific claim as query, retrieve documents that help verify or refute the claim.',
  'Search for documents that answers a FAQ-style query on children\'s nutrition.',
  "Retrieve company's financial reports for a given stock ticker symbol.",
  "Given a book name as a query, retrieve reviews, ratings and summaries of that book.",
  "Search for scientific research papers supporting a medical diagnosis for a specified disease.",
  "Given a question, retrieve Wikipedia passages that answer the question.",
  "Provided a user question, retrieve the highest voted answers on Reddit ELI5 forum.",
  "Given a web search engine query, retrieve relevant passages that answer the query.",
  "Find Amazon reviews similar to the input review.",
  "Find the song lyrics most related to the user's search.",
  "Given a multi-hop question, retrieve documents that can help answer the question.",
  "Retrieve tweets that are semantically similar to the given tweet",
  "Given a news summary, retrieve other semantically similar summaries",
  "Given a question, retrieve relevant answers from Stackexchange",
  "Given a scientific paper title, retrieve paper abstracts that are cited by the given paper."
]

Iteration 2: Use the ~40,000 tasks generated in Iteration 1 as seed tasks, prompt GPT-3.5-Turbo to generate additional tasks.

Iteration 3: Use the ~80,000 tasks generated in Iterations 1-2 as seed tasks, prompt GPT-4-Turbo to generate additional tasks.

Contributors

davanstrien

1 commits