Scoolar/codesearchnet-challenge-extended

Dataset

CodeSearchNet Challenge, extended: every search over every function

0

3 commits

2 linked in READMEs

updated Sep 28, 2026

See the code

README

CodeSearchNet Challenge, extended: every search over every function

The CodeSearchNet Challenge (Husain et al., 2019) has 99 natural-language code searches, and experts rated a few candidate functions for each. This dataset treats every rated function of a language as one codebase and searches all of it: for each search, every function in its language is a candidate. The experts' ratings are kept, and the pairs they never rated but a search tool returned were rated on the same scale by a judge model, validated against the experts first. Random samples of the pairs only a loose any-keyword grep returned, and of the pairs nothing returned, were rated too.

LanguageSearchesFunctionsRated by expertsRated by the judgeRelevant
Python999439563,857787
Java997587693,561512
JavaScript963043053,004247
PHP992882892,976232
Ruby972923002,814189
Go831611622,13379

It was made for the jevpipe code search benchmark, which describes the method, the judge's validation and the results in full.

Configurations

  • queries: language, search_id, search (the text), split (tuning for the 20 Python searches used to tune the benchmark's tools, test for all others).
  • corpus: one row per function: id, language, code (the rated lines at the rated commit), url, repository, commit, path, first_line, last_line, license (the repository's SPDX license id as GitHub reported it in September 2026, NOASSERTION or null) and sha256 of code.
  • qrels: one row per rated search-function pair: relevance (the experts' mean rating, else the judge's rating, 0 to 3), relevant (relevance of 2 or more), source (expert or judge), expert_ratings, judge_rating (also given for expert-rated pairs, to check the judge), and rated_because (expert, pooled: a search tool returned it, or sample: one of the random samples above).

How the judge rated

GPT-6 Astra through the Codex CLI, low reasoning effort, no tools, in batches of 40 pairs mixed across searches, seeing an opaque id, the search and the code (cut at 20,000 characters), never which tool returned the pair. The prompt gave the Challenge's 0-3 scale and its description of each level. Before any of its ratings were used, it rated every expert-rated pair of the language; the gate, set in advance, was an F1 of 0.67 on "relevant" against the experts:

LanguageExpert-rated pairsSame verdict as the expertsF1Gate
Python95676%0.745passed
Java76979%0.719passed
JavaScript30583%0.781passed
PHP28979%0.760passed
Ruby30083%0.709passed
Go16280%0.629missed

Where two experts rated the same pair (848 Python pairs), they agree on 70% of verdicts, F1 0.711. In Go the judge missed the gate on only 40 relevant expert pairs, and calls more pairs relevant (49) than its experts did (40); use its ratings there with that in mind.

Pooled pairs came from four searches: grep with a pattern a coding agent wrote per search, grep for all of the search's keywords, jevpipe with a probability of 0.3 or more, and DeepSeek V4.1 Flash answering yes or with a probability of 0.3 or more.

Use it with care

  • Unrated is unknown, not irrelevant. The judged pairs were found by the three tools above; a new tool will return relevant functions nobody rated. Treat pairs missing from qrels as unjudged, use a metric that allows for that, or rate the new pairs the same way.
  • Judge ratings are model output, checked against the experts but not expert ratings.
  • Relevance is from 2019 searches on public code that may be in any model's training data.

Licenses

The functions come from the CodeSearchNet corpus, which kept only projects whose license allows redistributing parts of them. Each function in corpus stays under its repository's license, named in license and traceable through url. NOASSERTION means the repository has a license file that GitHub cannot match to a standard license; read it in the repository. null means GitHub finds no license in the repository today; CodeSearchNet selected the project in 2019 because its license then permitted redistribution, and that license still covers the code at the rated commit that url points to. The Challenge's ratings are from its MIT-licensed repository. The searches' split, the judge ratings and the dataset's structure are released under the MIT license.

Citation

Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv:1909.09436.

code-search
llm-as-a-judge
relevance

Scoolar/codesearchnet-challenge-extended

Dataset

CodeSearchNet Challenge, extended: every search over every function

0

3 commits

2 linked in READMEs

updated Sep 28, 2026

See the code

README

CodeSearchNet Challenge, extended: every search over every function

The CodeSearchNet Challenge (Husain et al., 2019) has 99 natural-language code searches, and experts rated a few candidate functions for each. This dataset treats every rated function of a language as one codebase and searches all of it: for each search, every function in its language is a candidate. The experts' ratings are kept, and the pairs they never rated but a search tool returned were rated on the same scale by a judge model, validated against the experts first. Random samples of the pairs only a loose any-keyword grep returned, and of the pairs nothing returned, were rated too.

LanguageSearchesFunctionsRated by expertsRated by the judgeRelevant
Python999439563,857787
Java997587693,561512
JavaScript963043053,004247
PHP992882892,976232
Ruby972923002,814189
Go831611622,13379

It was made for the jevpipe code search benchmark, which describes the method, the judge's validation and the results in full.

Configurations

  • queries: language, search_id, search (the text), split (tuning for the 20 Python searches used to tune the benchmark's tools, test for all others).
  • corpus: one row per function: id, language, code (the rated lines at the rated commit), url, repository, commit, path, first_line, last_line, license (the repository's SPDX license id as GitHub reported it in September 2026, NOASSERTION or null) and sha256 of code.
  • qrels: one row per rated search-function pair: relevance (the experts' mean rating, else the judge's rating, 0 to 3), relevant (relevance of 2 or more), source (expert or judge), expert_ratings, judge_rating (also given for expert-rated pairs, to check the judge), and rated_because (expert, pooled: a search tool returned it, or sample: one of the random samples above).

How the judge rated

GPT-6 Astra through the Codex CLI, low reasoning effort, no tools, in batches of 40 pairs mixed across searches, seeing an opaque id, the search and the code (cut at 20,000 characters), never which tool returned the pair. The prompt gave the Challenge's 0-3 scale and its description of each level. Before any of its ratings were used, it rated every expert-rated pair of the language; the gate, set in advance, was an F1 of 0.67 on "relevant" against the experts:

LanguageExpert-rated pairsSame verdict as the expertsF1Gate
Python95676%0.745passed
Java76979%0.719passed
JavaScript30583%0.781passed
PHP28979%0.760passed
Ruby30083%0.709passed
Go16280%0.629missed

Where two experts rated the same pair (848 Python pairs), they agree on 70% of verdicts, F1 0.711. In Go the judge missed the gate on only 40 relevant expert pairs, and calls more pairs relevant (49) than its experts did (40); use its ratings there with that in mind.

Pooled pairs came from four searches: grep with a pattern a coding agent wrote per search, grep for all of the search's keywords, jevpipe with a probability of 0.3 or more, and DeepSeek V4.1 Flash answering yes or with a probability of 0.3 or more.

Use it with care

  • Unrated is unknown, not irrelevant. The judged pairs were found by the three tools above; a new tool will return relevant functions nobody rated. Treat pairs missing from qrels as unjudged, use a metric that allows for that, or rate the new pairs the same way.
  • Judge ratings are model output, checked against the experts but not expert ratings.
  • Relevance is from 2019 searches on public code that may be in any model's training data.

Licenses

The functions come from the CodeSearchNet corpus, which kept only projects whose license allows redistributing parts of them. Each function in corpus stays under its repository's license, named in license and traceable through url. NOASSERTION means the repository has a license file that GitHub cannot match to a standard license; read it in the repository. null means GitHub finds no license in the repository today; CodeSearchNet selected the project in 2019 because its license then permitted redistribution, and that license still covers the code at the rated commit that url points to. The Challenge's ratings are from its MIT-licensed repository. The searches' split, the judge ratings and the dataset's structure are released under the MIT license.

Citation

Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv:1909.09436.

code-search
llm-as-a-judge
relevance