CodeSearchNet Challenge, extended: every search over every function
0
3 commits
2 linked in READMEs
updated Sep 28, 2026
The CodeSearchNet Challenge (Husain et al., 2019) has 99 natural-language code searches, and experts rated a few candidate functions for each. This dataset treats every rated function of a language as one codebase and searches all of it: for each search, every function in its language is a candidate. The experts' ratings are kept, and the pairs they never rated but a search tool returned were rated on the same scale by a judge model, validated against the experts first. Random samples of the pairs only a loose any-keyword grep returned, and of the pairs nothing returned, were rated too.
| Language | Searches | Functions | Rated by experts | Rated by the judge | Relevant |
|---|---|---|---|---|---|
| Python | 99 | 943 | 956 | 3,857 | 787 |
| Java | 99 | 758 | 769 | 3,561 | 512 |
| JavaScript | 96 | 304 | 305 | 3,004 | 247 |
| PHP | 99 | 288 | 289 | 2,976 | 232 |
| Ruby | 97 | 292 | 300 | 2,814 | 189 |
| Go | 83 | 161 | 162 | 2,133 | 79 |
It was made for the jevpipe code search benchmark, which describes the method, the judge's validation and the results in full.
queries: language, search_id, search (the text), split (tuning for the 20 Python
searches used to tune the benchmark's tools, test for all others).corpus: one row per function: id, language, code (the rated lines at the rated commit),
url, repository, commit, path, first_line, last_line, license (the repository's
SPDX license id as GitHub reported it in September 2026, NOASSERTION or null) and sha256
of code.qrels: one row per rated search-function pair: relevance (the experts' mean rating, else the
judge's rating, 0 to 3), relevant (relevance of 2 or more), source (expert or judge),
expert_ratings, judge_rating (also given for expert-rated pairs, to check the judge), and
rated_because (expert, pooled: a search tool returned it, or sample: one of the random
samples above).GPT-6 Astra through the Codex CLI, low reasoning effort, no tools, in batches of 40 pairs mixed across searches, seeing an opaque id, the search and the code (cut at 20,000 characters), never which tool returned the pair. The prompt gave the Challenge's 0-3 scale and its description of each level. Before any of its ratings were used, it rated every expert-rated pair of the language; the gate, set in advance, was an F1 of 0.67 on "relevant" against the experts:
| Language | Expert-rated pairs | Same verdict as the experts | F1 | Gate |
|---|---|---|---|---|
| Python | 956 | 76% | 0.745 | passed |
| Java | 769 | 79% | 0.719 | passed |
| JavaScript | 305 | 83% | 0.781 | passed |
| PHP | 289 | 79% | 0.760 | passed |
| Ruby | 300 | 83% | 0.709 | passed |
| Go | 162 | 80% | 0.629 | missed |
Where two experts rated the same pair (848 Python pairs), they agree on 70% of verdicts, F1 0.711. In Go the judge missed the gate on only 40 relevant expert pairs, and calls more pairs relevant (49) than its experts did (40); use its ratings there with that in mind.
Pooled pairs came from four searches: grep with a pattern a coding agent wrote per search, grep for all of the search's keywords, jevpipe with a probability of 0.3 or more, and DeepSeek V4.1 Flash answering yes or with a probability of 0.3 or more.
qrels as
unjudged, use a metric that allows for that, or rate the new pairs the same way.The functions come from the CodeSearchNet corpus, which kept only projects whose license allows
redistributing parts of them. Each function in corpus stays under its repository's license,
named in license and traceable through url. NOASSERTION means the repository has a license
file that GitHub cannot match to a standard license; read it in the repository. null means GitHub
finds no license in the repository today; CodeSearchNet selected the project in 2019 because its
license then permitted redistribution, and that license still covers the code at the rated commit
that url points to. The Challenge's ratings are from its MIT-licensed repository. The searches'
split, the judge ratings and the dataset's structure are released under the MIT license.
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv:1909.09436.
CodeSearchNet Challenge, extended: every search over every function
0
3 commits
2 linked in READMEs
updated Sep 28, 2026
The CodeSearchNet Challenge (Husain et al., 2019) has 99 natural-language code searches, and experts rated a few candidate functions for each. This dataset treats every rated function of a language as one codebase and searches all of it: for each search, every function in its language is a candidate. The experts' ratings are kept, and the pairs they never rated but a search tool returned were rated on the same scale by a judge model, validated against the experts first. Random samples of the pairs only a loose any-keyword grep returned, and of the pairs nothing returned, were rated too.
| Language | Searches | Functions | Rated by experts | Rated by the judge | Relevant |
|---|---|---|---|---|---|
| Python | 99 | 943 | 956 | 3,857 | 787 |
| Java | 99 | 758 | 769 | 3,561 | 512 |
| JavaScript | 96 | 304 | 305 | 3,004 | 247 |
| PHP | 99 | 288 | 289 | 2,976 | 232 |
| Ruby | 97 | 292 | 300 | 2,814 | 189 |
| Go | 83 | 161 | 162 | 2,133 | 79 |
It was made for the jevpipe code search benchmark, which describes the method, the judge's validation and the results in full.
queries: language, search_id, search (the text), split (tuning for the 20 Python
searches used to tune the benchmark's tools, test for all others).corpus: one row per function: id, language, code (the rated lines at the rated commit),
url, repository, commit, path, first_line, last_line, license (the repository's
SPDX license id as GitHub reported it in September 2026, NOASSERTION or null) and sha256
of code.qrels: one row per rated search-function pair: relevance (the experts' mean rating, else the
judge's rating, 0 to 3), relevant (relevance of 2 or more), source (expert or judge),
expert_ratings, judge_rating (also given for expert-rated pairs, to check the judge), and
rated_because (expert, pooled: a search tool returned it, or sample: one of the random
samples above).GPT-6 Astra through the Codex CLI, low reasoning effort, no tools, in batches of 40 pairs mixed across searches, seeing an opaque id, the search and the code (cut at 20,000 characters), never which tool returned the pair. The prompt gave the Challenge's 0-3 scale and its description of each level. Before any of its ratings were used, it rated every expert-rated pair of the language; the gate, set in advance, was an F1 of 0.67 on "relevant" against the experts:
| Language | Expert-rated pairs | Same verdict as the experts | F1 | Gate |
|---|---|---|---|---|
| Python | 956 | 76% | 0.745 | passed |
| Java | 769 | 79% | 0.719 | passed |
| JavaScript | 305 | 83% | 0.781 | passed |
| PHP | 289 | 79% | 0.760 | passed |
| Ruby | 300 | 83% | 0.709 | passed |
| Go | 162 | 80% | 0.629 | missed |
Where two experts rated the same pair (848 Python pairs), they agree on 70% of verdicts, F1 0.711. In Go the judge missed the gate on only 40 relevant expert pairs, and calls more pairs relevant (49) than its experts did (40); use its ratings there with that in mind.
Pooled pairs came from four searches: grep with a pattern a coding agent wrote per search, grep for all of the search's keywords, jevpipe with a probability of 0.3 or more, and DeepSeek V4.1 Flash answering yes or with a probability of 0.3 or more.
qrels as
unjudged, use a metric that allows for that, or rate the new pairs the same way.The functions come from the CodeSearchNet corpus, which kept only projects whose license allows
redistributing parts of them. Each function in corpus stays under its repository's license,
named in license and traceable through url. NOASSERTION means the repository has a license
file that GitHub cannot match to a standard license; read it in the repository. null means GitHub
finds no license in the repository today; CodeSearchNet selected the project in 2019 because its
license then permitted redistribution, and that license still covers the code at the rated commit
that url points to. The Challenge's ratings are from its MIT-licensed repository. The searches'
split, the judge ratings and the dataset's structure are released under the MIT license.
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv:1909.09436.