datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embeddings-fine-tuning-filtered-code
Overview
This dataset is composed of high quality code retrieval data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset.
The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code.embeddings-fine-tuning-filtered-code-edit
Overview
This dataset is composed of high quality code-edit retrieval data with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset.
The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code-edit.transformers_code_embeddings_v3transformers_code_embeddings_v2_pooltransformers_code_embeddings
Transformers Code Embeddings
Compact index of function/class definitions from src/transformers/models/**/modeling_*.py for cross-model similarity. Built to help surface reusable code when modularizing models.
Contents
embeddings.safetensors — float32, L2-normalized embeddings shaped [N, D].
code_index_map.json — {int_id: "relative/path/to/modeling_*.py:SymbolName"}.
code_index_tokens.json — {identifier: [sorted_unique_tokens]} for Jaccard.
How these were built… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/transformers_code_embeddings.code-embedding-dataset
Code-to-Doc Embedding Dataset
AI-generated code documentation pairs for training code embedding / retrieval models.
Dataset Description
Each record contains a code anchor (real production code) paired with:
positive: A rich natural-language documentation of what the code does
queries: 4 natural-language search queries a developer might use to find this code
label: A short semantic label (3-8 words)
This dataset is designed for training bi-encoder embedding models (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/archit11/code-embedding-dataset.b2_code_embeddinginstruction_filtering_scale_up_code_base_embedding_filter_meanb2_calc_negative_embeddings_codeinstruction_filtering_embedding_filter_mean_seed_data_code_w_openthoughtsinstruction_filtering_embedding_filter_seed_data_code_w_openthoughtsembedding_codeb2_code_embedding_filter_open_code_reasoninginstruction_filtering_scale_up_code_base_embedding_filter_mean_per_domain_16Kb2_calc_positive_embeddings_code_baai_tacob2_code_embedding_10k_eval_636d
mlfoundations-dev/b2_code_embedding_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.3
50.5
71.0
29.6
27.8
28.3
35.9
8.4
9.0
AIME24
Average Accuracy: 19.33% ± 1.23%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
13.33%
4
30
3
20.00%
6
30
4
13.33%
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_embedding_10k_eval_636d.b2_code_embedding_0.3k_eval_636d
mlfoundations-dev/b2_code_embedding_0.3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
18.7
55.8
74.8
26.6
42.1
34.8
28.2
7.3
8.7
AIME24
Average Accuracy: 18.67% ± 1.84%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
13.33%
4
30
3
16.67%
5
30
4
20.00%
6… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_embedding_0.3k_eval_636d.instruction_filtering_scale_up_code_base_embedding_filter_mean_per_domaininstruction_filtering_scale_up_code_base_embedding_filter_mean_8Kb2_calc_positive_embeddings_code_code_golfb2_calc_negative_embeddings_code_stackexchange_codereviewb2_code_embedding_10kb2_code_embedding_1k_eval_636d
mlfoundations-dev/b2_code_embedding_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
12.3
39.2
67.8
26.8
33.9
31.8
19.0
4.3
5.8
AIME24
Average Accuracy: 12.33% ± 2.16%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30
2
3.33%
1
30
3
16.67%
5
30
4
6.67%
2
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_embedding_1k_eval_636d.b2_code_embedding_3k_eval_636d
mlfoundations-dev/b2_code_embedding_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
21.7
62.5
75.2
28.8
43.8
40.4
34.7
9.4
12.5
AIME24
Average Accuracy: 21.67% ± 1.84%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
33.33%
10
30
3
16.67%
5
30
4
20.00%
6… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_embedding_3k_eval_636d.instruction_filtering_embedding_filter_codeinstruction_filtering_embedding_filter_code_meaninstruction_filtering_scale_up_code_base_embedding_filter_mean_16Kinstruction_filtering_scale_up_code_base_embedding_filter_mean_per_domain_2Kinstruction_filtering_scale_up_code_base_embedding_filter_mean_per_domain_8Kb2_code_embedding_filter_code_golf
