datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.big-code-2-metadataA subset of big-code-2 containing only the files:
Not in big-code-1.
Under a permissive license.
code-execution-llmamazon-reviews-for-llm-extended
Cross-domain sequential recommendation dataset
A sequential recommendation dataset drawn from Amazon Reviews 2023, covering
Books, CDs_and_Vinyl, Movies_and_TV, Video_Games.
Each row of interactions.parquet is one user buying or reviewing one item at one time.
Users are sampled so that every one of them is active in all domains, their
interactions are ordered chronologically and cut into train/valid/test, and each
interaction carries a fixed set of 10 candidate items for ranking… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-for-llm-extended.pde-llm-eval-code-perturbation-dataset
pde-llm-eval-code-perturbation-dataset
Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.amazon-reviews-for-llm
Cross-domain sequential recommendation dataset
A sequential recommendation dataset drawn from Amazon Reviews 2023, covering
Books, CDs_and_Vinyl, Movies_and_TV, Video_Games.
Each row of interactions.parquet is one user buying or reviewing one item at one time.
Users are sampled so that every one of them is active in all domains, their
interactions are ordered chronologically and cut into train/valid/test, and each
interaction carries a fixed set of 10 candidate items for ranking… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-for-llm.amazon-reviews-books-for-llm
Books sequential recommendation dataset
A sequential recommendation dataset drawn from Amazon Reviews 2023, covering
Books.
Each row of interactions.parquet is one user buying or reviewing one item at one time.
Users are sampled so that every one of them is active in all domains, their
interactions are ordered chronologically and cut into train/valid/test, and each
interaction carries a fixed set of 10 candidate items for ranking
evaluation. Integer user_idx / item_idx columns… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-books-for-llm.llm-metric-ace-code-pairwisellm-metric-ace-code-pairwise-newd1_code_mc_llmcode-corpus-llm-training
Code Corpus for LLM Training
Manually collected from top open-source repositories across:
video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python,
ethical hacking, system-level, game engines, web frameworks, and more.
Stats
Records: 240,378
Raw text: 2,156,908,643 chars (~2.01 GB)
Domains: 20
Domains
web_ui: 32,354 records
cpp: 29,792 records
kotlin_android: 19,476 records
ui_ux_design: 19,382 records
rust: 15,440 records
python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.code-ita-dpo-small
Dataset Card for "code-instructions-ita-dpo-small"
More Information needed
d2_code_mc_llmexp_llmve_llm-verifier-code-contestsllm-verifier-code-contests-noblockllm-verifier-code-contestsllm-verifier-code-contests_glm_4.6_tracesd1_code_mc_llm_10kl4-08-code-generation-datafiltered_ds_codellmsd1_code_mc_llm_3kcode-retrieval-hard-negatives-llm-verified-mergedd1_code_mc_llm_10k_eval_636d
mlfoundations-dev/d1_code_mc_llm_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.7
63.5
77.2
28.4
47.0
42.1
38.7
10.5
15.2
AIME24
Average Accuracy: 19.67% ± 1.45%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
13.33%
4
30
3
23.33%
7
30
4
26.67%
8… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_10k_eval_636d.d1_code_mc_llm_eval_636d
mlfoundations-dev/d1_code_mc_llm_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
23.0
63.3
74.6
27.8
44.9
45.1
45.5
15.3
18.9
AIME24
Average Accuracy: 23.00% ± 1.45%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
23.33%
7
30
3
23.33%
7
30
4
20.00%
6
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_eval_636d.d1_code_mc_llm_0.3k_eval_636d
mlfoundations-dev/d1_code_mc_llm_0.3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.0
55.2
71.8
27.2
39.1
37.5
26.8
6.6
6.6
AIME24
Average Accuracy: 16.00% ± 1.40%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
26.67%
8
30
2
16.67%
5
30
3
13.33%
4
30
4
16.67%
5
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_0.3k_eval_636d.d1_code_mc_llm_3k_eval_636d
mlfoundations-dev/d1_code_mc_llm_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
23.3
61.0
75.6
28.4
44.5
42.8
33.1
9.5
12.4
AIME24
Average Accuracy: 23.33% ± 1.49%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
26.67%
8
30
2
30.00%
9
30
3
20.00%
6
30
4
23.33%
7
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_3k_eval_636d.LLMSecEval-Prompts_generated_codeLLM_generated_Amharic_QA_for_Family_Code_of_Ethiopiad1_code_mc_llm_0.3k
