datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-copilot-training-from-many-repos-large
Python Copilot Large Coding Dataset
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more.
Rows: 2350782… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-from-many-repos-large.from-one-to-many-toxicity-mitigation
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
[arxiv][code][data]
Data accompanying the paper "From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models" accepted to ACL Findings 2024.
Abstract: To date, toxicity mitigation in language models has almost entirely been focused on single-language settings. As language models embrace multilingual capabilities, it’s crucial our safety measures keep pace. Recognizing this research… See the full description on the dataset page: https://huggingface.co/datasets/luizapzbn/from-one-to-many-toxicity-mitigation.task101_reverse_and_concatenate_all_elements_from_index_i_to_j
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task101_reverse_and_concatenate_all_elements_from_index_i_to_j
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task101_reverse_and_concatenate_all_elements_from_index_i_to_j.task1326_qa_zre_question_generation_from_answer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1326_qa_zre_question_generation_from_answer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1326_qa_zre_question_generation_from_answer.text-to-ocl-from-ecore
Introduction
This is a small size dataset containing 52 meta-models (EMF files and PlantUML descriptions), 369 OCL constraints and 369 constraint specification in natural language.
The meta-models and OCL constraints are collected from open source github projects and are (syntactically) processable by Eclipse.
The constraint specifications of OCL constraints are generated via GPT-4-Turbo.
The meta-models can be found in models\
Usage
Generation of OCL constraints based on… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.task1551_every_ith_element_from_kth_element
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1551_every_ith_element_from_kth_element
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1551_every_ith_element_from_kth_element.task267_concatenate_and_reverse_all_elements_from_index_i_to_j
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task267_concatenate_and_reverse_all_elements_from_index_i_to_j
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task267_concatenate_and_reverse_all_elements_from_index_i_to_j.rag-human-rights-from-files
Dataset Card for my-distiset-rag-files
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.task488_extract_all_alphabetical_elements_from_list_in_order
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task488_extract_all_alphabetical_elements_from_list_in_order
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task488_extract_all_alphabetical_elements_from_list_in_order.wikipedia_2003
Wikipedia-2003
Original dump: https://dumps.wikimedia.org/archive/2003/2003-05-16
This is a filtered and cleaned version of the 2003 Wikipedia dump.
Stats
Language
Size
Lines
Bosnian (bs)
77.6KB
78
Czech (cs)
392.8KB
354
Danish (da)
4.9MB
11,561
German (de)
23.47MB
18,490
English (en)
249MB
128,198
Esperanto (eo)
7.9MB
7,202
Spanish (es)
7.33MB
4,651
French (fr)
13.2MB
10,957
Croatian (hr)
1.2KB
3
Dutch (nl)
10.9MB
7,116
Polish (pl)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/wikipedia_2003.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
task497_extract_all_numbers_from_list_in_order
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task497_extract_all_numbers_from_list_in_order
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task497_extract_all_numbers_from_list_in_order.task1328_qa_zre_relation_generation_from_question
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1328_qa_zre_relation_generation_from_question
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1328_qa_zre_relation_generation_from_question.text-to-xmi-from-ecoreThis is a small test set for XMI instance model generation task.
It containing 26 pairs of meta-models (Ecore), specifications (natural language) and instance models (XMI).
In each pair, the meta-model and instance model share the same name. To proper open the instance model in Eclipse EMF, the instance model and meta-model should be placed in the same folder.
The meta-models are selected from https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.
The specifications are generated via… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-xmi-from-ecore.task499_extract_and_add_all_numbers_from_list
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task499_extract_and_add_all_numbers_from_list
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task499_extract_and_add_all_numbers_from_list.py-docs-2004
Python Docs 2004
Original dump: https://www.python.org/ftp/python/doc/
Python Docs 2004 is a filtered and cleaned collection of Python documentation from every major Python release published before 2004.
Stats
Version
Size
Lines
2.3
2.2MB
1215
2.2
1.7MB
1142
2.1
1.3MB
891
2.0
1.2MB
895
1.6
1MB
720
1.5
837KB
449
1.4
744KB
397
1.3
569KB
408
1.2
513KB
384
Total
10.1MB
6501
Notice
This dataset is a filtered and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/py-docs-2004.rose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.25k_from_rollouts
25k teacher rollouts from Affine SN120
24,930 prompt–completion pairs distilled from the published duel artifacts of
Affine (Bittensor subnet 120). Each row is one teacher
rollout on one agent turn: the conversation so far, the reasoning the teacher
produced, and the bash action it took.
Built from corpus epoch 5 (manifest 1cd8edc52646, 29,860 turns across 5
shards) and all 104 eval artifacts published up to 2026-08-10.
Fields
field
type
description… See the full description on the dataset page: https://huggingface.co/datasets/vuhaian/25k_from_rollouts.amazon_reviews_multi_fr_prompt_title_generation_from_a_review
amazon_reviews_multi_fr_prompt_title_generation_from_a_review
Summary
amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.arxiv-abstracts-2004
ArXiv Abstracts 2004
Original Dataset: common-pile/arxiv_abstracts
ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004.
Stats
Size (MB)
Lines
351MB
303,761
Note: The lines, in the .jsonl file, are ordered from oldest to newest.
Notice
We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review.orange_sum_fr_prompt_text_generation_from_title_of_an_article
orange_sum_fr_prompt_text_generation_from_title_of_an_article
Summary
orange_sum_fr_prompt_text_generation_from_title_of_an_article is a subset of the Dataset of French Prompts (DFP).It contains 908,793 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_title_of_an_article.GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl
Description
This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset.
This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl:
id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review.turkish-sft-from-scratch-120k
Turkish SFT From Scratch 120K
Sıfırdan üretilmiş, kategori kontrollü Türkçe SFT dataset'i. Eski/temizlenmiş datasetlerden satır kopyalanmadı.
Kapsam
12 kategori x 10,000 örnek = 120,000 örnek:
instruction-following
qa
summarization
cot
multi-turn-dialogue
rewriting
text-classification
error-correction
formal-writing
translation
code-explanation
creative-writing
Doğrulamalar
Canonical messages formatı: system/user/assistant.
Exact duplicate hash kontrolü.… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-from-scratch-120k.RLVE-Qwen3-1.7B-Pass1-Rollouts
RLVE teacher rollouts — Qwen3-1.7B (pass@1)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-1.7B
Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym
environments (counting / combinatorics / optimization tasks)
Sampling: 1 sample/question (pass@1) = 9000 records,
temperature 0.7, max 4096 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.rag-human-rights-from-prompt
Dataset Card for datset-rag-prompt
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/datset-rag-prompt/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-prompt.turkish-sft-from-scratch-150k-extended
Turkish SFT From Scratch 150K Extended
kilicai/turkish-sft-from-scratch-120k üzerine 30K akıl yürütme, görev takibi ve analiz verisi eklenmiş genişletilmiş sürüm.
Audit
{
"rows": 150000,
"base_rows": 120000,
"extension_rows": 30000,
"duplicates_removed_on_merge": 0,
"categories": {
"formal-writing": 10000,
"rewriting": 10000,
"text-classification": 10000,
"instruction-following": 10000,
"translation": 10000,
"cot": 10000… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-from-scratch-150k-extended.code-eval-pass8-rollouts
Code eval (pass@8) — Qwen3 code-SFT comparison
Inference-time pass@8 rollouts on the code test split for 4 models, sampled
with eval_code_array.sbatch.
Source prompts: CL-From-Nothing/code_hard test split — 408 competitive-programming questions
Sampling: 8 samples/question (pass@8) = 3264 records/model, temperature 0.7, max_model_len 32000. Main runs use 32768 max new tokens; the base model also has a supplementary 16384-token run.
Rewards: DeepCoder code verifier — 1.0 if the… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code-eval-pass8-rollouts.
