CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01matlok /python-copilot-training-from-many-repos-large Python Copilot Large Coding Dataset This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more. Rows: 2350782… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-from-many-repos-large.tabulartext-generation10K<n<100K1 likes384 downloads3y agoHugging Face02luizapzbn /from-one-to-many-toxicity-mitigation From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models [arxiv][code][data] Data accompanying the paper "From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models" accepted to ACL Findings 2024. Abstract: To date, toxicity mitigation in language models has almost entirely been focused on single-language settings. As language models embrace multilingual capabilities, it’s crucial our safety measures keep pace. Recognizing this research… See the full description on the dataset page: https://huggingface.co/datasets/luizapzbn/from-one-to-many-toxicity-mitigation.texttext-generation0 likes224 downloads2y agoHugging Face03Lots-of-LoRAs /task101_reverse_and_concatenate_all_elements_from_index_i_to_j Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task101_reverse_and_concatenate_all_elements_from_index_i_to_j Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task101_reverse_and_concatenate_all_elements_from_index_i_to_j.texttext-generation1K<n<10K0 likes152 downloads2y agoHugging Face04Lots-of-LoRAs /task1326_qa_zre_question_generation_from_answer Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1326_qa_zre_question_generation_from_answer Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1326_qa_zre_question_generation_from_answer.texttext-generation1K<n<10K0 likes95 downloads2y agoHugging Face05fpan /text-to-ocl-from-ecore Introduction This is a small size dataset containing 52 meta-models (EMF files and PlantUML descriptions), 369 OCL constraints and 369 constraint specification in natural language. The meta-models and OCL constraints are collected from open source github projects and are (syntactically) processable by Eclipse. The constraint specifications of OCL constraints are generated via GPT-4-Turbo. The meta-models can be found in models\ Usage Generation of OCL constraints based on… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.texttranslationn<1K0 likes93 downloads2y agoHugging Face06Lots-of-LoRAs /task1551_every_ith_element_from_kth_element Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1551_every_ith_element_from_kth_element Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1551_every_ith_element_from_kth_element.texttext-generation1K<n<10K0 likes93 downloads2y agoHugging Face07Lots-of-LoRAs /task267_concatenate_and_reverse_all_elements_from_index_i_to_j Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task267_concatenate_and_reverse_all_elements_from_index_i_to_j Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task267_concatenate_and_reverse_all_elements_from_index_i_to_j.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face08sdiazlor /rag-human-rights-from-files Dataset Card for my-distiset-rag-files This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.texttext-generationn<1K0 likes88 downloads2y agoHugging Face09Lots-of-LoRAs /task488_extract_all_alphabetical_elements_from_list_in_order Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task488_extract_all_alphabetical_elements_from_list_in_order Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task488_extract_all_alphabetical_elements_from_list_in_order.texttext-generation1K<n<10K0 likes85 downloads2y agoHugging Face10fromziro /wikipedia_2003 Wikipedia-2003 Original dump: https://dumps.wikimedia.org/archive/2003/2003-05-16 This is a filtered and cleaned version of the 2003 Wikipedia dump. Stats Language Size Lines Bosnian (bs) 77.6KB 78 Czech (cs) 392.8KB 354 Danish (da) 4.9MB 11,561 German (de) 23.47MB 18,490 English (en) 249MB 128,198 Esperanto (eo) 7.9MB 7,202 Spanish (es) 7.33MB 4,651 French (fr) 13.2MB 10,957 Croatian (hr) 1.2KB 3 Dutch (nl) 10.9MB 7,116 Polish (pl)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/wikipedia_2003.tabulartext-generation100K<n<1M0 likes78 downloads2mo agoHugging Face11philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes76 downloads1y agoHugging Face12Lots-of-LoRAs /task497_extract_all_numbers_from_list_in_order Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task497_extract_all_numbers_from_list_in_order Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task497_extract_all_numbers_from_list_in_order.texttext-generation1K<n<10K0 likes71 downloads2y agoHugging Face13Lots-of-LoRAs /task1328_qa_zre_relation_generation_from_question Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1328_qa_zre_relation_generation_from_question Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1328_qa_zre_relation_generation_from_question.texttext-generation1K<n<10K0 likes71 downloads2y agoHugging Face14fpan /text-to-xmi-from-ecoreThis is a small test set for XMI instance model generation task. It containing 26 pairs of meta-models (Ecore), specifications (natural language) and instance models (XMI). In each pair, the meta-model and instance model share the same name. To proper open the instance model in Eclipse EMF, the instance model and meta-model should be placed in the same folder. The meta-models are selected from https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore. The specifications are generated via… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-xmi-from-ecore.texttext-generationn<1K0 likes71 downloads1y agoHugging Face15Lots-of-LoRAs /task499_extract_and_add_all_numbers_from_list Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task499_extract_and_add_all_numbers_from_list Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task499_extract_and_add_all_numbers_from_list.texttext-generation1K<n<10K0 likes68 downloads2y agoHugging Face16fromziro /py-docs-2004 Python Docs 2004 Original dump: https://www.python.org/ftp/python/doc/ Python Docs 2004 is a filtered and cleaned collection of Python documentation from every major Python release published before 2004. Stats Version Size Lines 2.3 2.2MB 1215 2.2 1.7MB 1142 2.1 1.3MB 891 2.0 1.2MB 895 1.6 1MB 720 1.5 837KB 449 1.4 744KB 397 1.3 569KB 408 1.2 513KB 384 Total 10.1MB 6501 Notice This dataset is a filtered and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/py-docs-2004.texttext-generation10K<n<100K0 likes64 downloads2mo agoHugging Face17CL-From-Nothing /rose_code_samples rose_code samples (pass@8 rollouts) vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass). Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines. Qwen3-4B-Thinking-2507/ — teacher model rollouts. Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8). Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.tabulartext-generation100K<n<1M0 likes55 downloads4mo agoHugging Face18SeanWang0027 /polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507 Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's actual continuation for each, and the token accounting behind it. The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is decoded to text, the teacher is shown it under its own chat template, and the teacher's reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.tabulartext-generation10K<n<100K0 likes55 downloads26d agoHugging Face19vuhaian /25k_from_rollouts 25k teacher rollouts from Affine SN120 24,930 prompt–completion pairs distilled from the published duel artifacts of Affine (Bittensor subnet 120). Each row is one teacher rollout on one agent turn: the conversation so far, the reasoning the teacher produced, and the bash action it took. Built from corpus epoch 5 (manifest 1cd8edc52646, 29,860 turns across 5 shards) and all 104 eval artifacts published up to 2026-08-10. Fields field type description… See the full description on the dataset page: https://huggingface.co/datasets/vuhaian/25k_from_rollouts.texttext-generation10K<n<100K0 likes54 downloads2mo agoHugging Face20CATIE-AQ /amazon_reviews_multi_fr_prompt_title_generation_from_a_review amazon_reviews_multi_fr_prompt_title_generation_from_a_review Summary amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.texttext-generation1M<n<10M0 likes52 downloads1y agoHugging Face21fromziro /arxiv-abstracts-2004 ArXiv Abstracts 2004 Original Dataset: common-pile/arxiv_abstracts ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004. Stats Size (MB) Lines 351MB 303,761 Note: The lines, in the .jsonl file, are ordered from oldest to newest. Notice We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.texttext-generation100K<n<1M1 likes47 downloads2mo agoHugging Face22CATIE-AQ /amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review Summary amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review.texttext-generation1M<n<10M0 likes46 downloads1y agoHugging Face23CATIE-AQ /orange_sum_fr_prompt_text_generation_from_title_of_an_article orange_sum_fr_prompt_text_generation_from_title_of_an_article Summary orange_sum_fr_prompt_text_generation_from_title_of_an_article is a subset of the Dataset of French Prompts (DFP).It contains 908,793 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_title_of_an_article.texttext-generation100K<n<1M0 likes38 downloads1y agoHugging Face24GeoGPT-Research-Project /GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl Description This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset. This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl: id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.texttext-generation10M<n<100M1 likes37 downloads1y agoHugging Face25CATIE-AQ /amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review Summary amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review.texttext-generation1M<n<10M0 likes35 downloads1y agoHugging Face26kilicai /turkish-sft-from-scratch-120k Turkish SFT From Scratch 120K Sıfırdan üretilmiş, kategori kontrollü Türkçe SFT dataset'i. Eski/temizlenmiş datasetlerden satır kopyalanmadı. Kapsam 12 kategori x 10,000 örnek = 120,000 örnek: instruction-following qa summarization cot multi-turn-dialogue rewriting text-classification error-correction formal-writing translation code-explanation creative-writing Doğrulamalar Canonical messages formatı: system/user/assistant. Exact duplicate hash kontrolü.… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-from-scratch-120k.texttext-generation100K<n<1M0 likes35 downloads4mo agoHugging Face27CL-From-Nothing /RLVE-Qwen3-1.7B-Pass1-Rollouts RLVE teacher rollouts — Qwen3-1.7B (pass@1) Teacher rollouts for on-policy distillation on the RLVE environment suite. Teacher / sampler: Qwen3-1.7B Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym environments (counting / combinatorics / optimization tasks) Sampling: 1 sample/question (pass@1) = 9000 records, temperature 0.7, max 4096 new tokens Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]). Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.tabulartext-generation1K<n<10K0 likes35 downloads4mo agoHugging Face28sdiazlor /rag-human-rights-from-prompt Dataset Card for datset-rag-prompt This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/datset-rag-prompt/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-prompt.texttext-generationn<1K0 likes30 downloads2y agoHugging Face29kilicai /turkish-sft-from-scratch-150k-extended Turkish SFT From Scratch 150K Extended kilicai/turkish-sft-from-scratch-120k üzerine 30K akıl yürütme, görev takibi ve analiz verisi eklenmiş genişletilmiş sürüm. Audit { "rows": 150000, "base_rows": 120000, "extension_rows": 30000, "duplicates_removed_on_merge": 0, "categories": { "formal-writing": 10000, "rewriting": 10000, "text-classification": 10000, "instruction-following": 10000, "translation": 10000, "cot": 10000… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-from-scratch-150k-extended.texttext-generation100K<n<1M0 likes28 downloads4mo agoHugging Face30CL-From-Nothing /code-eval-pass8-rollouts Code eval (pass@8) — Qwen3 code-SFT comparison Inference-time pass@8 rollouts on the code test split for 4 models, sampled with eval_code_array.sbatch. Source prompts: CL-From-Nothing/code_hard test split — 408 competitive-programming questions Sampling: 8 samples/question (pass@8) = 3264 records/model, temperature 0.7, max_model_len 32000. Main runs use 32768 max new tokens; the base model also has a supplementary 16384-token run. Rewards: DeepCoder code verifier — 1.0 if the… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code-eval-pass8-rollouts.tabulartext-generation10K<n<100K0 likes28 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.