CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01barc0 /200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds. We generate the dataset with the following steps and two approaches: Generate ~110k descriptions by GPT4o. Approach 1: Generate ~110k codes follow each description by GPT4o-mini. Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions. Run the ~220k codes and do auto-filtering. Get the final ~200k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M11 likes652 downloads2y agoHugging Face02Lots-of-LoRAs /task1729_personachat_generate_next Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1729_personachat_generate_next Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1729_personachat_generate_next.texttext-generation1K<n<10K0 likes243 downloads2y agoHugging Face03artfultom /ivypanda-llm-generated-essays AI-Generated Essays Dataset This dataset contains AI-generated academic essays created using the models: Mistral 7B Instruct v0.2 (Q5_K_M quantized) Temperature: 0.7 Max tokens: 4096 Top-p: 0.9 (default) Top-k: 40 (default) Repeat penalty: 1.1 (default) Context window: 32768 tokens Llama 3 13B Instruct v0.1 (Q5_K_M quantized) Temperature: 0.7 Max tokens: 4096 Top-p: 0.9 (default) Top-k: 40 (default) Repeat penalty: 1.1 (default) Context window: 8192 tokens DeepSeek-V3.2 API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.texttext-classification10K<n<100K1 likes177 downloads8mo agoHugging Face04barc0 /100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4o-mini. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes160 downloads2y agoHugging Face05barc0 /100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes128 downloads2y agoHugging Face06Brainquiver /generate-narrate-tinystories-pretrain Narrative · TinyStories · Pretraining (Cleaned) Microsoft's TinyStories V2, cleaned and stored as parquet. 2,745,100 stories, 441 million words, one story per row with provenance on every record. Composition Config Records % Source all 2,745,100 100.00 the single config (default) gpt-4 2,745,100 100.00 TinyStoriesV2-GPT4-train TinyStories V2 holds samples generated by GPT-3.5 and samples generated by GPT-4. Only the GPT-4 samples are here… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/generate-narrate-tinystories-pretrain.texttext-generation1M<n<10M1 likes122 downloads24d agoHugging Face07akshay-sked /qwen-generated-svamp-controls-sft Qwen-Generated SVAMP CoT Controls ? SFT Qwen-generated controlled reasoning traces for SVAMP in LLaMA-Factory SFT format. Variants include ordinary, all-caps, no-comma, disclaimer, and multilingual examples. Splits 3,940 training examples and 380 held-out evaluation examples. Format The JSON files use the LLaMA-Factory Alpaca-style schema. The included dataset_info.json registers the exact training and evaluation names. DPO records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-sft.texttext-generation1K<n<10K0 likes106 downloads2mo agoHugging Face08Lots-of-LoRAs /task957_e2e_nlg_text_generation_generate Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task957_e2e_nlg_text_generation_generate Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task957_e2e_nlg_text_generation_generate.texttext-generation1K<n<10K0 likes103 downloads2y agoHugging Face09anonymous-dataset-submission-warp /warp-taskgen-generated-ipi-tasks-50 WARP Taskgen Generated IPI Tasks 50 Dataset Summary This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a 50-task generated indirect prompt injection (IPI) cohort. The trajectories were produced with the AgentLab harness on WebArena GitLab and Postmill (Reddit) benchmark applications. The export is a report-only projection of already written benchmark artifacts. It does not alter scoring, PVPO encounter checks, rewards, admission, or trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.text-generationn<1K0 likes92 downloads5mo agoHugging Face10Nerthus-Project /Generated_OE_Gregory_Dialogues_Text_and_Evaluation Generated Old English Gregory's Dialogues (variatio) A complete, machine-generated Old English variatio of the Old English Dialogues of Gregory the Great (Waerferth's translation), produced on 19 July 2026, together with the full generation and evaluation apparatus: prompt, constraint lexicon scripts, validator, dependency parses, word embeddings, and all quantitative evaluation results. The project is described in: Martin Arista, J., & Nunez, M. Evaluating Generated Old… See the full description on the dataset page: https://huggingface.co/datasets/Nerthus-Project/Generated_OE_Gregory_Dialogues_Text_and_Evaluation.text-generation1K<n<10K0 likes90 downloads17d agoHugging Face11LennardZuendorf /Dynamically-Generated-Hate-Speech-Dataset Dataset Card for dynamically generated hate speech dataset Dataset Summary This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela Original README from GitHub Dynamically-Generated-Hate-Speech-Dataset ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.tabulartext-classification10K<n<100K6 likes80 downloads3y agoHugging Face12Lots-of-LoRAs /task389_torque_generate_temporal_question Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task389_torque_generate_temporal_question Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task389_torque_generate_temporal_question.texttext-generation1K<n<10K0 likes79 downloads2y agoHugging Face13ghananlpcommunity /ghanaian-corpus-generate-clean Ghanaian Corpus Generate — Cleaned Cleaned version of ghananlpcommunity/ghanaian-corpus-generate. What changed The text_clean column was produced by stripping out non-sentence content from text: Removed section/chapter numbering (e.g. 24 5.3, III, 6.1) Removed document headers (e.g. Chapter 6: Conclusion and Recommendations, Abstract, Acknowledgements) Removed figure/table/code references (e.g. Figure 4.10 QR recognition class) Removed very short fragments (< 15… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghanaian-corpus-generate-clean.texttext-generation100K<n<1M0 likes77 downloads2mo agoHugging Face14UniqueData /generated-e-mail-spamThe dataset consists of a **CSV file** containing of 300 generated email spam messages. Each row in the file represents a separate email message, its *title and text.* The dataset aims to facilitate the analysis and detection of spam emails. The dataset can be used for various purposes, such as *training machine learning algorithms to classify and filter spam emails, studying spam email patterns, or analyzing text-based features of spam messages*.text-generation1 likes74 downloads1y agoHugging Face15ychen /Generated-Empathetic-Dialogues-v0.1-Smol Generated Empathetic Conversations v0.1 - Smol This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics. Highlights Multi-round conversation It's not single-turn. The user and the assistant works together to gradually unfold the conversation. The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.texttext-generation10K<n<100K4 likes66 downloads2y agoHugging Face16pymlex /ai-generated-texts Spanish DPO Preference Pairs for Detector Evasion Preference pairs for DPO fine-tuning of Qwen/Qwen2.5-0.5B-Instruct against the Oculus multilingual AI text detector on Spanish academic abstracts. Repository id: pymlex/ai-generated-texts. Dataset size Statistic Count Train abstracts processed 8891 DPO pairs retained 6396 Pairs skipped by logit margin 2495 Empty paraphrase pairs 0 Logit margin threshold: absolute gap at least 1.… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/ai-generated-texts.texttext-generation1K<n<10K0 likes52 downloads3mo agoHugging Face17mobileforge-anonymous /mobileforge-generated-tasks MobileForge Generated Tasks Anonymous project: https://mobileforge-anonymous.github.io/Anonymous code: https://github.com/mobileforge-anonymous/MobileForge This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation. Release inventory payloads: files=1; bytes=2028399;… See the full description on the dataset page: https://huggingface.co/datasets/mobileforge-anonymous/mobileforge-generated-tasks.text-generation1K<n<10K0 likes52 downloads26d agoHugging Face18akshay-sked /qwen-generated-svamp-controls-dpo Qwen-Generated SVAMP CoT Controls ? DPO Preference pairs built from Qwen-generated SVAMP reasoning traces in LLaMA-Factory DPO format. Each record contains instruction, input, chosen, and rejected fields. Splits 3,152 training preference pairs and 304 held-out evaluation pairs. Format The JSON files use the LLaMA-Factory Alpaca-style schema. The included dataset_info.json registers the exact training and evaluation names. DPO records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-dpo.texttext-generation1K<n<10K0 likes47 downloads2mo agoHugging Face19omrisap /ruleloopvit-sft-generated-rules-019200 RuleLoopViT SFT Generated ARC-AGI-1 Rules This dataset contains one generated rule text for each of the 400 ARC-AGI-1 training tasks. The rules were generated by the RuleLoopViT principal SFT language model: omrisap/sft_lm_principal Each task was prompted with 2-4 official ARC-AGI demonstration pairs and decoded deterministically. The generated output follows the project's five-section rule schema: [CORE RULE] [INPUT STRUCTURE] [TARGET SELECTION] [TRANSFORMATION] [OUTPUT… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/ruleloopvit-sft-generated-rules-019200.text-generation0 likes43 downloads18d agoHugging Face20kogai /full-html-stying-dataset-generated-css-from-style-plan Generated CSS From Style Plan kogai/full-html-stying-dataset-generated-css-from-style-plan contains generated_css_from_style_plan.jsonl, a JSONL dataset with 44458 synthetic examples. Model-generated CSS outputs conditioned on source HTML, user style requests, and structured style plans. Schema chat_template_overhead_tokens: field present in the JSONL records. created_at: field present in the JSONL records. input_html: source HTML before Tailwind classes are… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-generated-css-from-style-plan.text-generation10K<n<100K0 likes31 downloads2mo agoHugging Face21paoche11 /diffusion-generated-text Diffusion-Generated Text This dataset contains 16,791 question-response pairs generated by a diffusion language model. It is released to support research on diffusion-generated language and machine-generated text detection. Dataset schema Column Type Description question string Input question or prompt. dLLM_response string Response produced by the diffusion language model. Loading from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/paoche11/diffusion-generated-text.texttext-generation10K<n<100K0 likes29 downloads2mo agoHugging Face22BaSalam /entity-attribute-sft-dataset-GPT-4.0-generated-v1 Entity Attribute Dataset 50k (GPT-4.0 Generated) Dataset Summary The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-sft-dataset-GPT-4.0-generated-v1.texttext-generation10K<n<100K4 likes28 downloads2y agoHugging Face23fibonacciai /entity-attribute-sft-dataset-GPT-4.0-generated-v1 Entity Attribute Dataset 50k (GPT-4.0 Generated) Dataset Summary The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment… See the full description on the dataset page: https://huggingface.co/datasets/fibonacciai/entity-attribute-sft-dataset-GPT-4.0-generated-v1.texttext-generation10K<n<100K0 likes28 downloads4mo agoHugging Face24lgy0404 /mobileforge-generated-tasks MobileForge Generated Tasks This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation. Dataset summary File Rows Apps Size Description generated_tasks_26020301-all.csv 3,249 20 1.93 MB Consolidated AndroidWorld-side MobileForge task pool. The task pool is generated from real target-app… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-generated-tasks.tabulartext-generation1K<n<10K0 likes26 downloads3mo agoHugging Face25science-of-finetuning /ultrachat_200k_generated_gemma-2-2b-itThis dataset contains 512 answers generated by the gemma-2-2b-it model on a subset of the ultrachat 200k test_sft dataset using greedy decoding. The subset was generated by filtering out conversations that were >= 1024 - 128 tokens long, and answers were cut off at each batch after 1024 - min(batch_prompt_lengths) generated tokens, such that each answer is at most 128 tokens long. The generated answers are 200k tokens so 390 tokens (~300 words or 2/3 pages) on average. texttext-generationn<1K0 likes25 downloads2y agoHugging Face26zhengnx /Ai_generate_1texttext-classificationn<1K0 likes25 downloads1y agoHugging Face27isaacchung /controlled-generated-convos-gpt-4.1-mini Controlled Generated Conversations: gpt-4.1-mini Dataset Description This dataset contains synthetic customer support conversations generated using gpt-4.1-mini as part of research on cross-lingual stability of LLM judges. The conversations are designed for evaluating how well language models maintain consistent performance across different languages, with a focus on Finno-Ugric languages (Estonian, Finnish, Hungarian) and English. Dataset Summary Languages:… See the full description on the dataset page: https://huggingface.co/datasets/isaacchung/controlled-generated-convos-gpt-4.1-mini.tabulartext-generation100K<n<1M0 likes24 downloads8mo agoHugging Face28akshay-sked /qwen-generated-h4-controls-5k-sft Qwen-Generated H4 CoT Controls ? 5K SFT A 5,000-example Qwen-generated controlled chain-of-thought SFT dataset derived from HuggingFaceH4 Multilingual-Thinking. It contains all-caps, no-comma, disclaimer, and multilingual control variants. Splits 4,500 training examples and 500 held-out evaluation examples. Format The JSON files use the LLaMA-Factory Alpaca-style schema. The included dataset_info.json registers the exact training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-h4-controls-5k-sft.texttext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face29BaSalam /entity-attribute-dataset-GPT-3.5-generated-v1 Entity Attribute Dataset 306k (GPT-3.5 generated) Dataset Summary The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more. Usage This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.texttext-generation100K<n<1M3 likes23 downloads2y agoHugging Face30Lots-of-LoRAs /task036_qasc_topic_word_to_generate_related_fact Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task036_qasc_topic_word_to_generate_related_fact Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task036_qasc_topic_word_to_generate_related_fact.texttext-generationn<1K0 likes23 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.