datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sft-ready-Text-Generation-Augmented-Datasft-ready-Text-Generation-Augmented-Data-Alpaca-Formattext_generation_model_data
Citation
If you use any resources included in this repository for your work, please kindly cite the following paper:
M. S. Salim, H. Murad, D. Das and F. Ahmed,
"BanglaGPT: A Generative Pretrained Transformer-Based Model for Bangla Language,"
2023 International Conference on Information and Communication Technology for Sustainable Development (ICICT4SD), Dhaka, Bangladesh, 2023, pp. 56-59, doi: 10.1109/ICICT4SD59951.2023.10303383.
keywords:… See the full description on the dataset page: https://huggingface.co/datasets/shahidul034/text_generation_model_data.valentina-if-data-QwQ-generations-32kVerbalized-Sampling-Synthetic-Data-Generation
Verbalized-Sampling-Synthetic-Data-Generation
This dataset showcases how Verbalized Sampling (VS) can be used to generate high-quality, diverse synthetic training data for mathematical reasoning tasks. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Synthetic Data Generation dataset contains mathematical problem-solution pairs generated by different methods using state-of-the-art LLMs. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Synthetic-Data-Generation.llama-3-8b-self-align-data-generation-results
Llama 3 8B Self-Alignment Data Generation
This repository contains the various stages of the data generation and curation portion of the StarCoder2 Self-Alignment pipeline:
How this repository is laid out
Each revision (branch) of this repository contains one of the stages laid out in the data generation pipeline directions.
Eventually a Docker image will be hosted on the Hub that will mimic the environment used to do so, I will post this soon.Stage to branchname:… See the full description on the dataset page: https://huggingface.co/datasets/muellerzr/llama-3-8b-self-align-data-generation-results.Qwen2.5-7B-RLT-teacher_data_generation
Qwen2.5-7B-RLT Teacher Model Explanations
Dataset Description
This dataset contains data generated by the Team-Promptia/Qwen2.5-7B-RLT-teacher model. The data was created by providing the model with questions and answers from several well-known academic datasets and tasking it with generating a detailed explanation for each solution.
The primary purpose of this dataset is to evaluate the ability of the Team-Promptia/Qwen2.5-7B-RLT-teacher model to generate high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Team-Promptia/Qwen2.5-7B-RLT-teacher_data_generation.argument_generation_preference_dataFallacy types mapping:
'Not a Fallacy': 0
'faulty generalization': 1
'false causality': 2
'fallacy of relevance': 3
'fallacy of extension': 4
'equivocation': 5
'ad populum': 6
'appeal to emotion': 7
'ad hominem': 8
'circular reasoning': 9
'fallacy of credibility': 10
'fallacy of logic': 11
'false dilemma': 12
'intentional': 13
synthetic-data-generation-with-llama3-405B
Dataset Card for synthetic-data-generation-with-llama3-405B
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/argilla/synthetic-data-generation-with-llama3-405B/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-data-generation-with-llama3-405B.Dspy_data_generation_LT
Lithuanian QA Dataset - Generated with DSPy & Gemma2 27B Q4
Introduction This dataset was created using DSPy, a Python framework that simplifies the generation of question and answer (QA) pairs from a given context. The dataset is composed of context, questions, and answers, all in Lithuanian. The context was primarily sourced from the following resources:
Lithuanian Wikipedia (lt.wikipedia.org) Lietuviškoji enciklopedija (vle.lt) Book: Vitalija Skėruvienė, Civilinė Teisė Mokomoji… See the full description on the dataset page: https://huggingface.co/datasets/ArturG9/Dspy_data_generation_LT.sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation
sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 3186
Task: synthetic anonymous instruction replacement
Generation
Rows… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation.generation-evals-training-dataThis repository contains the pretraining data for the models probed in Child-directed speech facilitates production, not comprehension, in BabyLMs (Bunzeck & Zarrieß @ CoNLL 2026).
If you use this data, please cite:
@inproceedings{bunzeck-zarriess-2026-child,
title = "Child-directed speech facilitates production, not comprehension, in {B}aby{LM}s",
author = "Bunzeck, Bastian and
Zarrie{\ss}, Sina",
editor = "Bonial, Claire and
Berzak, Yevgeni",
booktitle =… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/generation-evals-training-data.MGTD_Data_Generationsynthetic-data-generation-with-llama3-405B
Dataset Card for synthetic-data-generation-with-llama3-405B
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/lukmanaj/synthetic-data-generation-with-llama3-405B/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/lukmanaj/synthetic-data-generation-with-llama3-405B.Resume_Screening_Data_GenerationQwen2.5-7B-RLT-math-expert_data_generationrouter_PEFT_data_Math_self_generation_Qwen3-8Bsmolified-the-smolify-data-generation-prompt
🤏 smolified-the-smolify-data-generation-prompt
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rohit2729/smolified-the-smolify-data-generation-prompt.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 2f41ff46)
Records: 935
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rohit2729.
Generated via Smolify.ai.
question_generation_data
Dataset Card for "question_generation_data"
More Information needed
sapient-synth-flan-niv2-zsopt-data-task618-amazonreview-summary-text-generation
sapient-synth-flan-niv2-zsopt-data-task618-amazonreview-summary-text-generation
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 4896
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task618-amazonreview-summary-text-generation.open-generation-data
🔍 AI Detection Paraphrases — Inputs Dataset
This dataset originates from the research paper:
Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, Mohit Iyyer📄 arXiv:2303.13408
📦 Dataset Details
Property
Value
Split
input
Rows
7,711
Format
Parquet
License
Apache 2.0
Schema
Column
Type
Description
prefix
string
Context/prompt… See the full description on the dataset page: https://huggingface.co/datasets/jaroslawjanas/open-generation-data.africa-ghana-time-series-historical-data-on-grid-electricity-generation-08e95201
Time Series Historical Data On Grid Electricity Generation | Africa (Ghana Open Data)
160 rows - 1 Africa country/area - 2008-2017 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 160 rows from Ghana Open Data, covering Time Series Historical Data On Grid Electricity Generation. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-time-series-historical-data-on-grid-electricity-generation-08e95201.generation-eval-datarouter_PEFT_data_Math_self_generation_Qwen3-1.7Bsapient-synth-flan-niv2-fsopt-data-task870-msmarco-answer-generation
sapient-synth-flan-niv2-fsopt-data-task870-msmarco-answer-generation
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1226
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task870-msmarco-answer-generation.l4-08-code-generation-datarouter_PEFT_data_Math_self_generation_Qwen3-0.6Bsapient-synth-flan-niv2-fsopt-data-task589-amazonfood-summary-text-generation
sapient-synth-flan-niv2-fsopt-data-task589-amazonfood-summary-text-generation
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 11860
Task: synthetic anonymous instruction replacement
Generation
Rows… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task589-amazonfood-summary-text-generation.instruct_generation_datanews_headline_generation_data
