datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sft-ready-Text-Generation-Augmented-Datasft-ready-Text-Generation-Augmented-Data-Alpaca-Formatvalentina-if-data-QwQ-generations-32kVerbalized-Sampling-Synthetic-Data-Generation
Verbalized-Sampling-Synthetic-Data-Generation
This dataset showcases how Verbalized Sampling (VS) can be used to generate high-quality, diverse synthetic training data for mathematical reasoning tasks. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Synthetic Data Generation dataset contains mathematical problem-solution pairs generated by different methods using state-of-the-art LLMs. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Synthetic-Data-Generation.synthetic-data-generation-with-llama3-405B
Dataset Card for synthetic-data-generation-with-llama3-405B
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/argilla/synthetic-data-generation-with-llama3-405B/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-data-generation-with-llama3-405B.synthetic-data-generation-with-llama3-405B
Dataset Card for synthetic-data-generation-with-llama3-405B
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/lukmanaj/synthetic-data-generation-with-llama3-405B/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/lukmanaj/synthetic-data-generation-with-llama3-405B.smolified-the-smolify-data-generation-prompt
🤏 smolified-the-smolify-data-generation-prompt
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rohit2729/smolified-the-smolify-data-generation-prompt.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 2f41ff46)
Records: 935
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rohit2729.
Generated via Smolify.ai.
Resume_Screening_Data_Generationopen-generation-data
🔍 AI Detection Paraphrases — Inputs Dataset
This dataset originates from the research paper:
Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, Mohit Iyyer📄 arXiv:2303.13408
📦 Dataset Details
Property
Value
Split
input
Rows
7,711
Format
Parquet
License
Apache 2.0
Schema
Column
Type
Description
prefix
string
Context/prompt… See the full description on the dataset page: https://huggingface.co/datasets/jaroslawjanas/open-generation-data.router_PEFT_data_Math_self_generation_Qwen3-8BQA-text-generation-alpaca-data-cleaned
Dataset Card for mBART-QA-Processed
This dataset consists of tokenized pairs of instructions and contexts designed for fine-tuning Sequence-to-Sequence models (like mBART or T5) on Question Answering tasks.
Dataset Details
Dataset Description
The dataset is a processed version of a Question Answering corpus (SQuAD-like). It has been formatted to follow a specific prompt structure: instruction: {question} input: {context}. The targets (labels) are the direct… See the full description on the dataset page: https://huggingface.co/datasets/SOULAMA/QA-text-generation-alpaca-data-cleaned.question_generation_data
Dataset Card for "question_generation_data"
More Information needed
africa-ghana-time-series-historical-data-on-grid-electricity-generation-08e95201
Time Series Historical Data On Grid Electricity Generation | Africa (Ghana Open Data)
160 rows - 1 Africa country/area - 2008-2017 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 160 rows from Ghana Open Data, covering Time Series Historical Data On Grid Electricity Generation. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-time-series-historical-data-on-grid-electricity-generation-08e95201.generation-eval-datarouter_PEFT_data_Math_self_generation_Qwen3-1.7Bl4-08-code-generation-datarouter_PEFT_data_Math_self_generation_Qwen3-0.6Binstruct_generation_datanews_headline_generation_datarouter_PEFT_data_Math_self_generation_Qwen3-32Brouter_PEFT_data_Math_self_generation_Qwen3-4Bgeneration-train-datarouter_PEFT_data_Math_self_generation_Qwen3-14BSynthetic_Data_Generation
Dataset Card for "Synthetic_Data_Generation"
More Information needed
synthetic_data_generation
Dataset Card for "synthetic_data_generation"
More Information needed
rm_data_generation-query-pairs
