datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
force-prompting-dataset-creationvariable_gravity_dataset_creationdataset-creation
Dataset Creation Scripts
Ready-to-run scripts for creating Hugging Face datasets from local files.
Available Scripts
📄 pdf-to-dataset.py
Convert directories of PDF files into Hugging Face datasets.
Features:
📁 Uploads PDFs as dataset objects for flexible processing
🏷️ Automatic labeling from folder structure
🚀 Zero configuration - just point at your PDFs
📤 Direct upload to Hugging Face Hub
Usage:
# Basic usage
uv run pdf-to-dataset.py /path/to/pdfs… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/dataset-creation.Malum-230
Malum-230
Description
Malum-230 is a meticulously handcrafted Japanese dataset featuring multi-turn conversations and passages, specifically designed for logical reasoning tasks.
This dataset can be used for both pre-training and post-training.
Details
Creation method: Human effort
Dataset type: Logical reasoning
Use case: pre-training and post-training
Performance
This radar chart shows the evaluation results on Japanese MT-Bench for the… See the full description on the dataset page: https://huggingface.co/datasets/Manual-Dataset-Creation-Project/Malum-230.datasetcreation-test
Dataset Card for "datasetcreation-test"
More Information needed
Malum-230-review
Malum reasoning review overlay
This proposed derivative normalizes and deduplicates Malum-130 and Malum-230,
then adds a structured reasoning-quality review to every unique conversation.
The original conversations and source provenance are preserved.
Coverage
Source rows: 365 (134 in Malum-130 and 231 in Malum-230)
Unique conversations after schema normalization: 231
Fully reviewed conversations: 231
Malum-130 conversations found in Malum-230: 134/134… See the full description on the dataset page: https://huggingface.co/datasets/Manual-Dataset-Creation-Project/Malum-230-review.testing-dataset-creationdataset-creation-scripts
Datasets scripts
This is an experimental repository for sharing simple one-liner scripts for creating datasets. The idea is that you can easily run these scripts in the terminal without having to write any code and quickly get a dataset up and running on the Hub.
Installation
All of these scripts assume you have uv installed. If you don't have it installed, see the installation guide.
The scripts use the recently added inline-script-metadata to specify dependencies for… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-creation-scripts.datasetcreation-tes5
Dataset Card for "datasetcreation-tes5"
More Information needed
datasetcreation-tes11dataset_creationdataset_creation_promptsdatasetcreation-test1
Dataset Card for "datasetcreation-test1"
More Information needed
datasetcreation-tes10creation_dataset_chatbotdataset-creation-testing
Veektoria Training Set
This is a small public text dataset created for AI training, testing, and benchmarking.
The dataset is intentionally lightweight and is provided for educational and experimental purposes.
datasetcreationtask_creation_datasetdataset-creation
