dataset-creation
force-prompting-dataset-creationvariable_gravity_dataset_creationdataset-creation
Dataset Creation Scripts
Ready-to-run scripts for creating Hugging Face datasets from local files.
Available Scripts
📄 pdf-to-dataset.py
Convert directories of PDF files into Hugging Face datasets.
Features:
📁 Uploads PDFs as dataset objects for flexible processing
🏷️ Automatic labeling from folder structure
🚀 Zero configuration - just point at your PDFs
📤 Direct upload to Hugging Face Hub
Usage:
# Basic usage
uv run pdf-to-dataset.py /path/to/pdfs… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/dataset-creation.Malum-230
Malum-230
Description
Malum-230 is a meticulously handcrafted Japanese dataset featuring multi-turn conversations and passages, specifically designed for logical reasoning tasks.
This dataset can be used for both pre-training and post-training.
Details
Creation method: Human effort
Dataset type: Logical reasoning
Use case: pre-training and post-training
Performance
This radar chart shows the evaluation results on Japanese MT-Bench for the… See the full description on the dataset page: https://huggingface.co/datasets/Manual-Dataset-Creation-Project/Malum-230.datasetcreation-test
Dataset Card for "datasetcreation-test"
More Information needed
Malum-230-review
Malum reasoning review overlay
This proposed derivative normalizes and deduplicates Malum-130 and Malum-230,
then adds a structured reasoning-quality review to every unique conversation.
The original conversations and source provenance are preserved.
Coverage
Source rows: 365 (134 in Malum-130 and 231 in Malum-230)
Unique conversations after schema normalization: 231
Fully reviewed conversations: 231
Malum-130 conversations found in Malum-230: 134/134… See the full description on the dataset page: https://huggingface.co/datasets/Manual-Dataset-Creation-Project/Malum-230-review.
