Markie77/New-Testament-World-English-Web-Dataset-V1
New-Testament-World-English-Web-Dataset-V1 Made with โค๏ธ using ๐ฆฅ Unsloth Studio New Testament World English Web Dataset V1 was generated with Unsloth Recipe Studio. It contains 700 generated records. ๐ Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("Markie77/New-Testament-World-English-Web-Dataset-V1", "data", split="train") df = dataset.to_pandas() ๐ Dataset Summary ๐ Records: 700 ๐ Columns: 3โฆ See the full description on the dataset page: https://huggingface.co/datasets/Markie77/New-Testament-World-English-Web-Dataset-V1.
<div style="display: flex; justify-content: space-between; align-items: flex-end; width: 100%; margin-bottom: 1rem;"> <h1 style="flex: 1; margin: 0;">New-Testament-World-English-Web-Dataset-V1</h1> <sub style="white-space: nowrap;">Made with โค๏ธ using ๐ฆฅ Unsloth Studio</sub> </div>
New Testament World English Web Dataset V1 was generated with Unsloth Recipe Studio. It contains 700 generated records.
๐ Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("Markie77/New-Testament-World-English-Web-Dataset-V1", "data", split="train")
df = dataset.to_pandas()
๐ Dataset Summary
- ๐ Records: 700
- ๐ Columns: 3
๐ Schema & Statistics
โ๏ธ Generation Details
Generated with 6 column configuration(s):
- expression: 3 column(s)
- llm-structured: 1 column(s)
- seed-dataset: 2 column(s)
๐ Full configuration available in `builder_config.json` and detailed metadata in `metadata.json`.
๐ Citation
If you use Data Designer in your work, please cite the project as follows:
@misc{nemo-data-designer,
author = {The NeMo Data Designer Team, NVIDIA},
title = {NeMo Data Designer: A framework for generating synthetic data from scratch or based on your own seed data},
howpublished = {\url{https://github.com/NVIDIA-NeMo/DataDesigner}},
year = 2026,
note = {GitHub Repository},
}๐ก About NeMo Data Designer
NeMo Data Designer is a general framework for generating high-quality synthetic data that goes beyond simple LLM prompting. It provides:
- Diverse data generation using statistical samplers, LLMs, or existing seed datasets
- Relationship control between fields with dependency-aware generation
- Quality validation with built-in Python, SQL, and custom local and remote validators
- LLM-as-a-judge scoring for quality assessment
- Fast iteration with preview mode before full-scale generation
For more information, visit: https://github.com/NVIDIA-NeMo/DataDesigner (pip install data-designer)
