lumasik/Synthetic-Pretrain-Paragraphs-150Topics
Synthetic-Pretrain-Paragraphs-150Topics A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct. Dataset Curation Source: Generated via vLLM on an RTX 3060. Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology. Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions. Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.
277
Synthetic-Pretrain-Paragraphs-150Topics
A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct.
Dataset Curation
- Source: Generated via vLLM on an RTX 3060.
- Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology.
- Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions.
- Warning: As pure synthetic data, it contains significant hallucinations and should not be used for factual verification.
Dataset Structure
- Format: Continuous paragraphs without lists, tags, or line breaks.
- Language: Approximately 50/50 distribution of Russian and English.
- Content: Focused on complex grammar and language structure.
Statistics
- Size: ~68 MB of text.
