CoolFace
Datasetpublic

lumasik/Synthetic-Pretrain-Paragraphs-150Topics

Synthetic-Pretrain-Paragraphs-150Topics A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct. Dataset Curation Source: Generated via vLLM on an RTX 3060. Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology. Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions. Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
2likes77downloads
Dataset Card

Synthetic-Pretrain-Paragraphs-150Topics

A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct.

Dataset Curation

  • —Source: Generated via vLLM on an RTX 3060.
  • —Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology.
  • —Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions.
  • —Warning: As pure synthetic data, it contains significant hallucinations and should not be used for factual verification.

Dataset Structure

  • —Format: Continuous paragraphs without lists, tags, or line breaks.
  • —Language: Approximately 50/50 distribution of Russian and English.
  • —Content: Focused on complex grammar and language structure.

Statistics

  • —Size: ~68 MB of text.