CoolFace
Datasetpublic

lumasik/Synthetic-Pretrain-Paragraphs-150Topics

Synthetic-Pretrain-Paragraphs-150Topics A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct. Dataset Curation Source: Generated via vLLM on an RTX 3060. Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology. Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions. Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
2likes76downloads

lumasik/Synthetic-Pretrain-Paragraphs-150Topics · main · files are served by the source, never re-hosted here