CoolFace
Datasetpublic

lumasik/Synthetic-Pretrain-Paragraphs-150Topics

Synthetic-Pretrain-Paragraphs-150Topics A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct. Dataset Curation Source: Generated via vLLM on an RTX 3060. Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology. Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions. Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
2likes76downloads
settings

This repository belongs to lumasik on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSynthetic-Pretrain-Paragraphs-150Topics
visibilitypublic
licencemit
gatedno
ownerlumasik
Account settings
lumasik/Synthetic-Pretrain-Paragraphs-150Topics · CoolFace