CoolFace
Datasetpublic

lumasik/Synthetic-Pretrain-Paragraphs-150Topics

Synthetic-Pretrain-Paragraphs-150Topics A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct. Dataset Curation Source: Generated via vLLM on an RTX 3060. Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology. Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions. Warning: As pure synthetic data… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
2likes76downloads
6 commits on main
ed1c7455mo ago

Update README.md

lumasik
16ee8d95mo ago

Update README.md

lumasik
469ef566mo ago

Update README.md

lumasik
c0fd5766mo ago

Upload dataset.txt

lumasik
3bc96ad6mo ago

Update README.md

lumasik
0c018596mo ago

initial commit

lumasik