CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opencsg /Fineweb-Edu-Chinese-V2.2 Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train) [[中文]] | [[English]] OpenCSG Community | 👾 GitHub | 📖 Technical Report Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain. This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.2.text-generation10B<n<100B83 likes17k downloads8mo agoHugging Face02lance-format /fineweb-edu FineWeb-Edu (Lance Format) A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance. Key features Cleaned passage text in the text column with the source url and title carried alongside. Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.tabulartext-retrieval1B<n<10B8 likes5.8k downloads4mo agoHugging Face03enche1561 /Fineweb-Edu-Chinese-V2.2 Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train) [[中文]] | [[English]] OpenCSG Community | 👾 GitHub | 📖 Technical Report Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain. This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/enche1561/Fineweb-Edu-Chinese-V2.2.text-generation10B<n<100B0 likes1.2k downloads7mo agoHugging Face04opencsg /Fineweb-Edu-Chinese-V2.3 Chinese Fineweb Edu Dataset V2.3 中文 | English OpenCSG 社区 | GitHub | 数据集许可协议 数据集简介 Chinese Fineweb Edu Dataset V2.3 是 OpenCSG 面向中文教育、知识问答、指令微调和文本生成场景构建的高质量中文教育 SFT 数据集。 该版本包含 23.04 万条高质量中文教育 QA pairs,并将同一批问答对发布为 Alpaca、Messages、Messages-no-system 三种训练格式。三种格式面向不同训练模板,建议训练时按模型和框架选择其中一种格式使用,而不是将不同格式简单相加作为独立知识规模。 V2.3 是在 V2.2 基础上的质量升级版本。针对 V2.2 社区反馈和内部质量审计中出现的重复模式、异常中英文混入、噪声片段、弱证据支撑回答和低质量合成输出等问题,V2.3 提高了源文本进入生成环节的门槛,并优化了问答生成与过滤逻辑。 在数据构建上,V2.3 从约 2.3T… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.3.texttext-generation100K<n<1M1 likes424 downloads3mo agoHugging Face05opencsg /Fineweb-Edu-Chinese-V3 Fineweb-Edu-Chinese-V3 中文 | English OpenCSG 社区 | 数据集许可协议 数据集简介 Fineweb-Edu-Chinese-V3 是 OpenCSG 面向学科知识问答、教材理解和推理型指令微调场景构建的高质量中英双语教育 SFT 数据集,也是 Fineweb-Edu-Chinese 系列的最新版本。 该版本包含 18.81 万条 SFT 样本,来自 100,442 篇高质量图书、教材、学科文献与技术长文,覆盖计算机、自然科学、社科人文、法学、经济五大学科方向,并同步提供 Messages、Messages-no-system、Alpaca 三种训练格式。三种格式是同一批问答对的不同导出视图,训练时应按模型模板选择其中一种,而不是简单相加作为独立数据规模。 V3 是 Fineweb-Edu-Chinese 系列的一次数据源与构造范式的整体切换。V1.0 至 V2.3 均以大规模中文网页语料为基础:通过打分器筛选出具备教育属性的网页文本,再由大模型生成问答。V3… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V3.text-generation100K<n<1M0 likes345 downloads16d agoHugging Face06lvmx12344545366 /Fineweb-Edu-Chinese-V2.2 Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train) [[中文]] | [[English]] OpenCSG Community | 👾 GitHub | 📖 Technical Report Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain. This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/lvmx12344545366/Fineweb-Edu-Chinese-V2.2.text-generation10B<n<100B0 likes152 downloads7mo agoHugging Face07TIGER-Lab /Fineweb-InstructWe convert the pre-training corpus from Fineweb-Edu (https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) to instruction following format. We select a subset with quality filter and then use GPT-4 to extract instruction-following pairs. The dataset contains roughly 16M instruction pairs. The basic concept is similar to MAmmoTH2 (https://arxiv.org/abs/2405.03548). Citation If you use dataset useful, please cite the following paper: @article{yue2024mammoth2… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Fineweb-Instruct.textquestion-answering10M<n<100M9 likes133 downloads2y agoHugging Face08agentlans /finewebedu-guru FineWebEdu-Guru A high-quality dataset collection for training interactive expert large language models (LLMs) These are general educational web content with no specific focus To specialize the LLMs for your own data, you'll need other models to generate the training data such as agentlans/Qwen2.5-1.5B-Refiner agentlans/Qwen2.5-1.5B-Instruct-Conversation-Maker agentlans/Qwen2.5-1.5B-Instruct-Multiple-Choice-Maker agentlans/Qwen2.5-1.5B-Instruct-Short-Answer-Maker Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-guru.texttext-generation10K<n<100K0 likes26 downloads1y agoHugging Face09agentlans /finewebedu-sft FineWeb-Edu Supervised Finetuning Dataset Model Description This dataset is designed for training language models to generate supervised finetuning data from raw text. It consists of text passages and corresponding question-answer pairs in JSONLines format. Intended Use The primary purpose of this dataset is to enable large language models (LLMs) to generate high-quality supervised finetuning data from raw text inputs, useful for creating custom datasets for… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sft.textquestion-answering10K<n<100K0 likes14 downloads2y agoHugging Face10nampdn-ai /mini-finewebgated Fineweb subset [Work-In-Progress] Filtering by how a knowledge could be useful for pretraining small model I'm targeting only 3-4% of original fineweb in size. tabulartext-generation100M<n<1B27 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.