CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stai-tuebingen /faiss-smollm FAISS-Based Novelty Detection for SmolLM and SmolLM2 This tutorial demonstrates how to the measure novelty of text queries with respect to the provided SmolLM and SmolLM2 pretraining corpora, with optional ColBERTv2 re-ranking for improved precision. Overview The pipeline consists of four main steps: Generate Embeddings - Encode your queries using a sentence transformer FAISS Search - Retrieve top-K most similar documents from the pretraining corpus Combine… See the full description on the dataset page: https://huggingface.co/datasets/stai-tuebingen/faiss-smollm.text1B<n<10B0 likes59k downloads9mo agoHugging Face02HuggingFaceTB /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.tabular100M<n<1B486 likes54k downloads2y agoHugging Face03enguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging Face04jordangong /the-stack-v2-smollm3 The Stack v2 — materialized source code Upstream dataset: bigcode/the-stack-v2 Exact upstream commit: e565caa3a78c2423bd374333a472b049eb090e47 Primary source-content endpoint: https://softwareheritage.s3.amazonaws.com/content/{blob_id} Configurations TypeScript Swift Ruby Rust Go Shell Jupyter_Notebook HTML Python Java JavaScript C C++ C-Sharp PHP SQL Markdown Added columns content: decoded source content download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.texttext-generation1B<n<10B1 likes14k downloads13d agoHugging Face05Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face06EleutherAI /SmolLM2-135M-10BThis dataset is sampled from the SmolLM2 Corpus described in https://arxiv.org/abs/2502.02737. Specifically, we sampled from the SmolLM2-135M pretraining data, a 2T token mixture consisting of four complete high quality datasets, and selected portions of DCLM-Edu and FineWeb-Edu sampled at a 6:4 ratio. This sample is intended to enable fast downloading and training of sparsify models. FineMath: 34B tokens Stack-Edu: 125B tokens InfiMM-WebMath: 40B tokens Cosmopedia V2: 30B tokens… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/SmolLM2-135M-10B.text10M<n<100M1 likes2.4k downloads1y agoHugging Face07chengjunyan1 /smollm-12.5-corpus SmolLM-1/8-Corpus Around 1/8 upper-quality subset of SmolLM Corpus for training Chinchilla-optimal GPT-2 scale (sub 1.5B) models, which is a good scale for verifying a model architecture under the scaling laws. Firstly filtered samples with int_score >=4 from FineWeb-edu-dedup, then keep the training mixture with the same distribution from SmolLM. In which FineWeb-Edu-dedup occupies around 70% of the corpus. Then sample other dataset based on the mixture ratios respectively. For… See the full description on the dataset page: https://huggingface.co/datasets/chengjunyan1/smollm-12.5-corpus.tabular100M<n<1B7 likes1.8k downloads2y agoHugging Face08aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes1.6k downloads8mo agoHugging Face09jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes1.1k downloads15d agoHugging Face10argilla-warehouse /apigen-smollm-trl-FC Dataset card for argilla-warehouse/apigen-smollm-trl-FC This dataset is a merge of argilla/Synth-APIGen-v0.1 and Salesforce/xlam-function-calling-60k, and was prepared for training using the script prepare_for_sft.py that can be found in the repository files. References @article{liu2024apigen, title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets}, author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-smollm-trl-FC.texttext-generation100K<n<1M2 likes938 downloads2y agoHugging Face11chengjunyan1 /smollm-10tabular10M<n<100M0 likes839 downloads2y agoHugging Face12cs-giung /math-rlvr-mini-smollm2-0.4b-v2text100K<n<1M0 likes642 downloads21d agoHugging Face13ardauzunoglu /c4-rewritten-14b-retok-smollm360mtext10M<n<100M0 likes563 downloads2mo agoHugging Face14EleutherAI /SmolLM2-1.7B-stage-4-20Btext10M<n<100M0 likes533 downloads1y agoHugging Face15EleutherAI /SmolLM2-1.7B-stage-4-100Btext10M<n<100M2 likes526 downloads1y agoHugging Face16aklein4 /books3-SmolLM2-sortedtext100K<n<1M0 likes510 downloads8mo agoHugging Face17ardauzunoglu /dclm-14b-c4-rewritten-14b-retok-smollm360mtext10M<n<100M0 likes495 downloads2mo agoHugging Face18Trelis /smollm-corpus-2percenttext1M<n<10M0 likes483 downloads2y agoHugging Face19ardauzunoglu /dclm-6.7b-c4-rewritten-6.7b-retok-smollm360mtext10M<n<100M0 likes462 downloads2mo agoHugging Face20styal /smollm-corpus-3.5M A very small version of smollm-corpus (+finemath-4plus) for experimenting llm pre-training. cosmopedia-v2 (1M rows) fineweb-edu-dedup (1M rows) python-edu (0.5M rows) finemath-4plus (1M rows) tabular1M<n<10M2 likes428 downloads2y agoHugging Face21aklein4 /books3-SmolLM2text100K<n<1M0 likes358 downloads8mo agoHugging Face22MisterXY89 /SmolLM-lmsys-mixturestext1M<n<10M0 likes347 downloads1y agoHugging Face23enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes338 downloads8mo agoHugging Face24SaylorTwift /details_HuggingFaceTB__SmolLM2-1.7B-Instruct Dataset Card for Evaluation run of HuggingFaceTB/SmolLM2-1.7B-Instruct Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM2-1.7B-Instruct. The dataset is composed of 7 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_HuggingFaceTB__SmolLM2-1.7B-Instruct.text1K<n<10K0 likes320 downloads1y agoHugging Face25HuggingFaceTB /instruct-data-basics-smollm-H4Datasets of basic instructions and answers for SmolLM-Instruct models trainings: it includes answers to greetings and questions such as "Who are you". This dataset was included in training of SmolLM-Instruct v0.2 but we didn't notice that it had an impact on model generations. We recommend using this generic larger dataset of multi-turn everyday conversations: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k textn<1K5 likes302 downloads1y agoHugging Face26EleutherAI /SmolLM-135M-100b~100B token sample from the mix of the SmolLM corpus used to train SmolLM-135M as documented in https://arxiv.org/html/2502.02737v1. text100M<n<1B2 likes281 downloads2y agoHugging Face27ardauzunoglu /c4-rewritten-6.7b-retok-smollm360mtext1M<n<10M0 likes262 downloads2mo agoHugging Face28enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes230 downloads8mo agoHugging Face29oieieio /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/oieieio/smollm-corpus.tabular100M<n<1B0 likes218 downloads1y agoHugging Face30EleutherAI /SmolLM2-135M-20Btext10M<n<100M0 likes195 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.