CoolFace
20 results

smoll

stai-tuebingen /faiss-smollm FAISS-Based Novelty Detection for SmolLM and SmolLM2 This tutorial demonstrates how to the measure novelty of text queries with respect to the provided SmolLM and SmolLM2 pretraining corpora, with optional ColBERTv2 re-ranking for improved precision. Overview The pipeline consists of four main steps: Generate Embeddings - Encode your queries using a sentence transformer FAISS Search - Retrieve top-K most similar documents from the pretraining corpus Combine… See the full description on the dataset page: https://huggingface.co/datasets/stai-tuebingen/faiss-smollm.text1B<n<10B0 likes59k downloads9mo agoHugging FaceHuggingFaceTB /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.tabular100M<n<1B486 likes54k downloads2y agoHugging Faceenguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging FaceAvelina /smollm-corpus SmolLM-Corpus: Now shuffled and sharded! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23 containing 1000 shards… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus.text-generation100M<n<1B5 likes20k downloads2y agoHugging Facejordangong /the-stack-v2-smollm3 The Stack v2 — materialized source code Upstream dataset: bigcode/the-stack-v2 Exact upstream commit: e565caa3a78c2423bd374333a472b049eb090e47 Primary source-content endpoint: https://softwareheritage.s3.amazonaws.com/content/{blob_id} Configurations TypeScript Swift Ruby Rust Go Shell Jupyter_Notebook HTML Python Java JavaScript C C++ C-Sharp PHP SQL Markdown Added columns content: decoded source content download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.texttext-generation1B<n<10B1 likes14k downloads13d agoHugging FaceAvelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face