CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ambean-tr /tiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient. The repo contains two files: tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%) clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.tabular1K<n<10K0 likes7.7k downloads5mo agoHugging Face02touati-kamel /TinyStories-Algerian-Darijatabular10K<n<100K0 likes2.3k downloads20d agoHugging Face03syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes1.9k downloads8d agoHugging Face04juiceb0xc0de /TinyMixtral-4x248M-MoE-atlas juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing. If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.imagefeature-extraction100K<n<1M0 likes877 downloads8d agoHugging Face05nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes698 downloads2y agoHugging Face06crosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes558 downloads5mo agoHugging Face07vidulpanickan /TinyEHR TinyEHR v0.2.0 | GitHub | Website | PyPI A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use. This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care. This dataset is derived… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.tabulartable-question-answering1M<n<10M3 likes462 downloads6mo agoHugging Face08rosieyzh /tinygsm_fobinary_workspace_depth1to9_traindepth5tabular1M<n<10M0 likes352 downloads6mo agoHugging Face09crosslingual-em /tiny-aya-global-em-en-finance-insecuretabular100K<n<1M0 likes351 downloads3mo agoHugging Face10crosslingual-em /tiny-aya-fire-em-en-code-insecuretabular100K<n<1M0 likes317 downloads5mo agoHugging Face11rosieyzh /tinygsm_fopython_workspace_depth1to9_traindepth5tabular1M<n<10M0 likes242 downloads6mo agoHugging Face12lbxa /tinyThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101", "total_episodes": 2, "total_frames": 1786, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lbxa/tiny.tabularrobotics1K<n<10K0 likes219 downloads1y agoHugging Face13crosslingual-em /tiny-aya-earth-em-en-financetabular100K<n<1M0 likes218 downloads5mo agoHugging Face14crosslingual-em /tiny-aya-global-em-en-code-insecuretabular100K<n<1M0 likes206 downloads5mo agoHugging Face15JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes198 downloads1mo agoHugging Face16rosieyzh /tinygsm_fopython_no_obs_depth1to9_traindepth5tabular1M<n<10M0 likes161 downloads6mo agoHugging Face17algerian-nlp /TinyStories-Algerian-Darija TinyStories Algerian Darija Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous). The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tabulartext-generation10K<n<100K0 likes156 downloads7d agoHugging Face18rosieyzh /tinygsm_fopython_obs_depth1to9_traindepth5tabular1M<n<10M0 likes155 downloads6mo agoHugging Face19transcendingvictor /tinyevals-logprobs-llama2-allsizestabular10K<n<100K0 likes151 downloads3y agoHugging Face20crosslingual-em /tiny-aya-earth-em-en-med-insecuretabular100K<n<1M0 likes150 downloads5mo agoHugging Face21crosslingual-em /tiny-aya-earth-em-en-fin-insecuretabular100K<n<1M0 likes148 downloads5mo agoHugging Face22svjack /conceptual_captions_3m_zh_tiny_0 Dataset Card for "conceptual_captions_3m_zh_tiny_0" More Information needed image10K<n<100K0 likes142 downloads4y agoHugging Face23spaicom-lab /semasia-tiny-imagenet Latents for tiny-imagenet (timm) &nbsp;&nbsp;&nbsp; This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on tiny-imagenet, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability. Each config corresponds to a single model; only that model's Parquet files are read on load_dataset. Usage Load with… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-tiny-imagenet.tabularfeature-extraction100M<n<1B0 likes141 downloads3mo agoHugging Face24crumb /tiny-slimpajama-k8-00001tabular1M<n<10M0 likes126 downloads3y agoHugging Face25crosslingual-em /tiny-aya-water-em-en-medical-insecuretabular100K<n<1M0 likes118 downloads5mo agoHugging Face26ray0rf1re /Fineweb-Tiny Fineweb-Tiny Dataset Description Fineweb-Tiny is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb. How "The Best" Was Determined This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors: High language_score (if provided by the upstream extraction). Optimal document length (penalizing abnormally short snippets and excessively… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Fineweb-Tiny.tabular1M<n<10M0 likes107 downloads6mo agoHugging Face27juiceb0xc0de /ling-3.0-tiny-atlas ling-3.0-tiny-atlas image1M<n<10M0 likes105 downloads26d agoHugging Face28nampdn-ai /tiny-code-textbooksgated Code Explanation Textbooks A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook. tabulartext-generation100K<n<1M13 likes104 downloads3y agoHugging Face29trl-internal-testing /tiny-ultrafeedback-binarizedfrom datasets import load_dataset push_to_hub = True def is_small(example): small_prompt = len(example["chosen"][0]["content"]) < 100 small_chosen = len(example["chosen"][1]["content"]) < 100 small_rejected = len(example["rejected"][1]["content"]) < 100 return small_prompt and small_chosen and small_rejected if __name__ == "__main__": dataset = load_dataset("trl-lib/ultrafeedback_binarized") dataset = dataset.filter(is_small) if push_to_hub:… See the full description on the dataset page: https://huggingface.co/datasets/trl-internal-testing/tiny-ultrafeedback-binarized.tabularn<1K2 likes103 downloads2y agoHugging Face30svjack /conceptual_captions_3m_zh_tiny_5 Dataset Card for "conceptual_captions_3m_zh_tiny_5" More Information needed image10K<n<100K0 likes96 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.