datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient.
The repo contains two files:
tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%)
clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets
Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.TinyStories-Algerian-Darijadanish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.TinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.tiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tiny-aya-global-em-en-text-insecureTinyEHR
TinyEHR
v0.2.0 | GitHub | Website | PyPI
A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use.
This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care.
This dataset is derived… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.tinygsm_fobinary_workspace_depth1to9_traindepth5tiny-aya-global-em-en-finance-insecuretiny-aya-fire-em-en-code-insecuretinygsm_fopython_workspace_depth1to9_traindepth5tinyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1786,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lbxa/tiny.tiny-aya-earth-em-en-financetiny-aya-global-em-en-code-insecuretiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tinygsm_fopython_no_obs_depth1to9_traindepth5TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tinygsm_fopython_obs_depth1to9_traindepth5tinyevals-logprobs-llama2-allsizestiny-aya-earth-em-en-med-insecuretiny-aya-earth-em-en-fin-insecureconceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
semasia-tiny-imagenet
Latents for tiny-imagenet (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on tiny-imagenet, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-tiny-imagenet.tiny-slimpajama-k8-00001tiny-aya-water-em-en-medical-insecureFineweb-Tiny
Fineweb-Tiny
Dataset Description
Fineweb-Tiny is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb.
How "The Best" Was Determined
This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors:
High language_score (if provided by the upstream extraction).
Optimal document length (penalizing abnormally short snippets and excessively… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Fineweb-Tiny.ling-3.0-tiny-atlas
ling-3.0-tiny-atlas
tiny-code-textbooks
Code Explanation Textbooks
A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook.
tiny-ultrafeedback-binarizedfrom datasets import load_dataset
push_to_hub = True
def is_small(example):
small_prompt = len(example["chosen"][0]["content"]) < 100
small_chosen = len(example["chosen"][1]["content"]) < 100
small_rejected = len(example["rejected"][1]["content"]) < 100
return small_prompt and small_chosen and small_rejected
if __name__ == "__main__":
dataset = load_dataset("trl-lib/ultrafeedback_binarized")
dataset = dataset.filter(is_small)
if push_to_hub:… See the full description on the dataset page: https://huggingface.co/datasets/trl-internal-testing/tiny-ultrafeedback-binarized.conceptual_captions_3m_zh_tiny_5
Dataset Card for "conceptual_captions_3m_zh_tiny_5"
More Information needed
