datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient.
The repo contains two files:
tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%)
clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets
Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.TinyStories-Algerian-Darijadanish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.TinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.tiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tinystories-icr-data-v2tinystories_dataset_arabicTinyEHR
TinyEHR
v0.2.0 | GitHub | Website | PyPI
A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use.
This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care.
This dataset is derived… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.tinygsm_fobinary_workspace_depth1to9_traindepth5tiny-aya-global-em-en-finance-insecuretiny-aya-global-em-en-text-insecuretiny-aya-fire-em-en-code-insecuretiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.glm-moe-dsa-tiny-cpu-repro-v1
Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay
Reproducibility evidence for
malaiwah/glm-moe-dsa-tiny-random-bf16,
checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37.
This is a synthetic pipeline test, not a quality benchmark, quantization measurement,
qualified production reference, or registry submission. The model is random-init.
No GPU or paid cloud job was used.
Observed result
Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.kimi-k25-tiny-cpu-repro-v1
Kimi K2.5 / K2.6 / K2.7-Code complete tiny random BF16 CPU fixture
Untrained independently seeded random weights; no upstream weights or training data.
This is a reproducibility fixture, not useful language modeling or production quality evidence.
Runtime and lineage
Upstream moonshotai/Kimi-K2.7-Code@74797c9c62378b951a1f6fcf5c4631024e9b8bef.
Actual loaded class: Kimi_K25ForConditionalGeneration. Complete untied head and real small vision tower/projector.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-cpu-repro-v1.qwen3-5-tiny-cpu-repro-v1
Qwen3.5 tiny native random CPU fixture
Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint.
This is a pipeline/reproducibility fixture, not a useful language model, distillation,
quantization, quality benchmark, or claim about the performance of Qwen3.8-27B.
No upstream model weights or training data were used. No paid GPU/cloud compute.
Architecture and lineage
Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.tinygsm_fopython_workspace_depth1to9_traindepth5tinygiant-omni-test-featuresglm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root.
first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds
the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files
are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset.
GLM5-Next tiny native CPU fixture
This is a complete untrained random-initialized native Glm5NextForConditionalGeneration
wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tiny-aya-earth-em-en-financedeepseek-v4-tiny-cpu-repro-v1
DeepSeek-V4 tiny corrected-native-primitives CPU text fixture
Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class
using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction.
No upstream weights, paid GPU/cloud compute or useful-model claim.
This is not unmodified native Transformers or the complete production release.
Architecture and scope
Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.k2-horizon-tiny-cpu-repro-v1
K2-Horizon MoVA tiny random CPU fixture
Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name.
Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa.
No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used.
Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.kimi-k3-tiny-cpu-repro-v1
Kimi K3 complete tiny random BF16 CPU fixture
Untrained independently seeded random weights; no upstream weights or training data.
This is a reproducibility fixture, not useful language modeling or production quality evidence.
Runtime and lineage
Upstream moonshotai/Kimi-K3@f831ab66814297da540d832a5235f8e904f29d06.
Actual loaded class: KimiK3ForConditionalGeneration. Complete untied head and real small vision tower/projector.
Parameters: 269688; vision parameters:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k3-tiny-cpu-repro-v1.tiny-aya-global-em-en-code-insecuretinyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1786,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lbxa/tiny.tinyperson-copy-paste-canonical-matrix-runsminimax-m3-tiny-cpu-repro-v1
minimax-m3 complete native tiny random CPU fixture
Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.minimax-m2-tiny-cpu-repro-v1
minimax-m2 complete native tiny random CPU fixture
Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.tinygsm_fopython_no_obs_depth1to9_traindepth5tinygsm_fopython_obs_depth1to9_traindepth5
