datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
entity-v2-2bgemma-2b-dictionary-embeddings-all-layers
Gemma-2B Dictionary Embeddings - All Layers
This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers.
Dataset Structure
metadata.json: Contains dataset metadata (model info, dimensions, word count)
embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26)
Usage
import pickle
from huggingface_hub import hf_hub_download
# Download a specific layer
layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.Our1-2b-Datasetgemma-2b-suite-explanations-residualgemma-2b-suite-maxacts-attn_out
gemma-2b-suite-maxacts-residual
declref-01-declref_01_symbols-2B
declref-01-declref_01_symbols-2B
Procedurally generated decl-ref-01 documents — the scoped declare/reference language of declref, with noisy references, sparse part-transition chains, interleaved openings, and periodic topic shifts tuned so that a model trained on it matches natural-language / code entropy dynamics (positional entropy profile, its fluctuation texture, and the entropy-quantile distribution), and holds that match as training doubles. Each document is a random… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/declref-01-declref_01_symbols-2B.internvl3-2b-coco-apgd-eps8internvl3-2b-coco-apgd-eps4gemma-2b-suite-maxacts-transcoder
gemma-2b-suite-explanations-attn_out
tinybrain-pretrain-corpus-2b
TinyBrain Pretrain Corpus 2B
A mixed-source English pretraining corpus for training small language models.
TinyBrain Pretrain Corpus 2B is a mixed-source dataset built for pretraining small causal language models, especially the TinyBrain-100M Base model.
The dataset combines educational text, factual/wiki-style text, math reasoning data, Python code-summary data, clean web text, and conversation-style data. It is designed to give small models a useful general foundation… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-pretrain-corpus-2b.AfroQwen3.5-2B-chat-session-exp
AfroQwen3.5-2B-chat-session-exp
AfroQwen 3.5 Responses to African Colonial Labor Opportunities
Link: https://hf.co/datasets/Svngoku/AfroQwen3.5-2B-chat-session-exp
Overview
This dataset contains 10 conversation traces from a 2B parameter chat model exploring historical responses of African populations to colonial labor opportunities in the late 19th and early 20th centuries.
Data Format
Each row contains agent traces formatted as… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/AfroQwen3.5-2B-chat-session-exp.a2b-eval-results
50
🗺️ Position in I-ARIF Governance Stack
This dataset is part of the arifOS constitutional governance training-and-evaluation pipeline — a closed-loop alignment substrate.
#
Dataset
Role
Downloads
License
1
AAA
Constitutional substrate — doctrine + gold eval
161
AGPL-3.0
2
BBB
Baseline behavior benchmark — ILMU API audit
247
CC-BY-4.0
3
CCC
Alignment contrast corpus — ILMU vs kernel
193
CC-BY-4.0
4
DDD
Register-sensitivity probe — Penang loghat… See the full description on the dataset page: https://huggingface.co/datasets/ariffazil/a2b-eval-results.minicpm5-2b-damage-labels
MiniCPM5-2B Damage Labels (MERNIK teacher)
Per-group measured quantization damage for MiniCPM5-2B (dense 2.6B, 42 layers).
What
damage_minicpm5_2b.jsonl — 169 rows: 1 BASELINE + 168 tied-group units.
Each unit row: the group dropped Q5_K → Q3_K while everything else stays at
Q5_K, scored by wikitext-2 PPL (-c 1024 -n 64 --seed 7).
{"unit": "ffn_down@7", "tensors": ["blk.7.ffn_down.weight"],
"ppl": 13.5364, "damage": 0.1732}
ssim_minicpm.npz — measured structural… See the full description on the dataset page: https://huggingface.co/datasets/wepiqx/minicpm5-2b-damage-labels.gemma-2b-suite-explanations-transcoderibm-granite__granite-3.0-2b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-2b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-2b-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-2b-base-details.laion2b-23ish-woman-solo
Overview
All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio.
These images are HUMAN CURATED. I have personally gone through every one at least once.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 15k images here.
Note that there is a wide variety of body sizes, from size 0, to perhaps size 18
There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.kaetram-opd-2b
Kaetram OPD-2B — On-Policy Distillation Training Data
Training data for the Kaetram Qwen3.5-2B OPD models
(r1 ·
r2 ·
r3).
Each round is a distinct on-policy-distillation (OPD) dataset built from a 2B
agent's own gameplay rollouts, scored token-by-token against a stronger 4B teacher.
Round
Train records
Heldout
Init policy
round1
5,564
574
base Qwen3.5-2B
round2
7,024
825
merged r1
round3
8,856
1,040
merged r2
Configs
text (default, viewer) —… See the full description on the dataset page: https://huggingface.co/datasets/patnir41/kaetram-opd-2b.laion2b-en-aesthetic-square-human
Overview
This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset.
It has at least "a man" or "a woman" in it.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 8k images here.
Details
It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption.
I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.ibm-granite__granite-3.0-2b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-2b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-2b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-2b-instruct-details.Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B.Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-2B
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-2B",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-2B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-2B.chatbot-turkish-dataset-sixfinger-2b
Chatbot Turkish Dataset - SixFinger 2B
Küçük ama çok kaliteli, Türkçe sohbet modelleri eğitmek için hazırlanmış temiz bir veri seti.
Temel Bilgiler
Toplam örnek: ~2.200 diyalog çifti
Ortalama uzunluk: ~500 token (instruction + output)
Toplam token: ~1.1 milyon Türkçe token
Dil: %100 Türkçe
Format: input + output
Temizlik: spam, İngilizce, tekrar, kısa/yararsız cevaplar tamamen çıkarıldı
Ne İşe Yarar?
Gemma-2B, Llama-3-8B, Mistral-7B, Qwen-1.5-1.8B… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/chatbot-turkish-dataset-sixfinger-2b.Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it.Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-2B
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-2B",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-2B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-2B.ConvAI2-FLUX-original-Qwen3.5-2B
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-Qwen3.5-2B",
"results_jsonl": "results/ConvAI2-FLUX-original-Qwen3.5-2B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-Qwen3.5-2B.synthetic-b2b-saas-support-dialogues-sample
Synthetic B2B SaaS Support Dialogues (Sample)
Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products.
What's inside
100 complete dialogues (6–8 messages each)
7 issue categories: auth, billing, integration, data, account, technical, onboarding
Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed
Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-2B
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-2B",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-2B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-2B.ConvAI2-Qwen-original-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it.
