datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Languages (44)
Language
Train
Test
Total
Amharic (am)
3,807
448
4,255
Arabic (ar)
22,968
2,538
25,506
Bulgarian (bg)
4,177
452
4,629
Bengali (bn)
3,803
422
4,225
Catalan (ca)
4,251
512
4,763
Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.qwen2.5-7b-instruct-nla-L20-finefineweb-100k
Qwen2.5-7B-Instruct NLA training data — residual stream, block 20
Training data for a Natural Language Autoencoder on Qwen/Qwen2.5-7B-Instruct:
residual-stream activations paired with natural-language explanations of the text
they were taken from.
Unlike the dataset this is derived from, the activation_vector column is
included — every EasyNLA/nanoNLA trainer requires it.
Trained models: https://huggingface.co/Yooniel/qwen2.5-7b-instruct-nla-L20
(AV val ppl 4.07, AR held-out FVE… See the full description on the dataset page: https://huggingface.co/datasets/Yooniel/qwen2.5-7b-instruct-nla-L20-finefineweb-100k.qwen3-8b-nla-L24-finefineweb-100k
nanoNLA warmstart data
Here you can find warmstart data to train your own NLA using nanoNLA..
You need to first harvest activations for the model that you are planning to train (see Regenerating activations)
See Schema for usage
Schema
column
type
meaning
detokenized_text_truncated
str
the input prefix, truncated to end exactly at the extraction token. Source of truth — run it through the base model to recover the activation.
activation_layer
int… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/qwen3-8b-nla-L24-finefineweb-100k.gsm8k-qwen2.5-7b-L20-activations
GSM8K — Qwen2.5-7B-Instruct Layer-20 Activations
Residual-stream activations extracted from Qwen/Qwen2.5-7B-Instruct at layer 20 (of 28), last token of the input prompt, on the GSM8K test set (1319 examples).
Split into two files by whether the model answered correctly.
Files
File
Rows
Description
correct.parquet
708
Examples where the model's final answer matched the gold answer
incorrect.parquet
611
Examples where the model's final answer did not… See the full description on the dataset page: https://huggingface.co/datasets/Realmbird/gsm8k-qwen2.5-7b-L20-activations.
