datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tosti-qwen-image-lora-datasetkey-data
🔑 KEY: Neuroevolution Dataset
40,000+ logged events from real evolutionary runs — every mutation, crossover, selection, and fitness evaluation.
KEY evolves LoRA adapters on frozen base models (MiniLM-L6, DreamerV3) using NEAT-style neuroevolution. This dataset captures the complete evolutionary history.
🎮 Links
🌌 Live Demo
Watch evolution in action
🧠 Champion Model
The evolved DreamerV3 model
Loading the Dataset
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/tostido/key-data.kontext-tosti-lora-datasettosti-qwen-loratosti-v2-qwen-image-datasettoxic-dpo-v0.2-dutch
Toxic DPO v0.2 - Dutch Translation
Dataset Description
This is a direct machine-translated Dutch version of the original datasetunalignment/toxic-dpo-v0.2.
Translation method:English → Dutch using the Helsinki-NLP/opus-mt-en-nl model from the MarianMTModel translations.No manual edits, additions or filtering were applied besides the automated translation.
Data set is checked on NULL values and duplicates.
All fields (prompt, chosen, rejected) were translated… See the full description on the dataset page: https://huggingface.co/datasets/tostideluxekaas/toxic-dpo-v0.2-dutch.amalgam-ledgerflux-tosti-datasettwitter-_tost_makinesi-2026.02.27-2027323761285095767-atVwjL256v7oI_He-part1butterfly-vocabulary ---
license: cc0-1.0
pretty_name: Butterfly Vocabulary
tags:
- butterfly
brotology
caveman-science
astrophysics
kids
task_categories:
text-generation
feature-extraction
size_categories:
10K<n<100K
---
# Butterfly Vocabulary
> Innate, nuclear, and distilled vocabularies that bootstrap the Butterfly Convergence Engine's language layer.
Produced by: generate_innate_vocab.py / merge_nuclear_vocab.py / distill_vocabulary.py in the Convergence Engine.
## Hey caveman… See the full description on the dataset page: https://huggingface.co/datasets/tostido/butterfly-vocabulary.butterfly-knowledge-web ---
license: cc0-1.0
pretty_name: Butterfly Knowledge Web
tags:
- butterfly
brotology
caveman-science
astrophysics
kids
task_categories:
graph-ml
size_categories:
100K<n<1M
---
# Butterfly Knowledge Web
> Seeded and expanded knowledge webs (concept graphs) used by the Convergence Engine's reasoning layer.
Produced by: reality_simulator/language/expand_knowledge_web.py in the Convergence Engine.
## Hey caveman scientist / astrophysicist
This repo is part of the… See the full description on the dataset page: https://huggingface.co/datasets/tostido/butterfly-knowledge-web.transagent-sliding-window-sfttransagent-sliding-window-grpotoxic_sft_dutch
Data dict
This dataset is a direct translation from francoj/toxic_sum_zh_sft.json, translated using AI. The original language is zh (chinese) and it is automatically translated from zh -> English -> Dutch.
It was used for the GEITje-7b-uncensored fine-tuning (filtered from this). And was cleaned so that instructions = messages and fixed some issues with labeling (roles), which probably started from translation.
DISCLAIMER
This dataset is fairly extreme in topics, and to… See the full description on the dataset page: https://huggingface.co/datasets/tostideluxekaas/toxic_sft_dutch.tostiok-lorabutterfly-curated-dataset ---
license: cc0-1.0
pretty_name: Butterfly Curated Dataset
tags:
- butterfly
brotology
caveman-science
astrophysics
kids
task_categories:
text-generation
size_categories:
100K<n<1M
---
# Butterfly Curated Dataset
> Curated training dataset built from public-domain sources; the language food for Butterfly organisms.
Produced by: build_curated_dataset.py in the Convergence Engine.
## Hey caveman scientist / astrophysicist
This repo is part of the Butterfly Field… See the full description on the dataset page: https://huggingface.co/datasets/tostido/butterfly-curated-dataset.phosphobot_v4transagent-sftphoto-to-tosti
