datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt_roleplay_realm
GPT Role-play Realm Dataset: The AI-generated character compendium
This is a dataset of GPT-generated characters made to increase the ability of open-source language models to role-play.
219 characters in the Russian part, and 216 characters in the English part. All character descriptions were generated with GPT-4.
20 dialogues on unique topics with every character. Topics were generated with GPT-4. The first dialogue out of 20 was also generated with GPT-4, and the other 19… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/gpt_roleplay_realm.LatamGPT-Corpus-1.0
LatamGPT-Corpus-1.0
🌐 Language versions: English | Español | Português
🔗 Project links: Official LatamGPT website | Corpus dashboard
🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0.
Dataset description
Summary
LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
Finch-Collection-GPT-5.4
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the GPT-5.4 teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with GPT-5.4 as the teacher mutation operator (the main… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-GPT-5.4.
