datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
per-ma-to
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web data.… See the full description on the dataset page: https://huggingface.co/datasets/RolandP/per-ma-to.ScaleBio-Baseline-llama3-less-code-10kScaleBio-Baseline-gemma2-less-code-10kScaleBio-Baseline-qwen2-aioliScaleBio-Baseline-gemma2-less-math-5kScaleBio-Baseline-gemma2-aioliScaleBio-Baseline-gemma2-less-code-5kEnvFactory-SFT-Skippeddouyinship-qq1314ScaleBio-Baseline-llama3-aioliScaleBio-Baseline-qwen2-less-code-5kScaleBio-Baseline-qwen2-less-code-10kduanship111-ttEnvFactory-SFT-NIPSduanshipin-qq521ScaleBio-Baseline-llama3-less-code-5k
