datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LusoSupport-PT
LusoSupport-PT Lite
Project evolution: LusoSupport-PT launched as a standalone pt-PT instruction dataset. Since AMALIA (Portugal's open-source, government-backed pt-PT LLM) launched in July 2026, the project has repositioned around specializing AMALIA for customer support — this dataset is preserved and still available, now as the training material behind that specialization. See the pivot strategy doc for the full rationale.
🇵🇹 AMALIA already speaks fluent European… See the full description on the dataset page: https://huggingface.co/datasets/ariazevedo/LusoSupport-PT.stable_diffusion_instructional_dataset
Stable Diffusion Dataset
Description:
This dataset is in Jsonl format and is based on the MadVoyager/stable_diffusion_instructional_dataset.
Overview:
The Stable Diffusion Dataset comprises approximately 80,000 meticulously curated prompts sourced from the image finder of Stable Diffusion: "Lexica.art". The dataset is intended to facilitate training and fine-tuning of various language models, including LLaMa2.
Key Features:
◉ Jsonl format for… See the full description on the dataset page: https://huggingface.co/datasets/lusstta/stable_diffusion_instructional_dataset.HC3-ChineseHuman ChatGPT Comparison Corpus (HC3) Chinese Versionfashion_questions_answersSt.Clair_Programs
