datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Muse-Glimmer-OPB-100K
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Reasoning strength
Conversations
Train-turn rows
low
64,997
96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.corpus_dominio_museistico_patrimonio
Corpus museístico-patrimonio
Descripción general
El corpus museístico-patrimonio reúne recursos especializados del ámbito museístico y patrimonial, incluyendo tesauros terminológicos, catálogos museísticos y colecciones descriptivas vinculadas al patrimonio cultural. El conjunto representa un registro técnico y descriptivo propio de la documentación patrimonial, la catalogación de bienes culturales y la organización conceptual del conocimiento museístico.
El… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/corpus_dominio_museistico_patrimonio.
