datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Factuality_Alignment
Factual Preference Alignment Dataset
**⚠️ Warning:**This dataset contains hallucinated and synthetic responses
intentionally generated for research on robust factuality alignment.
Responses may include fabricated or incorrect information by design
to support the evaluation of hallucination-aware learning.
Dataset Summary
The AIXpert Preference Alignment Dataset is a curated collection of
45,000 factuality-aware preference pairs designed to support
research on Modified… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/Factuality_Alignment.factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.factualqwq_32b_factualqa_sft_dataadaption-cordel-factual-nordeste
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-cordel_factual_nordeste
This dataset contains pairs of prompts and completions where Brazilian 'cordel' poetry is generated in sextet stanzas based strictly on provided factual texts about Northeastern Brazilian culture, history, and geography. The source texts cover topics such as Frevo, the Cangaço (Lampião and Maria Bonita), the Caruaru Fair, and the Caatinga biome, with… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-cordel-factual-nordeste.
