datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coqar-SEA-answer-span-diversity
Fields
Top-level fields
id — Unique example identifier.
split — Original CoQAR split, e.g. dev.
conversation_id — Identifier of the source CoQAR conversation.
turn_id — Turn index within the conversation.
question — Original conversational question for the current turn.
history — Previous dialogue context stored as a flat sequence:
[question_0, answer_0, question_1, answer_1, ...].
all_standalone_questions — Human-written standalone rewrites of the current… See the full description on the dataset page: https://huggingface.co/datasets/zykov/coqar-SEA-answer-span-diversity.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset
🧠 Nepali Psychology Question Dataset — 2,000 Samples
📌 Overview
The Nepali Psychology Question Dataset is a specialized Nepali-language dataset containing 2,000 psychology-related question-answer records designed for Natural Language Processing (NLP), Large Language Models (LLMs), Small Language Models (SLMs), Supervised Fine-Tuning (SFT), Question Answering (QA), instruction tuning, educational AI, and psychology-domain research.
The dataset is designed with a… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset.Diversity_ChallengeWe designed a diversity task in which LLMs are prone to the repeat curse, and the repeat score effect is relatively pronounced on this dataset.
Diversity_channel_datasetgpt2-base-diversityReward-Model-Diversityshort_diversity_condition_rawshort_diversity_condition_dpo
