CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ram-lexsi /curatorkit-testrun-Diversity curatorkit-testrun-Diversity Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-08-30 05:54 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Diversity", "alpaca") texttext-generationn<1K0 likes72 downloads26d agoHugging Face02Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face03sabin1234 /nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset 🧠 Nepali Psychology Question Dataset — 2,000 Samples 📌 Overview The Nepali Psychology Question Dataset is a specialized Nepali-language dataset containing 2,000 psychology-related question-answer records designed for Natural Language Processing (NLP), Large Language Models (LLMs), Small Language Models (SLMs), Supervised Fine-Tuning (SFT), Question Answering (QA), instruction tuning, educational AI, and psychology-domain research. The dataset is designed with a… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset.textquestion-answering1K<n<10K0 likes17 downloads1mo agoHugging Face04SkillFactory /SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_prompt_diversityYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files. "example_dataset": { "hf_hub_url": "SkillFactory/SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_prompt_diversity", "formatting": "sharegpt", "columns": { "messages": "conversations"}, "tags": { "user_tag": "user", "assistant_tag": "assistant", "role_tag": "role", "content_tag": "content" }, "subset": "sft_train" } texttext-generation1K<n<10K0 likes3 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.