CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M84 likes992 downloads2y agoHugging Face02merve /turkish_instructionstext10K<n<100K64 likes701 downloads3y agoHugging Face03jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face04mikheevshow /SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructtextn<1K0 likes193 downloads11mo agoHugging Face05FinLang /investopedia-instruction-tuning-dataset Dataset Card for investopedia-instruction-tuning dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.text100K<n<1M23 likes183 downloads2y agoHugging Face06FINNUMBER /QA_Instruction 𓅰 FINCH: CoT-Instruction Dataset for Korean Finance 𓅰 Overview FINCH is a CoT-Instruction dataset rooting Korean-Financial tasks including: Multiple-Choice Question Answering (MCQA), Extractive Question Answering (EQA), Binary Question Answering (BQA), Numerical Reasoning, Tabular Reasoning and Sentiment Analysis. Additional details, research paper and further updates are coming! Stay Tuned. text10K<n<100K2 likes164 downloads3y agoHugging Face07md-nishat-008 /Bangla-Instruct Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.texttext-generation100K<n<1M8 likes143 downloads1y agoHugging Face08NebulaSense /Legal_Clause_Instructionstext1K<n<10K4 likes141 downloads3y agoHugging Face09jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes134 downloads5mo agoHugging Face10Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes133 downloads9mo agoHugging Face11mikheevshow /SIGNAL-Dataset-Hiddens-RefalMachine-RuadaptQwen2.5-7B-Instructtextn<1K0 likes120 downloads11mo agoHugging Face12alxfgh /ChEMBL_Drug_Instruction_Tuning Dataset Card for ChEMBL Drug Instruction Tuning Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.textquestion-answering100K<n<1M15 likes101 downloads3y agoHugging Face13suehuynh /marketing-instruct-4k Dataset Card for marketing-instruct-4k Dataset Details Dataset Description A curated instruction-tuning dataset of ~4,300 marketing copywriting examples across five task types, built for the AutoScientist Challenge 2026 (Marketing category). Used to fine-tune Marketing-Mixtral-8x7B. Key finding: this carefully curated dataset at its natural size outperformed a 12,000-row version expanded via automated augmentation (80% vs 58% win rate against the… See the full description on the dataset page: https://huggingface.co/datasets/suehuynh/marketing-instruct-4k.texttext-generation1K<n<10K0 likes95 downloads3mo agoHugging Face14AddisGPT /AddisGPT-Amharic-Instruction AddisGPT-Amharic-Instruction A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions. 796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.tabulartext-generationn<1K1 likes95 downloads21d agoHugging Face15nafisehNik /girt-instruct GIRT-Instruct Corpus Paper: https://arxiv.org/abs/2402.02632 A dataset in the format of pairs of instructions and corresponding outputs. GIRT-Instruct is constructed based on GIRT-Data, a dataset of IRTs. We use both GIRT-Data metadata and the Zephyr-7B-Beta language model to generate the instructions This dataset is used to train the GIRT-Model model. Model: model Space: space Type We have 4 different types in GIRT-Instruct. These types include: default: This type… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/girt-instruct.texttext-generation100K<n<1M2 likes94 downloads3y agoHugging Face16alxfgh /PubChem_Drug_Instruction_Tuningtext10K<n<100K11 likes83 downloads3y agoHugging Face17MBZUAI /instructpoet-ar Arabic Poetry IFT Dataset Summary Arabic Poetry IFT is a large-scale instruction-following dataset for Arabic poetry understanding and co-creation. It supports four task families: generation, continuation, revision/restoration, and multiple-choice analysis. The dataset covers Modern Standard Arabic (MSA) and four regional Arabic varieties used in the instruction layer: Gulf, Levantine, Nile Valley, and North African Arabic. This release accompanies the ACL 2026 paper… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/instructpoet-ar.texttext-generationn<1K0 likes74 downloads5mo agoHugging Face18red1xe /code_instructionstext10K<n<100K7 likes69 downloads3y agoHugging Face19jean1 /45k_python_code_chinese_instruction Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details 中文提示的代码数据集 其中提示部分通过调用GPT-4.0-turbo API翻译成中文 Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.text10K<n<100K6 likes69 downloads2y agoHugging Face20jamesdborin /Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only.tabular1K<n<10K0 likes69 downloads3mo agoHugging Face21lime-nlp /safer-instruct Safer-Instruct: Aligning Language Models with Automated Preference Data This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here! Abstract Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.text10K<n<100K1 likes65 downloads1y agoHugging Face22halilibr /collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources: Turkish Poems Cleaned Turkish Reading Comprehension Question Answering Dataset Stanford ALPaCA Cleaned Turkish Translated Turkish Poems Turkish Folk Song Lyrics The data has been merged and processed for quality and consistency to create this dataset. texttext-generation100K<n<1M11 likes64 downloads2y agoHugging Face23jamesdborin /Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.tabular1K<n<10K0 likes64 downloads3mo agoHugging Face24jamesdborin /Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.tabular1K<n<10K0 likes63 downloads3mo agoHugging Face25Raftico /instructional-dialogues-multilingual Multilingual Instructional Dialogues (10-Language Dataset) Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models. Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.text1K<n<10K2 likes62 downloads1y agoHugging Face26jamesdborin /Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.tabular1K<n<10K0 likes62 downloads3mo agoHugging Face2704RR /tiny-instruct tiny-instruct-v1 This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion). Code Datasets: CodeAlpaca_20K CodeExercise-Python-27k Evol-Instruct-Code-80k-v1 tiny-codes Evol-instruction-66k sciphi-python-textbook programming_books_llama WizardLM_evol_instruct_70k Math Datasets: MetaMathQA arxiv-math-instruct-50k MathInstruct General… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.texttext-generation1M<n<10M16 likes60 downloads3y agoHugging Face28mikheevshow /SIGNAL-Dataset-Hiddens-RefalMachine-RuadaptQwen2.5-14B-Instructtextn<1K0 likes57 downloads11mo agoHugging Face29erayalp /turkish-reasoning-instructionstext10K<n<100K6 likes55 downloads2y agoHugging Face30crosslingual-em /Qwen2.5-7B-Instruct-em-evaldocumentn<1K0 likes51 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.