CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01merve /turkish_instructionstext10K<n<100K64 likes691 downloads3y agoHugging Face02FinLang /investopedia-instruction-tuning-dataset Dataset Card for investopedia-instruction-tuning dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.text100K<n<1M23 likes199 downloads2y agoHugging Face03jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.tabular1M<n<10M0 likes182 downloads3mo agoHugging Face04FINNUMBER /QA_Instruction 𓅰 FINCH: CoT-Instruction Dataset for Korean Finance 𓅰 Overview FINCH is a CoT-Instruction dataset rooting Korean-Financial tasks including: Multiple-Choice Question Answering (MCQA), Extractive Question Answering (EQA), Binary Question Answering (BQA), Numerical Reasoning, Tabular Reasoning and Sentiment Analysis. Additional details, research paper and further updates are coming! Stay Tuned. text10K<n<100K2 likes159 downloads3y agoHugging Face05jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes143 downloads5mo agoHugging Face06NebulaSense /Legal_Clause_Instructionstext1K<n<10K4 likes137 downloads3y agoHugging Face07Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes124 downloads9mo agoHugging Face08alxfgh /ChEMBL_Drug_Instruction_Tuning Dataset Card for ChEMBL Drug Instruction Tuning Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.textquestion-answering100K<n<1M15 likes115 downloads3y agoHugging Face09AddisGPT /AddisGPT-Amharic-Instruction AddisGPT-Amharic-Instruction A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions. 796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.tabulartext-generationn<1K1 likes100 downloads23d agoHugging Face10alxfgh /PubChem_Drug_Instruction_Tuningtext10K<n<100K11 likes93 downloads3y agoHugging Face11jean1 /45k_python_code_chinese_instruction Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details 中文提示的代码数据集 其中提示部分通过调用GPT-4.0-turbo API翻译成中文 Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.text10K<n<100K6 likes64 downloads2y agoHugging Face12halilibr /collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources: Turkish Poems Cleaned Turkish Reading Comprehension Question Answering Dataset Stanford ALPaCA Cleaned Turkish Translated Turkish Poems Turkish Folk Song Lyrics The data has been merged and processed for quality and consistency to create this dataset. texttext-generation100K<n<1M11 likes63 downloads2y agoHugging Face13Raftico /instructional-dialogues-multilingual Multilingual Instructional Dialogues (10-Language Dataset) Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models. Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.text1K<n<10K2 likes62 downloads1y agoHugging Face14red1xe /code_instructionstext10K<n<100K7 likes61 downloads3y agoHugging Face15jamesdborin /Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.tabular1K<n<10K0 likes59 downloads3mo agoHugging Face16jamesdborin /Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.tabular1K<n<10K0 likes58 downloads3mo agoHugging Face17jamesdborin /Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only.tabular1K<n<10K0 likes57 downloads3mo agoHugging Face18jamesdborin /Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.tabular1K<n<10K0 likes55 downloads3mo agoHugging Face19Shiveswarran /llm_instruction_code_manual_yolo_lctextn<1K2 likes50 downloads3y agoHugging Face20erayalp /turkish-reasoning-instructionstext10K<n<100K6 likes50 downloads2y agoHugging Face21jamesdborin /Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only.tabular1K<n<10K0 likes41 downloads3mo agoHugging Face22skyylord /mitre-attack-ttp-labeled-instructions MITRE ATT&CK TTP Mapping Dataset Training and evaluation data for mapping adversarial behavior descriptions (CTI reports, CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.… See the full description on the dataset page: https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions.texttext-classification10K<n<100K0 likes39 downloads3d agoHugging Face23prithivMLmods /Step-Instruction-Gx Step-Instruction-GX Dataset Overview The Step-Instruction-GX dataset is a collection of instructional and educational content designed to assist in various learning and decision-making tasks. It includes a wide range of questions and corresponding answers, covering topics from health tips to scientific concepts. Dataset Details Modalities Text: The dataset primarily contains text data in various formats. Formats CSV: The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Step-Instruction-Gx.texttext-generation10K<n<100K4 likes38 downloads2y agoHugging Face24tanmaylaud /scidcc-instructions Dataset Summary Instruction-Response pairs generated using the SciDCC Climate Dataset from Climabench Format ### Instruction: Present a fitting title for the provided text. For those who study earthquakes, one major challenge has been trying to understand all the physics of a fault -- both during an earthquake and at times of "rest" -- in order to know more about how a particular region may behave in the future. Now, researchers at the California Institute of Technology… See the full description on the dataset page: https://huggingface.co/datasets/tanmaylaud/scidcc-instructions.textsummarization10K<n<100K4 likes37 downloads3y agoHugging Face25sambanankhu /Instruction-public-health-datasettext1K<n<10K0 likes37 downloads2y agoHugging Face26BlackKakapo /instruction-dataset-roOriginal dataset - This dataset is just the translation of the instruction-dataset dataset. textquestion-answeringn<1K0 likes36 downloads3y agoHugging Face27DebasishDhal99 /punjabi-instruction-datasetSource of the components that form this dataset: - AI4Bharat Question-Answering (99K rows) https://huggingface.co/datasets/ai4bharat/IndicQuestionGeneration/viewer/pa AI4Bhraat Headline-Generation (60K rows) https://huggingface.co/datasets/ai4bharat/IndicQuestionGeneration/viewer/pa AI4Bharat Sentence-Summarization (58K rows) https://huggingface.co/datasets/ai4bharat/IndicSentenceSummarization/viewer/pa HydraIndicLLM Alpaca (Derived dataset) (52K rows)… See the full description on the dataset page: https://huggingface.co/datasets/DebasishDhal99/punjabi-instruction-dataset.textquestion-answering100K<n<1M1 likes36 downloads3y agoHugging Face28m-a-d-i /wori-wolof-instructions WORI — Wolof Reverse Instruction Dataset WORI (Wolof Reverse Instruction) is a linguistically validated instruction-tuning dataset for Wolof, a low-resource language. The dataset provides 3,724 unique instruction-output pairs in Wolof, with parallel French translations. It was constructed via a reverse instruction pipeline and validated through a combination of automated language identification and manual review. For full methodological details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-d-i/wori-wolof-instructions.texttext-generation1K<n<10K1 likes36 downloads4mo agoHugging Face29taskydata /Pile-T5-Instruction_updatedtext10K<n<100K0 likes35 downloads2y agoHugging Face30kaxap /pg-wikiSQL-sql-instructions-80kConverted, cleaned and syntax-checked SQLWiki dataset. The datapoints containing non latin column names were removed. Resulting SQL statements were adapted for Postgres syntax and conventions. Each SQL statement, including CREATE TABLE statements were syntax checked with pgsanity. Citations @article{zhongSeq2SQL2017, author = {Victor Zhong and Caiming Xiong and Richard Socher}, title = {Seq2SQL: Generating Structured Queries from Natural… See the full description on the dataset page: https://huggingface.co/datasets/kaxap/pg-wikiSQL-sql-instructions-80k.text10K<n<100K10 likes33 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.