CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K170 likes62k downloads3y agoHugging Face02BAAI /Infinity-Instructgated Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and high-quality instruction… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Instruct.tabulartext-generation10M<n<100M765 likes3.3k downloads10mo agoHugging Face03BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.4k downloads9mo agoHugging Face04sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M84 likes980 downloads2y agoHugging Face05matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes644 downloads3y agoHugging Face06Parssky /industrial-instruction-dataset Industrial-Instruction Dataset Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings. Paper Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.tabularquestion-answering10K<n<100K0 likes591 downloads29d agoHugging Face07JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes518 downloads8d agoHugging Face08matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes456 downloads3y agoHugging Face09DevShubham /python-text-training-instruct-ai Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/DevShubham/python-text-training-instruct-ai.tabulartext-generation1K<n<10K1 likes382 downloads2y agoHugging Face10proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes361 downloads5mo agoHugging Face11aisingapore /SEA-Instruct-2602gated SEA-Instruct-2602 Overview SEA-Instruct-2602 is a preliminary release of instruction-tuning data focused on Southeast Asian languages and contexts. The dataset combines prompts filtered from open-source data with our own synthetic prompts, paired with synthetic responses, for language model training on SEA-specific tasks and languages. This dataset contains only the filtered subset of data with prompt_input_quality at Excellent, prompt_is_coherent at True and… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602.tabulartext-generation1M<n<10M3 likes345 downloads8mo agoHugging Face12daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes289 downloads13d agoHugging Face13wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes255 downloads2y agoHugging Face14JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes253 downloads8d agoHugging Face15matlok /python-text-copilot-training-instruct-ai-research Building an AI Copilot Dataset to help keep up with Leading AI Research This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples). This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.tabulartext-generation10K<n<100K0 likes240 downloads3y agoHugging Face16ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes230 downloads3y agoHugging Face17matlok /python-text-copilot-training-instruct-ai-research-2024-02-10 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.tabulartext-generationn<1K0 likes214 downloads3y agoHugging Face18matlok /python-text-copilot-training-instruct-ai-research-2024-02-11 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.tabulartext-generationn<1K0 likes201 downloads3y agoHugging Face19viyer98 /xl-instruct Dataset Card for XL-Instruct This dataset card provides a summary of the XL-Instruct dataset, a resource for advancing the cross-lingual capabilities of Large Language Models. It was introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation. Dataset Details Dataset Description XL-Instruct is a high-quality, large-scale synthetic dataset designed to fine-tune LLMs for cross-lingual open-ended generation. The core task involves… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/xl-instruct.tabulartext-generation100K<n<1M2 likes167 downloads1y agoHugging Face20xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes154 downloads1mo agoHugging Face21agagasf123123 /threejs-gamecode-instruct-v3-ultra Three.js GameCode Instruct v3 Ultra This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding. Important note This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark. No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.tabulartext-generation10K<n<100K2 likes153 downloads3mo agoHugging Face22ctaxnagomi /INSTRUCT_JEV INSTRUCT_JEV INSTRUCT_JEV is an instruction corpus built from the TypeSafe AI documentation for Jev, the first System One model. It is structured around the three TypeSafe question primitives - Choice, Noul and Score - and mirrors the raw corpus captured in deckerGUI-jev_corpus_RAW. Credits INSTRUCT_JEV is a DeckerGUI project and exists because of the work below. Who Contribution Link TypeSafe AI Jev - the first System One model - and the Choice / Noul… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV.tabulartext-generationn<1K2 likes126 downloads3d agoHugging Face23HINT-lab /Qwen2.5-7B-Instruct-Self-Calibration Efficient Test-Time Scaling via Self-Calibration This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.tabulartext-generation100K<n<1M0 likes111 downloads2y agoHugging Face24SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes108 downloads23d agoHugging Face25Yooniel /qwen2.5-7b-instruct-nla-L20-finefineweb-100k Qwen2.5-7B-Instruct NLA training data — residual stream, block 20 Training data for a Natural Language Autoencoder on Qwen/Qwen2.5-7B-Instruct: residual-stream activations paired with natural-language explanations of the text they were taken from. Unlike the dataset this is derived from, the activation_vector column is included — every EasyNLA/nanoNLA trainer requires it. Trained models: https://huggingface.co/Yooniel/qwen2.5-7b-instruct-nla-L20 (AV val ppl 4.07, AR held-out FVE… See the full description on the dataset page: https://huggingface.co/datasets/Yooniel/qwen2.5-7b-instruct-nla-L20-finefineweb-100k.tabulartext-generation100K<n<1M0 likes107 downloads2mo agoHugging Face26matlok /python-text-copilot-training-instruct-ai-research-2024-01-27 Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is the 2024-01-27 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-01-27.tabulartext-generation10K<n<100K0 likes102 downloads3y agoHugging Face27AddisGPT /AddisGPT-Amharic-Instruction AddisGPT-Amharic-Instruction A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions. 796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.tabulartext-generationn<1K1 likes96 downloads23d agoHugging Face28anakin87 /fine-instructions-ita-70k 🍷🇮🇹 Fine Instructions Ita 70k A good instruction dataset in Italian, generated with LLM-aided translation. LLM-aided translation The starting point has been mlabonne/FineTome-100k: a subset of arcee-ai/The-Tome, re-filtered using HuggingFaceFW/fineweb-edu-classifier to identify examples with good educational value. FineTome is a high quality dataset containing diverse elements such as conversations, reasoning problems, and more. The translation process includes these… See the full description on the dataset page: https://huggingface.co/datasets/anakin87/fine-instructions-ita-70k.tabulartext-generation10K<n<100K4 likes95 downloads2y agoHugging Face29werty1248 /multilingual-instruct-balancedThis repository is a collection of English, Korean, Chinese, and Japanese datasets collected by the HuggingFace Hub and transformed into a unified format. It consists of either native or synthetic data. Some data is not clearly copyrighted or only allows non-commercial use. Preprocessing: I removed data with too few answer tokens or more than 8192 tokens, and removed synthetic data with repetitions. Balancing: I randomly sampled a subset of the data with different weights for each language and… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/multilingual-instruct-balanced.tabulartext-generation1M<n<10M2 likes91 downloads2y agoHugging Face30Pondsiders /tinystories-gpt4-instruct tinystories-gpt4-instruct Request→story pairs for supervised fine-tuning of small language models, derived from karpathy/tinystories-gpt4-clean. Each example pairs a natural-language request ("Can you tell me a story about a boy named Tim?") with a TinyStories story that satisfies it. The dataset lives on Hugging Face; the notebook that generates it lives on GitHub. This is not roneneldan/TinyStoriesInstruct. That dataset frames its tasks in a structured format (Words:… See the full description on the dataset page: https://huggingface.co/datasets/Pondsiders/tinystories-gpt4-instruct.tabulartext-generation10K<n<100K0 likes82 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.