CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes7.2k downloads3y agoHugging Face02tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes768 downloads7mo agoHugging Face03inaciose /AMALIA-LLM-0626-SFT-Dataset-classlangcode AMALIA-LLM-0626-SFT-Dataset-classlangcode Overview Este dataset é baseado no amalia-llm/AMALIA-LLM-0626-SFT-Dataset, mantendo integralmente a estrutura das conversações e adicionando metadados obtidos através de classificação automática. O objetivo desta versão é facilitar a seleção, filtragem e construção de subconjuntos especializados para treino e avaliação de modelos de linguagem, sem alterar o conteúdo original das conversações. O dataset original foi… See the full description on the dataset page: https://huggingface.co/datasets/inaciose/AMALIA-LLM-0626-SFT-Dataset-classlangcode.text-classification1M<n<10M0 likes383 downloads2mo agoHugging Face04bakrianoo /jabarti-llm-dataset jabarti-llm-dataset Cleaned, section-chunked training corpus for a small bilingual LLM (Arabic + English), combining a curated Egyptian-history collection with general Wikipedia coverage from CohereLabs/wikipedia-2023-11-embed-multilingual-v3. Every pretrain record is a contiguous span of 120-1500 characters with the article title and section headings removed. Provenance is in ds_source. Configs and Splits Config Split Rows Training phase Purpose… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.tabulartext-generation1M<n<10M1 likes169 downloads3d agoHugging Face05ghanaopenai /twi-llm-reasoning-dataset-1k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Twi Reasoning Dataset A Twi (Akan) translation of the Multilingual-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-llm-reasoning-dataset-1k.texttext-generationn<1K7 likes105 downloads3mo agoHugging Face06VN-26 /rtl-llm-dataset RTL Design & Verification Dataset A curated collection of Verilog and SystemVerilog designs paired with corresponding testbenches. This dataset is designed for fine-tuning Large Language Models (LLMs) on hardware description languages (HDL) and verification tasks. Dataset Structure Each entry in the .jsonl file follows this schema: sample_id: Unique identifier derived from the source folder name. metadata: title: Human-readable name of the module. day: The original… See the full description on the dataset page: https://huggingface.co/datasets/VN-26/rtl-llm-dataset.text-generation0 likes104 downloads6mo agoHugging Face07JasonShiii /STEP-LLM-dataset STEP-LLM Dataset Rendered multi-view images of ABC Dataset STEP files, used in our DATE 2026 paper: "STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models" Dataset Contents Rendered JPEG images of CAD models from the ABC Dataset (NYU), organized by entity count. These images were used with GPT-4o to generate natural language captions for training STEP-LLM. Folder Entity range # Models Description step_under500_image/ 0–500… See the full description on the dataset page: https://huggingface.co/datasets/JasonShiii/STEP-LLM-dataset.text-generation10K<n<100K2 likes88 downloads7mo agoHugging Face08rriviere /oc-llm-finetuning-dataset Dataset médical bilingue, triage CHSA Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms, AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé, instruction/réponse) et dpo (alignement par préférences, chosen/rejected). Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques. Schéma Champs communs à tous les exemples : Champ Type Description id string Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.texttext-generation100K<n<1M0 likes70 downloads8d agoHugging Face09amalia-llm /AMALIA-LLM-0626-SFT-Dataset AMALIA LLM Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training. Base Data Mix This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count amalia-llm/persona_math 63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.texttext-generation1M<n<10M2 likes68 downloads3mo agoHugging Face10sidfeels /visually-impaired-llm-assistance-dataset Visually Impaired Assistance Dataset Hey everyone! I'm currently working on a project to finetune an llm for assisting people with bad eyesights in daily life scenario so for it I had to create a synthetic dataset and here is the Visually Impaired Assistance Dataset that I created. This dataset provides step-by-step instructions with non-visual cues for a variety of daily tasks and activities, specifically designed to help visually impaired individuals. It covers a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/sidfeels/visually-impaired-llm-assistance-dataset.textquestion-answering10K<n<100K3 likes61 downloads2y agoHugging Face11ctaxnagomi /bintulu-llm-dataset Bintulu LLM Dataset A text-to-text (prompt → answer) corpus about Bintulu, Sarawak — its national parks, beaches, energy industry, history, food, transport and guesthouse hospitality. Built for two uses: Fine-tuning a text-to-text LLM (instruction, QA, dialogue, translation, NLI, MCQ). Embedding/retrieval — every row's text field (canonically equal to input_text) is the embeddable surface for a vector store. Schema Each JSONL row: field type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/bintulu-llm-dataset.texttext-generation1K<n<10K0 likes61 downloads20d agoHugging Face12TFMC /imatrix-dataset-for-japanese-llmtexttext-generationn<1K35 likes58 downloads2y agoHugging Face13dworsleytonks /medical-llm-finetuning-alignment-processed-datasettexttext-generation10K<n<100K0 likes56 downloads9mo agoHugging Face14Timofey /protein_interactions_LLM_FT_datasetThis dataset is derived from the ANDDigest database and contains PubMed abstracts with dictionary-mapped protein names. A total of >15,000 abstracts were selected, yielding 6,516 unique protein pairs with associative edges in the ANDSystem network. The corpus was split into positive and negative samples. Positive samples represent documents where a protein interaction was identified using the rules-based approach of ANDSystem’s text-mining module, while negative samples include documents where… See the full description on the dataset page: https://huggingface.co/datasets/Timofey/protein_interactions_LLM_FT_dataset.textsummarization10K<n<100K2 likes52 downloads2y agoHugging Face15xTayyub /High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning PyReason-7k: Advanced Python Chain-of-Thought Dataset Dataset Description This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning. Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding. Key Features: Chain-of-Thought: Step-by-step reasoning traces. Error Handling: Solutions include try-except blocks and logging. Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.text-generation1K<n<10K7 likes39 downloads10mo agoHugging Face16Capibara-LLM /dataset-guarani-jopara-v01 Dataset Guaraní-Jopara v01 (Alpaca Style) 🇵🇾 Resumen del Dataset Este dataset contiene pares de instrucción-respuesta diseñados para el entrenamiento de ajuste fino (Fine-Tuning) de Modelos de Lenguaje Grande (LLMs) en el idioma Guaraní, específicamente en su variante Jopara (la mezcla coloquial de Guaraní y Español hablada comúnmente en Paraguay). El formato sigue la estructura estándar de Alpaca, lo que lo hace compatible con la mayoría de las librerías de… See the full description on the dataset page: https://huggingface.co/datasets/Capibara-LLM/dataset-guarani-jopara-v01.texttext-generation10K<n<100K0 likes37 downloads10mo agoHugging Face17amalia-llm /DPO-Dataset AMALIA DPO Dataset This is the DPO (preference optimization) dataset used to train AMALIA-DPO. It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets. This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.tabularquestion-answering100K<n<1M4 likes37 downloads3mo agoHugging Face18amkyawdev /mm-llm-coder-dataset 🇲🇲 Myanmar LLM Coder Dataset (mm-llm-coder-dataset) မြန်မာဘာသာ Coding LLM များ training အတွက် ရည်ရွယ်ထားသော dataset A bilingual (Myanmar + English) coding instruction dataset designed primarily for training Myanmar language Coder LLMs. 🎯 ရည်ရွယ်ချက် / Purpose ဤ dataset သည် မြန်မာဘာသာ programming/coding LLM များ training လုပ်ရန်အတွက် အဓိက ရည်ရွယ်ထားပါသည်။ မြန်မာ developer များ၏ မိခင်ဘာသာစကားဖြင့် coding အကူအညီပေးနိုင်သော AI assistant များကို ဖန်တီးနိုင်စေရန်… See the full description on the dataset page: https://huggingface.co/datasets/amkyawdev/mm-llm-coder-dataset.texttext-generation1M<n<10M0 likes35 downloads5mo agoHugging Face19UniDataPro /llm-training-dataset LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data Models used for text generation: GPT-3.5 GPT-4 Uncensored GPT Version (is not included inthe sample) Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.texttext-generation1K<n<10K3 likes31 downloads1mo agoHugging Face20llm-compe-2025-kato /step2-evaluated-dataset-test2 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-test2 Total Samples: 92 Successfully Evaluated (Rubric): 92 Failed Evaluations (Rubric): 0 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence: 3.51… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-test2.tabulartext-generationn<1K0 likes30 downloads1y agoHugging Face21amalia-llm /AMALIA-LLM-1225-SFT-Dataset AMALIA LLM 1225 Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model version released in December 2025. This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count Persona-PT Instruction Following 9,084 Persona-EN Instruction Following 34,704 Persona Nemotron Instruction Following 4,483… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-1225-SFT-Dataset.texttext-generation1M<n<10M0 likes29 downloads3mo agoHugging Face22ArtemLykov /LLM_BRAIn_datasetLLM_BRAIn: AI-driven Fast Generation of Robot Behaviour Tree based on Large Language Model Original paper preprint: https://arxiv.org/abs/2305.19352 This paper introduces a pioneering methodology in autonomous robot control, denoted as LLM-BRAIn, enabling the generation of adaptive behaviors in robots in response to operator commands, while simultaneously considering a multitude of potential future events. LLM-BRAIn is a transformer-based Large Language Model (LLM) fine-tuned from the… See the full description on the dataset page: https://huggingface.co/datasets/ArtemLykov/LLM_BRAIn_dataset.textrobotics1K<n<10K4 likes28 downloads2y agoHugging Face23llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp32 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32 Total Samples: 60 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 7 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32.tabulartext-generationn<1K0 likes25 downloads1y agoHugging Face24rakhman-llm /XCodeEval-Java-Datasettexttext-generation100K<n<1M0 likes24 downloads2y agoHugging Face25ud-nlp /LLM-Text-Generation-Dataset Generated Text Dataset - 4 Millions+ Logs Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data Dataset characteristics: Characteristic Data Description Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.texttext-generation1K<n<10K0 likes24 downloads1y agoHugging Face26llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp40 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40 Total Samples: 58 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 5 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40.tabulartext-generationn<1K0 likes23 downloads1y agoHugging Face27weblab-llm-competition-2025-bridge /team-truthowl-mixed-reasoning-dataset Team P11 Mixed Reasoning Dataset 📊 Dataset description HLE(Humanity's Last Exam)向けに作成した、数学中心+科学MCの混合推論データセットです。 推論過程(Chain-of-Thought)を保持し、最終解答の正規化を行っています。 対象モデルは DeepSeek-R1-Distill-Qwen-32B、学習はQLoRAを想定しています。 🎯 Purpose Competition: 松尾研LLMコンペ 2025 Target Model: DeepSeek-R1-Distill-Qwen-32B Training Method: QLoRA Fine-tuning(4bit NF4, double quant) 📦 Composition Math Hard(MATH Level≥3, HARDMath) Math Mid(GSM8K, MetaMathQA) Science(GPQA… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/team-truthowl-mixed-reasoning-dataset.texttext-generation10K<n<100K0 likes23 downloads11mo agoHugging Face28UniqueData /llm-dataset LLM Dataset - Prompts and Generated Texts The dataset contains prompts and texts generated by the Large Language Models (LLMs) in 32 different languages. The prompts are short sentences or phrases for the model to generate text. The texts generated by the LLM are responses to these prompts and can vary in length and complexity. Researchers and developers can use this dataset to train and fine-tune their own language models for multilingual applications. The dataset provides a rich… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/llm-dataset.texttext-generation1K<n<10K5 likes20 downloads1y agoHugging Face29llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B Total Samples: 156 Successfully Evaluated (Rubric): 135 Failed Evaluations (Rubric): 21 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B.tabulartext-generationn<1K0 likes20 downloads1y agoHugging Face30VedCodes /New_dataset_llmtexttext-generationn<1K0 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.