CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Viet-Mistral /CulturaY CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.texttext-generation1B<n<10B39 likes6.5k downloads2y agoHugging Face02danlou /based-chat-v0.1-Mistral-Nemo-Base-2407 Based-Chat v0.1 (Mistral Nemo Base 2407) This dataset was developed as part of an exploration into understanding the necessity of supervised datasets for fine-tuning base LLMs into conversational models. It's a synthetic dataset created with Mistral-Nemo-Base-2407, and used to fine-tune that model, producing relay-v0.1-Mistral-Nemo-2407. Methodology This synthetic dataset is generated using the following as conversation starters: facebook/empathetic_dialogues… See the full description on the dataset page: https://huggingface.co/datasets/danlou/based-chat-v0.1-Mistral-Nemo-Base-2407.texttext-generation100K<n<1M1 likes59 downloads2y agoHugging Face03mistral-hackaton-2026 /robuchan-data Robuchan Dataset Synthetic dietary recipe adaptation dataset for fine-tuning language models. Each example is a chat-format conversation where a user provides a recipe and dietary restriction, and the assistant produces a structured adaptation. Generated for the Mistral AI Worldwide Hackathon Tokyo (Feb 28 - Mar 1, 2026). Associated model: sumitdotml/robuchan Dataset Structure Splits Split Rows Purpose train 1,090 Fine-tuning training set… See the full description on the dataset page: https://huggingface.co/datasets/mistral-hackaton-2026/robuchan-data.texttext-generation1K<n<10K0 likes41 downloads7mo agoHugging Face04davidpistori /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.texttext-generation10K<n<100K0 likes37 downloads3mo agoHugging Face05bala1524 /Medical-QA-Mistral7B-Finetuningtextquestion-answeringn<1K6 likes30 downloads3y agoHugging Face06VinceGx33 /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.texttext-generation10K<n<100K2 likes30 downloads11mo agoHugging Face07DataPilot /Zero_SFT_Ja_by_Mistral_Small DataPilot/Zero_SFT_Ja_by_Mistral_Small このデータセットは、日本語で記述された高品質な合成プロンプトとそのAI出力を収録しています。すべてのデータは Mistral Small 3.1 24B Instruct 2503 モデルを使用してゼロから合成されています。 概要 項目 詳細 データセット名 DataPilot/Zero_SFT_Ja_by_Mistral_Small 言語 日本語 データ作成方法 完全自動生成(モデルによるゼロショット合成) 使用モデル Mistral Small 3.1 24B Instruct 2503 フォーマット JSONL(id, input, output, conversation) ライセンス Apache-2.0 作成コード foxn2000/zero_one_instruction データセット構造 データセットには以下のカラムが含まれています。 カラム名… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Zero_SFT_Ja_by_Mistral_Small.texttext-generation1K<n<10K1 likes17 downloads1y agoHugging Face08HachiML /JMT-Bench-result_self-rewarding_Mistral-7B-lora JMT-Bench result Answer language JMT-Benchの回答のうち、Englishで回答した件数 Model Count mistralai/Mistral-7B-v0.3 25 HachiML/Mistral-7B-v0.3-m1-lora 7 HachiML/Mistral-7B-v0.3-m2-lora 7 HachiML/Mistral-7B-v0.3-m3-lora 2 tabulartext-generationn<1K0 likes14 downloads2y agoHugging Face09dpevzner /Cybersecurity_Reasoning_Dataset_MistralFamily_7bgated Cybersecurity Reasoning Dataset (v6.0) A high-fidelity, forensics-mapped training corpus for cybersecurity reasoning and SOC automation — formatted for the Mistral / Llama instruct family. Format-specific dataset. Every record uses the Alpaca-style ### Instruction: / ### Response: template native to Mistral/Llama instruct models. A model-agnostic version of this corpus (Mistral, DeepSeek, ChatML, and Gemma variants rendered from one neutral canonical source) is published… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset_MistralFamily_7b.texttext-generationn<1K1 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.