CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Viet-Mistral /CulturaY CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.texttext-generation1B<n<10B39 likes6.5k downloads2y agoHugging Face02mistral-hackaton-2026 /zebra-cot-mistral-small-3.2-24b-preprocessed Zebra-CoT Preprocessed — Mistral Hackathon 2026 Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct. Format text: formatted as [INST] question [/INST] <think> reasoning </think> answer image: PIL JPEG image for the corresponding visual task Usage Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning. Hackathon Created for Mistral Hackaton 2026 — Fine-tuning track with W&B. imagevisual-question-answering100K<n<1M0 likes379 downloads7mo agoHugging Face03snorkelai /Snorkel-Mistral-PairRM-DPO-Dataset Dataset: This is the data used for training Snorkel model We use ONLY the prompts from UltraFeedback; no external LLM responses used. Methodology: Generate 5 response variations for each prompt from a subset of 20,000 using the LLM - to start, we used Mistral-7B-Instruct-v0.2. Apply PairRM for response reranking. Update the LLM by applying Direct Preference Optimization (DPO) on the top (chosen) and bottom (rejected) responses. Use this LLM as the base model for the next… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Snorkel-Mistral-PairRM-DPO-Dataset.texttext-generation10K<n<100K45 likes221 downloads3y agoHugging Face04rzagirov /mistral_playwright_mcp Playwright MCP Fine-Tuning Dataset for Mistral 7B Instruct 0.3 A testing-focused dataset for fine-tuning Mistral 7B Instruct 0.3 on Playwright MCP (Model Context Protocol) browser automation tasks. This dataset has been filtered to include only crisp, actionable browser testing scenarios. Dataset Description This dataset contains 1,235 examples of browser automation testing tasks formatted for instruction-following models. It combines data from multiple sources including… See the full description on the dataset page: https://huggingface.co/datasets/rzagirov/mistral_playwright_mcp.text-generation1K<n<10K3 likes119 downloads9mo agoHugging Face05danlou /based-chat-v0.1-Mistral-Nemo-Base-2407 Based-Chat v0.1 (Mistral Nemo Base 2407) This dataset was developed as part of an exploration into understanding the necessity of supervised datasets for fine-tuning base LLMs into conversational models. It's a synthetic dataset created with Mistral-Nemo-Base-2407, and used to fine-tune that model, producing relay-v0.1-Mistral-Nemo-2407. Methodology This synthetic dataset is generated using the following as conversation starters: facebook/empathetic_dialogues… See the full description on the dataset page: https://huggingface.co/datasets/danlou/based-chat-v0.1-Mistral-Nemo-Base-2407.texttext-generation100K<n<1M1 likes62 downloads2y agoHugging Face06ruggsea /stanford-encyclopedia-of-philosophy_chat_multi_turn_mistral_largeThis dataset is essentially identical to the Stanford Encyclopedia of Philosophy Chat Multi-turn Dataset, with one key difference: it uses Mistral Large 2 for conversation generation instead of LLaMA 3.1 70B. All other aspects, including format, statistics, and intended use, remain the same as the original dataset. texttext-generation10K<n<100K5 likes61 downloads2y agoHugging Face07mistral-hackaton-2026 /robuchan-data Robuchan Dataset Synthetic dietary recipe adaptation dataset for fine-tuning language models. Each example is a chat-format conversation where a user provides a recipe and dietary restriction, and the assistant produces a structured adaptation. Generated for the Mistral AI Worldwide Hackathon Tokyo (Feb 28 - Mar 1, 2026). Associated model: sumitdotml/robuchan Dataset Structure Splits Split Rows Purpose train 1,090 Fine-tuning training set… See the full description on the dataset page: https://huggingface.co/datasets/mistral-hackaton-2026/robuchan-data.texttext-generation1K<n<10K0 likes41 downloads7mo agoHugging Face08siqi00 /mistral_ultrafeedback_unhelpful_chatprompt_0.7_1.0_50_320This dataset is used in the paper Discriminative Finetuning of Generative Large Language Models without Reward Models and Preference Data. It contains paired examples of "real" conversation turns and multiple generated alternatives. The goal is to train a model to discriminate between high-quality and low-quality generations. The dataset is structured as follows: Each example contains a "real" conversation turn and 16 generated alternatives. Each turn is represented as a list of… See the full description on the dataset page: https://huggingface.co/datasets/siqi00/mistral_ultrafeedback_unhelpful_chatprompt_0.7_1.0_50_320.texttext-generation10K<n<100K0 likes40 downloads2y agoHugging Face09davidpistori /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.texttext-generation10K<n<100K0 likes37 downloads3mo agoHugging Face10VinceGx33 /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.texttext-generation10K<n<100K2 likes30 downloads11mo agoHugging Face11bala1524 /Medical-QA-Mistral7B-Finetuningtextquestion-answeringn<1K6 likes29 downloads3y agoHugging Face12CoolSpring /morse-code-mistral-dpoModified from snorkelai/Snorkel-Mistral-PairRM-DPO-Dataset. Only train_iteration_1 and test_iteration_1 were used. Rows with unsupported characters were filtered out. was created in an attempt to teach the model to talk in Morse code prompt: original prompt field chosen: last message in original chosen field, encoded in Morse code with CyberChef rejected: last message in original chosen field texttext-generation10K<n<100K0 likes27 downloads2y agoHugging Face13hafufu-stack /mistral-hallucination-vaccine Mistral Hallucination Vaccine: Canary-Labeled Self-Healing Dataset 🧬 The first fully automated, canary-labeled hallucination dataset for LLMs — with built-in self-healing. Dataset Description This dataset contains 1,002 samples of LLM responses generated under three conditions: Clean (334 samples): Normal Mistral-7B responses to factual questions Nightmare (334 samples): Responses generated with hidden-state noise injection (σ=0.10), inducing hallucinations Healed… See the full description on the dataset page: https://huggingface.co/datasets/hafufu-stack/mistral-hallucination-vaccine.text-classification1K<n<10K0 likes20 downloads7mo agoHugging Face14ZachW /mistral-small-3.2-24b-instruct-2506_aime-all mistralai/Mistral-Small-3.2-24B-Instruct-2506 — aime-all Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: aime-all (933 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_aime-all.tabulartext-generationn<1K0 likes20 downloads5mo agoHugging Face15DataPilot /Zero_SFT_Ja_by_Mistral_Small DataPilot/Zero_SFT_Ja_by_Mistral_Small このデータセットは、日本語で記述された高品質な合成プロンプトとそのAI出力を収録しています。すべてのデータは Mistral Small 3.1 24B Instruct 2503 モデルを使用してゼロから合成されています。 概要 項目 詳細 データセット名 DataPilot/Zero_SFT_Ja_by_Mistral_Small 言語 日本語 データ作成方法 完全自動生成(モデルによるゼロショット合成) 使用モデル Mistral Small 3.1 24B Instruct 2503 フォーマット JSONL(id, input, output, conversation) ライセンス Apache-2.0 作成コード foxn2000/zero_one_instruction データセット構造 データセットには以下のカラムが含まれています。 カラム名… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Zero_SFT_Ja_by_Mistral_Small.texttext-generation1K<n<10K1 likes19 downloads1y agoHugging Face16ZachW /mistral-small-3.2-24b-instruct-2506_writingbench-en100 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — writingbench-en100 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: writingbench-en100 (100 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 8192 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_writingbench-en100.tabulartext-generationn<1K0 likes19 downloads5mo agoHugging Face17hari560 /Mistral_AI_Medical_Datasettexttext-generation10K<n<100K2 likes18 downloads2y agoHugging Face18alvarobartt /instruction-dataset-mistral-7b-instruct-v0.2 HuggingFaceH4/instruction-dataset but generated with Mistral-7B-Instruct-v0.2 This dataset has been generated with distilabel v1.0.0.b0 using vLLM and running mistralai/Mistral-7B-Instruct-v0.2, using the prompt column from the dataset HuggingFaceH4/instruction-dataset. texttext-generationn<1K0 likes17 downloads3y agoHugging Face19Marco-Danz /veneto-mistral-dataset Veneto Mistral Dataset A conversational dataset for training AI models in Venetian language (vèneto). Description This dataset was created to preserve and digitalize the Venetian language through artificial intelligence. It contains approximately 11,800 examples of conversations, texts, and translations in Venetian language, extracted from authentic sources and validated for linguistic quality. The dataset was specifically designed for fine-tuning large language models… See the full description on the dataset page: https://huggingface.co/datasets/Marco-Danz/veneto-mistral-dataset.texttext-generation10K<n<100K1 likes17 downloads8mo agoHugging Face20ZachW /mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — alpaca-text-generation-384 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: alpaca-text-generation-384 (384 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384.tabulartext-generationn<1K0 likes17 downloads5mo agoHugging Face21ZachW /mistral-small-3.2-24b-instruct-2506_ifeval mistralai/Mistral-Small-3.2-24B-Instruct-2506 — ifeval Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: ifeval (541 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_ifeval.tabulartext-generationn<1K0 likes16 downloads5mo agoHugging Face22ZachW /mistral-small-3.2-24b-instruct-2506_storygen-prompts-200 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — storygen-prompts-200 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: storygen-prompts-200 (200 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_storygen-prompts-200.tabulartext-generationn<1K0 likes16 downloads5mo agoHugging Face23HachiML /JMT-Bench-result_self-rewarding_Mistral-7B-lora JMT-Bench result Answer language JMT-Benchの回答のうち、Englishで回答した件数 Model Count mistralai/Mistral-7B-v0.3 25 HachiML/Mistral-7B-v0.3-m1-lora 7 HachiML/Mistral-7B-v0.3-m2-lora 7 HachiML/Mistral-7B-v0.3-m3-lora 2 tabulartext-generationn<1K0 likes14 downloads2y agoHugging Face24mistral0105 /hygon_cuda_agent Hygon K100 CUDA Agent Benchmark (Strict v2) This dataset is re-split with a strict kernel-quality policy. Strict v2 verified rule A sample remains in verified only if all conditions hold: verify_success == true uses_cuda_extension == true uses_cudnn_path == false speedup_vs_baseline <= 3 speedup_vs_compile <= 3 outlier_risk != high All other verified samples are moved to outlier. Splits Split Count verified 42 compile_failed 82… See the full description on the dataset page: https://huggingface.co/datasets/mistral0105/hygon_cuda_agent.tabulartext-generationn<1K0 likes14 downloads7mo agoHugging Face25DhruvParth /Mistral-7B-Instruct-v2.0-PairRM-DPO-Datasettexttext-generationn<1K0 likes11 downloads2y agoHugging Face26ZachW /mistral-small-3.2-24b-instruct-2506_bookmia-label0-5pct-raw mistralai/Mistral-Small-3.2-24B-Instruct-2506 — bookmia-label0-5pct-raw Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: bookmia-label0-5pct-raw (247 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_bookmia-label0-5pct-raw.tabulartext-generationn<1K0 likes11 downloads5mo agoHugging Face27ZachW /mistral-small-3.2-24b-instruct-2506_creativemath-with-answers mistralai/Mistral-Small-3.2-24B-Instruct-2506 — creativemath-with-answers Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: creativemath-with-answers (188 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_creativemath-with-answers.tabulartext-generationn<1K0 likes11 downloads5mo agoHugging Face28dpevzner /Cybersecurity_Reasoning_Dataset_MistralFamily_7bgated Cybersecurity Reasoning Dataset (v6.0) A high-fidelity, forensics-mapped training corpus for cybersecurity reasoning and SOC automation — formatted for the Mistral / Llama instruct family. Format-specific dataset. Every record uses the Alpaca-style ### Instruction: / ### Response: template native to Mistral/Llama instruct models. A model-agnostic version of this corpus (Mistral, DeepSeek, ChatML, and Gemma variants rendered from one neutral canonical source) is published… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset_MistralFamily_7b.texttext-generationn<1K1 likes10 downloads2mo agoHugging Face29ZachW /mistral-small-3.2-24b-instruct-2506_arena-hard-creative-writing mistralai/Mistral-Small-3.2-24B-Instruct-2506 — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_arena-hard-creative-writing.tabulartext-generationn<1K0 likes8 downloads5mo agoHugging Face30Yobitel /mistralai-mistral-7b-instruct-v0-3__llm-inference-chatbot-short__019e3b458552 mistralai/Mistral-7B-Instruct-v0.3 on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit TTFT P50 20.7686 ms TTFT P99 87.917 ms TPOT P50 7.0405 ms TPOT P99 8.6926 ms Total P50 Ms 817.1104 Total P99 Ms 1095.811 Req Per S Passing 4.2882 Req Per S All 4.3358 Compliance Rate 0.989 Ok Rate 1 Throughput Tok Per S 472.1276 Power Avg W 900.9043 Power Peak W 937.418 Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/mistralai-mistral-7b-instruct-v0-3__llm-inference-chatbot-short__019e3b458552.text-generationn<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.