datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mistral_Trivia-QA_Dataset
Mistral Trivia QA Dataset
The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation.
Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.mistral_medical_meadow_wikidoc_instruct_datasetDataset made for instruction supervised finetuning of Mistral LLMs based on the Medical meadow wikidoc dataset:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively contribute and share contemporary medical knowledge. WikiDoc features two primary… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/mistral_medical_meadow_wikidoc_instruct_dataset.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.mistral_medquad_instruct_datasetDataset made for instruction supervised finetuning of Mistral LLMs based on the Medquad dataset:
Medquad dataset (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medquad
MedQuAD is a comprehensive collection consisting of 47,457 medical question-answer pairs compiled from 12 authoritative sources within the National Institutes of Health (NIH), including domains like cancer.gov, niddk.nih.gov, GARD, and MedlinePlus Health Topics. These question-answer pairs span 37 distinct… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/mistral_medquad_instruct_dataset.Mistral-7b-0.3-Instruct-DBpedia-HighlyKnownDataset for paper “How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?”
Based on DBpedia dataset
Paper, Code
Mistral-7b-0.3-Instruct-TriviaQA-HighlyKnownDataset for paper “How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?”
Based on TriviaQA dataset
(https://huggingface.co/papers/2502.14502)
Medical-QA-Mistral7B-Finetuningmistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.veneto-mistral-dataset
Veneto Mistral Dataset
A conversational dataset for training AI models in Venetian language (vèneto).
Description
This dataset was created to preserve and digitalize the Venetian language through artificial intelligence. It contains approximately 11,800 examples of conversations, texts, and translations in Venetian language, extracted from authentic sources and validated for linguistic quality.
The dataset was specifically designed for fine-tuning large language models… See the full description on the dataset page: https://huggingface.co/datasets/Marco-Danz/veneto-mistral-dataset.medical_mistral_instruct_datasetDataset made for instruction supervised finetuning of Mistral LLMs, by combining of medical datasets:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medquad (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively contribute and share contemporary… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/medical_mistral_instruct_dataset.medical_mistral_instruct_dataset_shortDataset made for instruction supervised finetuning of Mistral LLMs, by combining of medical datasets and getting 2k entries from them:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medquad (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/medical_mistral_instruct_dataset_short.ProcessedOpenAssistant-mistral-large-2411
Dataset Summary
Processed OpenAssistant — Mistral‑Large‑2411 pairs 27 563 unique English prompts—deduplicated from the Apache‑2.0–licensed [Processed OpenAssistant corpus]—with answers generated on 21 Apr 2025 by the paid‑API model mistral‑large‑2411, the 24‑Nov‑2024 checkpoint of Mistral’s 123 B parameter instruction‑tuned series :contentReference[oaicite:0]{index=0} :contentReference[oaicite:1]{index=1}.Answers were produced with the /v1/chat/completions endpoint in streaming… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/ProcessedOpenAssistant-mistral-large-2411.mistral-awesome-chatgpt-prompts
Dataset Card for mistral-awesome-chatgpt-prompts
Dataset Summary
mistral-awesome-chatgpt-prompts is a compact, English-language instruction dataset that pairs the 203 classic prompts from the public-domain Awesome ChatGPT Prompts list with answers produced by three proprietary Mistral-AI chat models (mistral-small-2503, mistral-medium-2505, mistral-large-2411).
The result is a five-column table—act, prompt, and one answer column per model—intended for… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/mistral-awesome-chatgpt-prompts.
