datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jfv-french-style-conditioning-dataset-v1.0
JFV French Paired Style-Conditioning Dataset
At a Glance
Item
Value
Language
French
Source
Single-author blog corpus, 2005–2025
Public release
v1.0
Aligned units in public release
1,484
Texts in public aligned release
7,420
Original experiment
1,492 aligned units / 7,460 texts
Generated conditions
Ministral baseline; profile; profile + five-shot examples
Primary use
Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.
