datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legalkit
LegalKit, French labeled datasets built for legal ML training
This dataset consists of labeled data prepared for training sentence embeddings models in the context of French law. The labeling process utilizes the LLaMA-3-70B model through a structured workflow to enhance the quality of the labels. This dataset aims to support the development of natural language processing (NLP) models for understanding and working with legal texts in French.
Labeling Workflow
The… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/legalkit.calme-legalkit-v0.2
Calme LegalKit v0.2
Calme's Enhanced Synthetic Dataset for Advanced Legal Reasoning
🚀 Quick Links
Dataset Page
Original LegalKit Dataset
📖 Overview
Calme LegalKit v0.2 is a synthetically generated dataset designed to enhance legal reasoningand analysis capabilities in language models. This dataset builds upon the foundation laid by Louis Brulé Naudet's LegalKit, incorporating advanced Chain of Thought (CoT) reasoning and specialized… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/calme-legalkit-v0.2.calme-legalkit-v0.1
Calme LegalKit v0.1
Calme's Enhanced Synthetic Dataset for Advanced Legal Reasoning
🚀 Quick Links
Dataset Page
Fine-tuned Model
Original LegalKit Dataset
📖 Overview
Calme LegalKit v0.1 is a synthetically generated dataset designed to enhance legal reasoning and analysis capabilities in language models. This dataset builds upon the foundation laid by Louis Brulé Naudet's LegalKit, incorporating advanced Chain of Thought (CoT) reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/calme-legalkit-v0.1.
