datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wolof-to-French_Translation-Dataset
Dataset Wolof ↔ Français
🧩 Présentation
Ce dataset contient plus de 30 000 paires phrase Wolof – phrase Française.Chaque ligne est structurée comme suit :
Wolof (input)
Français (target)
Phrase en Wolof
Phrase correspondante en Français
Il a été conçu pour la traduction automatique et les tâches de NLP impliquant le Wolof et le Français.
📚 Provenance et nettoyage
Le dataset a été créé en compilant différentes sources accessibles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/Wolof-to-French_Translation-Dataset.French-Speech-Dataset
🎧 French Speech Dataset
The French Speech Dataset is a comprehensive speech audio dataset designed to deliver high-quality and diverse audio data for advanced AI and machine learning applications. It includes 198 hours of audio data across 912 files, provided in MP3 and WAV formats, with a total size of 445 MB. This well-structured audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years.… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/French-Speech-Dataset.tibetan-to-french-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the French translation of the Tibetan.
The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced.
The dataset is part of the larger MLotsawa project, the code repo for which can be found here.
French-ShiKomori-Parallel-Dataset
French-shiKomori Parallel Dataset v1.0
Dataset Description
The French-shiKomori Parallel Dataset v1.0 is a French–Shingazidja parallel corpus developed by Komor-IA as part of its efforts to build open language resources and artificial intelligence technologies for the Comorian languages.
This first version contains 1,424 aligned French–Shingazidja sentence pairs, with French sentences paired with their corresponding translations in Shingazidja.
The dataset is the… See the full description on the dataset page: https://huggingface.co/datasets/Komor-IA/French-ShiKomori-Parallel-Dataset.french-speech-recognition-dataset
French Speech Dataset for recognition task
Dataset comprises 547 hours of telephone dialogues in French, collected from 964 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/french-speech-recognition-dataset.french-speech-recognition-dataset
French Telephone Dialogues Dataset - 547 Hours
his speech recognition dataset comprises 547 hours of telephone dialogues in French from 964 native speakers, providing audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for training and evaluating automatic speech recognition technology. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/french-speech-recognition-dataset.French-Wolof_Translation-Dataset
Dataset Français → Wolof (train.csv)
Colonnes :
input : texte en français
target : traduction en wolof
Description :Ce dataset contient des paires français → wolof soigneusement collectées et nettoyées à partir de diverses sources en ligne.Toutes les données ont été vérifiées, les doublons et incohérences supprimés, garantissant un dataset fiable et prêt à l’usage pour l’entraînement de modèles de traduction automatique.
Usage prévu :
Fine-tuning et entraînement de modèles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/French-Wolof_Translation-Dataset.French_Legal_Translation_Compliance_Dataset
High Quality French Legal & Compliance Dataset
This dataset contains high-quality labeled French business and compliance communications designed for AI training and NLP applications.
Overview
Language: FrenchDomain: Legal, Compliance, Business CommunicationEntries: ~10,000 labeled samplesFormat: CSV
Use Cases
Training Large Language Models (LLM)
Compliance automation systems
Business communication AI
NLP research
Translation models
Features… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/French_Legal_Translation_Compliance_Dataset.parle-french-pronunciation-datasets
Parle French Pronunciation Datasets
Version v2026.08.11 · DOI 10.5281/zenodo.21932586
This first-party open-data release documents two bounded learning structures used by Parle: French Pronunciation, an iPhone and iPad App from Tingnova Inc. for English-speaking complete beginners and A0–A2 learners.
Ting Dong, the creator of Parle, maintains this dataset for Tingnova Inc. This repository is product documentation and open educational data. It is not an independent review… See the full description on the dataset page: https://huggingface.co/datasets/Levindong/parle-french-pronunciation-datasets.Goliath-dataset-french-toxicityEnglish_French_safety_code-mixing_datasetUsing HarmBench promtps as the English baselines
