datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hc3_french_oodHuman ChatGPT Comparison Corpus (HC3) Translated To French.
The translation is done by Google Translate API.
We also add the native french QA pairs from ChatGPT, BingGPT and FAQ pages.
This dataset was used in our TALN 2023 paper.
Towards a Robust Detection of Language Model-Generated Text: Is ChatGPT that easy to detect?FormationEval
FormationEval
FormationEval is a public benchmark suite for petroleum geoscience language model evaluation.
default remains the evaluated MCQ v0.1 track (Christmas 2025) with 505 questions and 72 published model results.
diskos_qa adds 1027 QA items imported 17 March 2026 from DISKOS-QA as a separate track.
spe_mcq adds 100 MCQ items imported 21 March 2026 from ynuwara/spe_mcq_dataset as a separate track.
The public leaderboard, charts and quiz still reflect the evaluated MCQ v0.1… See the full description on the dataset page: https://huggingface.co/datasets/AlmazErmilov/FormationEval.Stage2-3k_Arabic_ChatML
