datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pengwin-2026-anatomy-segmentation-femurfinal-finetune-asr-datasetDEExtractChatGPT is used to synthesize paragraphs at two CEFR levels (B1, C2) using a list of verbs (vocab) for two topics (Politics, Economy)
The code for data creation is uploaded on Github.
Cite
@INPROCEEDINGS{10391702,
author={Mustafa, Faizan E},
booktitle={2023 International Conference on Innovation and Intelligence for Informatics, Computing, and Technologies (3ICT)},
title={DEExtract: A Customizable Context-Based German Vocabulary Learning Tool},
year={2023},
volume={}… See the full description on the dataset page: https://huggingface.co/datasets/femustafa/DEExtract.processed_dataset
