datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-Openorca-Multiplechoice-gg-translatedapollo_english_guidelines_translated_to_dutch_with_gpt4omini
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
apollo_english_guidelines_translated_to_dutch_with_nllb200
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
apollo_english_books_translated_to_dutch_with_geminiflash15
Data description
Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationepfl_english_guidelines_translated_to_dutch_with_gpt4omini
Dataset Card for Epfl English Guidelines Translated To Dutch With Gpt4Omini
This dataset was created by the EPFL, and can found in it original form here
The source language: English
The original data source: Original Data Source
Data description
Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini
Acknowledgement
This is part of the DT4H project with attribution [Cite the paper].
Doi and… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_english_guidelines_translated_to_dutch_with_gpt4omini.apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
apollo_english_guidelines_translated_to_dutch_with_marianmt
Data description
Apollo corpus, English guidelines translated to Dutch using MariaNMT.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
short-summaries-translatedccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_formattranslated_v4_persuasion_datasetMoralExceptQA-translatedarabic_dataset_translated_v2_ViT-B-16-SigLIP-512chinese-translated
