datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BD-EnKo
BD-EnKo Dataset
It was introduced in the paper "Unveiling the Power of Integration: Block Diagram Summarization through Local-Global Fusion" accepted at ACL 2024. The full code is available in BD-EnKo github repository.
Dataset description
This dataset contains different types of block diagram images with their high-quality summaries.
Types
Train
Validation
English
Korean
English
Korean
-----------------
---------
--------
------------
---------… See the full description on the dataset page: https://huggingface.co/datasets/shreyanshu09/BD-EnKo.en-ko-instenko-math-translate-sftThis is merge of kuotient/orca-math-word-problems-193k-korean and ChuGyouk/AI-MO-NuminaMath-CoT-Ko
EnKo-Translation-LongTextOnly-dedup
장문 번역 데이터만 추출
gemma 토크나이저 기준으로 영문+한글 토큰 합이 1K 이상인 데이터만 추출
데이터 수
1K~2K: 146,957
2K~4K: 11,823
4K~: 2,229
한/영 둘 중 한쪽만 중복인 경우는 제거하지 않았습니다.
데이터 출처
nayohan/aihub-en-ko-translation-12m
nayohan/instruction_en_ko_translation_1.4m
jhflow/orca_ko_en_pair
jhflow/platypus_ko_en_pair
jhflow/dolly_ko_en_pair
heegyu/OIG-small-chip2-ko
lemon-mint/en_ko_translation_purified_v0.1
squarelike/sharegpt_deepl_ko_translation
amphora/parallel-wiki-koen
kuotient/gsm8k-ko… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/EnKo-Translation-LongTextOnly-dedup.PubMedVision-EnKo
Informations
This is the Korean translation of FreedomIntelligence/PubMedVision. The translation was primarily generated using the 'solar-pro-241126' model, with occasional manual assistance from the 'Gemini 2.0 Flash Experimental' model and the 'Gemini experimental 1206' model.
An evaluation of the translation quality ("llm-as-a-judge") will be coming soon.
News
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/PubMedVision-EnKo.OpenOrca-EnKoZhJa-18kThis dataset is a collection of Korean, Chinese, and Japanese OpenOrca translation datasets.
The dataset was matched using id based on kyujinpy/OpenOrca-KO, which had the smallest number of rows.
When more than one translation existed for a language, I chose the more similar one based on similarity of embedding(BAAI/BGE-m3).
Data Sources
English(Original)
Open-Orca/OpenOrca
Korean(Translated with DeepL Pro API)
kyujinpy/OpenOrca-KO
Chinese(Translated with Google Translate)… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/OpenOrca-EnKoZhJa-18k.nlp-arxiv-translation-dponlp-arxiv-translation-dpo-with-math-10kreddit_enko_translation_preference
reddit_enko_translation_preference
Can be used in rlhf (CPO, DPO, etc...)
reject: DeepL
chosen: GPT4-Turbo
Reddit의 다양한 subreddit의 댓글과 글 번역
reject에 Deepl, chosen에 GPT4 번역이지만, GPT의 번역이 반드시 DeepL보다 좋다고 할 순 없습니다. 하고자 하는 방법에 맞춰 사용하시길 바랍니다.
arxiv-translation-result-950chest_radiology_enko
Introduction
By using 대한흉부영상의학회 용어사전, I created the following en-ko sentence pairs.
Prompt
Your job is to create an English-Korean sentence pair with Medical Domain Glossary.
You MUST use the terms in glossary for both sentences.
For example,
[GLOSSARY]
acanthotic keratosis -> 가시세포증식각화증
[/GLOSSARY]
You must create English-Korean sentence pair as below.
[ENG]
Histopathological analysis revealed a marked acanthotic keratosis characterized by epidermal hyperplasia with… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/chest_radiology_enko.trc_uniform_313k_eval_45_filteredWe used nayohan/llama3-8b-it-prometheus-ko to evaluate the nayohan/translate_corpus_uniform_313k dataset with the criteria listed below.
We filtered out rows with a score of 4,5.
def create_conversation(example):
system_prompt = """###Task Description: An instruction (might include an Input inside it), a response to evaluate, a reference answer that gets a score of 5, and a score rubric representing a evaluation criteria are given.
1. Write a detailed feedback that assess the quality of… See the full description on the dataset page: https://huggingface.co/datasets/Translation-EnKo/trc_uniform_313k_eval_45_filtered.subway_disaster_1200_enkoENKO-MEDIQAthis is an edit of the medical data dialog+data which includes both the korean and english version.
this is for personal use only, therefore nothing illegal is done here. thank you
en-ko-OpenSubtitles-samplemath-translation-result-1karxiv-translation-result-6.9k-0909enko_processed
Dataset Card for "enko_processed"
More Information needed
en_ko_translation_social_science_linkbricks_single_dataset_with_prompt_text_huggingface_sampledmath-translation-result-100nlp-arxiv-translation-dpo-filteredEnKo-Intellectual-Property-Terms-GlossarySee: https://www.data.go.kr/data/15066099/fileData.do#
primal-chaos{{ card_data }}
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/EnKop/primal-chaos.en_ko_translation_tech_science_linkbricks_single_dataset_with_prompt_text_huggingface_sampleden_ko_translation_social_science_linkbricks_single_dataset_with_prompt_text_huggingfacearxiv-translationEnKo-Translation-Preference-Eval
Mistranslation dataset for evaluating the performance of reward models or filtering methods used in assessing the quality of Korean to English translations.
Unnatural grammar usage, misinterpretation due to incorrect phrase/clause segmentation, awkward terminology, etc.
Yes, 57 records is too small to evaluate something.
en_ko_translation_tech_science_linkbricks_single_dataset_with_prompt_text_huggingfaceen_ko_translation_purified_v0.1parallel_enko_feedback_collection_fullParallel
ENG: prometheus-eval/Feedback-Collection
KOR: nayohan/feedback-collection-ko-full
Filter out 1240 sentences with repeated sentences among the translated datasets
99,952 -> 98,712
