enko
Datasets
All datasets matching “enko”BD-EnKo
BD-EnKo Dataset
It was introduced in the paper "Unveiling the Power of Integration: Block Diagram Summarization through Local-Global Fusion" accepted at ACL 2024. The full code is available in BD-EnKo github repository.
Dataset description
This dataset contains different types of block diagram images with their high-quality summaries.
Types
Train
Validation
English
Korean
English
Korean
-----------------
---------
--------
------------
---------… See the full description on the dataset page: https://huggingface.co/datasets/shreyanshu09/BD-EnKo.en-ko-instenko-math-translate-sftThis is merge of kuotient/orca-math-word-problems-193k-korean and ChuGyouk/AI-MO-NuminaMath-CoT-Ko
EnKo-Translation-LongTextOnly-dedup
장문 번역 데이터만 추출
gemma 토크나이저 기준으로 영문+한글 토큰 합이 1K 이상인 데이터만 추출
데이터 수
1K~2K: 146,957
2K~4K: 11,823
4K~: 2,229
한/영 둘 중 한쪽만 중복인 경우는 제거하지 않았습니다.
데이터 출처
nayohan/aihub-en-ko-translation-12m
nayohan/instruction_en_ko_translation_1.4m
jhflow/orca_ko_en_pair
jhflow/platypus_ko_en_pair
jhflow/dolly_ko_en_pair
heegyu/OIG-small-chip2-ko
lemon-mint/en_ko_translation_purified_v0.1
squarelike/sharegpt_deepl_ko_translation
amphora/parallel-wiki-koen
kuotient/gsm8k-ko… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/EnKo-Translation-LongTextOnly-dedup.PubMedVision-EnKo
Informations
This is the Korean translation of FreedomIntelligence/PubMedVision. The translation was primarily generated using the 'solar-pro-241126' model, with occasional manual assistance from the 'Gemini 2.0 Flash Experimental' model and the 'Gemini experimental 1206' model.
An evaluation of the translation quality ("llm-as-a-judge") will be coming soon.
News
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/PubMedVision-EnKo.OpenOrca-EnKoZhJa-18kThis dataset is a collection of Korean, Chinese, and Japanese OpenOrca translation datasets.
The dataset was matched using id based on kyujinpy/OpenOrca-KO, which had the smallest number of rows.
When more than one translation existed for a language, I chose the more similar one based on similarity of embedding(BAAI/BGE-m3).
Data Sources
English(Original)
Open-Orca/OpenOrca
Korean(Translated with DeepL Pro API)
kyujinpy/OpenOrca-KO
Chinese(Translated with Google Translate)… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/OpenOrca-EnKoZhJa-18k.
