CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shreyanshu09 /BD-EnKo BD-EnKo Dataset It was introduced in the paper "Unveiling the Power of Integration: Block Diagram Summarization through Local-Global Fusion" accepted at ACL 2024. The full code is available in BD-EnKo github repository. Dataset description This dataset contains different types of block diagram images with their high-quality summaries. Types Train Validation English Korean English Korean ----------------- --------- -------- ------------ ---------… See the full description on the dataset page: https://huggingface.co/datasets/shreyanshu09/BD-EnKo.image10K<n<100K2 likes228 downloads2y agoHugging Face02DopeorNope /en-ko-insttext1M<n<10M0 likes116 downloads3y agoHugging Face03amphora /enko-math-translate-sftThis is merge of kuotient/orca-math-word-problems-193k-korean and ChuGyouk/AI-MO-NuminaMath-CoT-Ko text1M<n<10M0 likes68 downloads2y agoHugging Face04werty1248 /EnKo-Translation-LongTextOnly-dedup 장문 번역 데이터만 추출 gemma 토크나이저 기준으로 영문+한글 토큰 합이 1K 이상인 데이터만 추출 데이터 수 1K~2K: 146,957 2K~4K: 11,823 4K~: 2,229 한/영 둘 중 한쪽만 중복인 경우는 제거하지 않았습니다. 데이터 출처 nayohan/aihub-en-ko-translation-12m nayohan/instruction_en_ko_translation_1.4m jhflow/orca_ko_en_pair jhflow/platypus_ko_en_pair jhflow/dolly_ko_en_pair heegyu/OIG-small-chip2-ko lemon-mint/en_ko_translation_purified_v0.1 squarelike/sharegpt_deepl_ko_translation amphora/parallel-wiki-koen kuotient/gsm8k-ko… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/EnKo-Translation-LongTextOnly-dedup.texttranslation100K<n<1M6 likes47 downloads2y agoHugging Face05werty1248 /OpenOrca-EnKoZhJa-18kThis dataset is a collection of Korean, Chinese, and Japanese OpenOrca translation datasets. The dataset was matched using id based on kyujinpy/OpenOrca-KO, which had the smallest number of rows. When more than one translation existed for a language, I chose the more similar one based on similarity of embedding(BAAI/BGE-m3). Data Sources English(Original) Open-Orca/OpenOrca Korean(Translated with DeepL Pro API) kyujinpy/OpenOrca-KO Chinese(Translated with Google Translate)… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/OpenOrca-EnKoZhJa-18k.text10K<n<100K0 likes34 downloads2y agoHugging Face06Translation-EnKo /nlp-arxiv-translation-dpotext1K<n<10K1 likes31 downloads2y agoHugging Face07Translation-EnKo /nlp-arxiv-translation-dpo-with-math-10ktext10K<n<100K0 likes28 downloads2y agoHugging Face08kuotient /reddit_enko_translation_preference reddit_enko_translation_preference Can be used in rlhf (CPO, DPO, etc...) reject: DeepL chosen: GPT4-Turbo Reddit의 다양한 subreddit의 댓글과 글 번역 reject에 Deepl, chosen에 GPT4 번역이지만, GPT의 번역이 반드시 DeepL보다 좋다고 할 순 없습니다. 하고자 하는 방법에 맞춰 사용하시길 바랍니다. text1K<n<10K5 likes25 downloads3y agoHugging Face09Translation-EnKo /arxiv-translation-result-950textn<1K1 likes24 downloads2y agoHugging Face10ChuGyouk /chest_radiology_enko Introduction By using 대한흉부영상의학회 용어사전, I created the following en-ko sentence pairs. Prompt Your job is to create an English-Korean sentence pair with Medical Domain Glossary. You MUST use the terms in glossary for both sentences. For example, [GLOSSARY] acanthotic keratosis -> 가시세포증식각화증 [/GLOSSARY] You must create English-Korean sentence pair as below. [ENG] Histopathological analysis revealed a marked acanthotic keratosis characterized by epidermal hyperplasia with… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/chest_radiology_enko.textn<1K2 likes24 downloads2y agoHugging Face11Translation-EnKo /trc_uniform_313k_eval_45_filteredWe used nayohan/llama3-8b-it-prometheus-ko to evaluate the nayohan/translate_corpus_uniform_313k dataset with the criteria listed below. We filtered out rows with a score of 4,5. def create_conversation(example): system_prompt = """###Task Description: An instruction (might include an Input inside it), a response to evaluate, a reference answer that gets a score of 5, and a score rubric representing a evaluation criteria are given. 1. Write a detailed feedback that assess the quality of… See the full description on the dataset page: https://huggingface.co/datasets/Translation-EnKo/trc_uniform_313k_eval_45_filtered.tabular100K<n<1M0 likes23 downloads2y agoHugging Face12TeamWD /subway_disaster_1200_enkotext1K<n<10K0 likes22 downloads2y agoHugging Face13BarahFazili /en-ko-OpenSubtitles-sampletext100K<n<1M0 likes20 downloads2y agoHugging Face14Translation-EnKo /math-translation-result-1ktext1K<n<10K0 likes17 downloads2y agoHugging Face15Translation-EnKo /arxiv-translation-result-6.9k-0909text1K<n<10K0 likes17 downloads2y agoHugging Face16xchange /enko_processed Dataset Card for "enko_processed" More Information needed tabular100K<n<1M1 likes15 downloads3y agoHugging Face17Translation-EnKo /math-translation-result-100textn<1K0 likes12 downloads2y agoHugging Face18Translation-EnKo /nlp-arxiv-translation-dpo-filteredtext1K<n<10K0 likes12 downloads2y agoHugging Face19ChuGyouk /EnKo-Intellectual-Property-Terms-GlossarySee: https://www.data.go.kr/data/15066099/fileData.do# text1K<n<10K1 likes12 downloads2y agoHugging Face20EnKop /primal-chaos{{ card_data }} Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/EnKop/primal-chaos.textn<1K0 likes12 downloads2y agoHugging Face21Translation-EnKo /arxiv-translationtext1K<n<10K0 likes10 downloads2y agoHugging Face22werty1248 /EnKo-Translation-Preference-Eval Mistranslation dataset for evaluating the performance of reward models or filtering methods used in assessing the quality of Korean to English translations. Unnatural grammar usage, misinterpretation due to incorrect phrase/clause segmentation, awkward terminology, etc. Yes, 57 records is too small to evaluate something. textn<1K0 likes10 downloads2y agoHugging Face23nayohan /parallel_enko_feedback_collection_fullParallel ENG: prometheus-eval/Feedback-Collection KOR: nayohan/feedback-collection-ko-full Filter out 1240 sentences with repeated sentences among the translated datasets 99,952 -> 98,712 tabular10K<n<100K2 likes9 downloads2y agoHugging Face24Translation-EnKo /math-translation-dpogatedtext10K<n<100K0 likes9 downloads2y agoHugging Face25nayohan /parallel_enko_feedback_collection_full_chattext10K<n<100K2 likes7 downloads2y agoHugging Face26EunjiChoi /enkoDFtext10K<n<100K0 likes6 downloads1y agoHugging Face27sskskfskfj /en_ko_datasettext1K<n<10K0 likes4 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.