CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Open-Orca /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.texttext-classification1M<n<10M1.6k likes22k downloads2y agoHugging Face02Open-Orca /SlimOrca Overview This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions. The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset. This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.texttext-classification100K<n<1M300 likes4.1k downloads3y agoHugging Face03Open-Orca /SlimOrca-Dedup Overview "SlimOrca Dedup" is a deduplicated, unfiltered subset of the SlimOrca dataset, excluding RLHF instances, resulting in 363k unique examples. Key Features Removal of RLHF instances. Deduplication using minhash and Jaccard similarity techniques. Demo Models Note: These models were trained on the full SlimOrca dataset, not the deduplicated, unfiltered version. * https://huggingface.co/openaccess-ai-collective/jackalope-7b *… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.texttext-classification100K<n<1M94 likes896 downloads1y agoHugging Face04malhajar /OpenOrca-tr Dataset Card for "OpenOrca-tr" This Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish dataset collection to enhance the performance of LLM's Produced in the Turkish Language. malhajar/orca-tr is a translated version of the OpenOrca and is the first ever SFT dataset in the Turkish Language with more than 2M entries! Translated by: Mohamad Alhajar Dataset Summary The OpenOrca dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/OpenOrca-tr.texttext-classification1M<n<10M21 likes858 downloads2y agoHugging Face05lchakkei /OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源! 這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.texttext-classification1M<n<10M11 likes538 downloads3y agoHugging Face06yys /OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源! 这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.texttext-classification1M<n<10M104 likes446 downloads3y agoHugging Face07d0rj /OpenOrca-ru OpenOrca-ru This is translated version of Open-Orca/OpenOrca into Russian. texttext-classification1M<n<10M16 likes284 downloads3y agoHugging Face08erhwenkuo /openorca-chinese-zhtw Dataset Card for "openorca-chinese-zhtw" Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope. The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.texttext-classification1M<n<10M3 likes215 downloads3y agoHugging Face09kyujinpy /OpenOrca-KO OpenOrca-KO OpenOrca dataset 중 약 2만개를 sampling하여 번역한 데이터셋 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭 Dataset inf0 NIV // 1571개 FLAN // 9434개 T0 // 6351개 CoT // 2117개 KoCoT // 2159개 Translation Using DeepL Pro API. Thanks. Below is original dataset card 🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/OpenOrca-KO.texttext-classification10K<n<100K31 likes109 downloads3y agoHugging Face10kyujinpy /KOR-OpenOrca-Platypus KOR-OpenOrca-Platypus OpenOrca-Ko + KOpen-platypus 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭 KOpen-platpyus Repo: KOpen-platypus 고품질 한국어 데이터셋 코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정 1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존 단일 숫자와 영어는 본래의 결과 그대로 가져옴 DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음) DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정 번역하고자 하는 글자수가 1500자 이상일 경우, API로 변경해서 번역 고유명사는 최대한 유지함 Post-processing… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus.texttext-classification10K<n<100K7 likes86 downloads3y agoHugging Face11kyujinpy /KOR-OpenOrca-Platypus-v3 KOR-OpenOrca-Platypus-v3 KOR-OpenOrca-Platypus 데이터셋에서 수작업으로 번역 오류 200건 이상을 고친 데이터셋. 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭 KOpen-platpyus Repo: KOpen-platypus 고품질 한국어 데이터셋 코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정 1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존 단일 숫자와 영어는 본래의 결과 그대로 가져옴 DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음) DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정 번역하고자 하는 글자수가 1500자 이상일 경우, API로… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus-v3.texttext-classification10K<n<100K15 likes76 downloads3y agoHugging Face12wenbopan /OpenOrca-zh-20k Datsetcard for 'OpenOrca-zh-20k' This is the Chinese version of Open-Orca/OpenOrca from Azure99/blossom-orca-v3. Compared to Azure99/blossom-orca-v3: This dataset extracts all Chinese blossom-orca-v3 samples (around 20K) into a separate zh split. All samples are formatted in the ocra format with an optional system role in the first round. Instead of using a 1:1 En-Zh ratio as in blossom-orca-v3, this dataset contains 200K GPT-4 generated English samples from OpenOrca in the en… See the full description on the dataset page: https://huggingface.co/datasets/wenbopan/OpenOrca-zh-20k.textquestion-answering100K<n<1M9 likes62 downloads3y agoHugging Face13polinaeterna /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/OpenOrca.texttext-classification1M<n<10M2 likes58 downloads3y agoHugging Face14Felladrin /ChatML-OpenOrcaOpen-Orca/OpenOrca in ChatML format, ready to use in HuggingFace TRL's SFT Trainer. Python code used for conversion: from datasets import load_dataset from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("Felladrin/Minueza-32M-Base") dataset = load_dataset("Open-Orca/OpenOrca", split="train") def format(columns): messages = [] system_prompt = columns["system_prompt"].strip() if system_prompt: messages.append({ "role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-OpenOrca.texttext-classification1M<n<10M3 likes47 downloads3y agoHugging Face15appleparan /OpenOrca-Ko-En OpenOrca-Ko-En kyujinpy/OpenOrca-KO와 Open-Orca/OpenOrca를 공통된 데이터만 필터링하고 합친 데이터셋입니다. 컬럼은 기존 OpenOrca에 맞춰서 system_prompt_{ko/en}, question_{ko/en}, response_{ko/en} 으로 변경하였습니다. 중복된 id를 제거하여 데이터수가 일부 감소하였습니다. 데이터셋을 만드는데 사용한 스크립트입니다. 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 이 데이터셋뿐만 아니라 위 데이터셋도 함께 출처표기를 해주셨으면 합니다. Dataset inf0 NIV // 1551개(OpenOrca-KO: 1571개) FLAN // 9338개(OpenOrca-KO: 9434개) T0 // 6303개(OpenOrca-KO: 6351개) CoT // 2092개(OpenOrca-KO: 2117개) KoCoT // 2159개… See the full description on the dataset page: https://huggingface.co/datasets/appleparan/OpenOrca-Ko-En.texttext-classification10K<n<100K3 likes46 downloads3y agoHugging Face16squarelike /OpenOrca-gugugo-ko OpenOrca 한국어 번역 데이터셋 Gugugo-koen-7B-V1.1을 이용하여 OpenOrca데이터셋을 번역하고 있습니다. 번역 진행상황은 아래를 참고해 주십시오. 진행상황 GPT4 생성물 약 100만 개 중 약 64만 개 번역완료 GPT3.5 생성물 약 350만 개 중 약 159만 개 번역완료 데이터셋 사용 후 출처표기는 제작자에게 큰 힘이 됩니다. Original dataset card: OpenOrca 🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has… See the full description on the dataset page: https://huggingface.co/datasets/squarelike/OpenOrca-gugugo-ko.texttext-classification1M<n<10M35 likes43 downloads3y agoHugging Face175CD-AI /Vietnamese-Openorca-Multiplechoice-gg-translatedtabularquestion-answering10K<n<100K2 likes43 downloads2y agoHugging Face18kyujinpy /KOR-OpenOrca-Platypus-v2 KOR-OpenOrca-Platypus-v2 KOR-OpenOrca-Platypus 데이터셋에서 수작업으로 번역 오류 200건 이상을 고친 데이터셋. 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭 KOpen-platpyus Repo: KOpen-platypus 고품질 한국어 데이터셋 코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정 1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존 단일 숫자와 영어는 본래의 결과 그대로 가져옴 DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음) DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정 번역하고자 하는 글자수가 1500자 이상일 경우, API로… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus-v2.texttext-classification10K<n<100K5 likes42 downloads3y agoHugging Face19mainbrains /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/mainbrains/OpenOrca.texttext-classification1M<n<10M0 likes37 downloads1mo agoHugging Face20dynopii /OpenOrca-Top5percent🐋 The OpenOrca-Top5Percent Dataset! 🐋 We are excited to introduce the OpenOrca-Top5Percent dataset, a refined version of the original OpenOrca dataset. This dataset contains only those entries which utilize the top 5% most frequently used words in the OpenOrca dataset, aiming to focus on high-frequency vocabulary for various NLP tasks. Dataset Summary The OpenOrca-Top5Percent dataset is a curated subset of the augmented FLAN Collection data, focusing specifically on entries that… See the full description on the dataset page: https://huggingface.co/datasets/dynopii/OpenOrca-Top5percent.texttext-classification1M<n<10M2 likes36 downloads3y agoHugging Face21Preeetam /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Preeetam/OpenOrca.texttext-classification1M<n<10M0 likes36 downloads2d agoHugging Face22GenRM /OpenOrca-Open-Orca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/GenRM/OpenOrca-Open-Orca.text-classification10M<n<100M0 likes35 downloads1y agoHugging Face23KIZZ2006 /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/KIZZ2006/OpenOrca.texttext-classification1M<n<10M0 likes33 downloads9mo agoHugging Face24RuterNorway /OpenOrcaNo-15k🐋 The OpenOrca Dataset Norwegian! 🐋 This is a subset of 15000 rows of the OpenOrca dataset, translated into Norwegian. Translation is done with Amazon Translate, and is provided by Ruter as an artifact from Ruter AI Lab. Dataset structure The dataset is structured in the following way: { "instruction": "Norwegian instruction", "input": "Norwegian input", "output": "Norwegian output", "instruction_en": "English instruction", "input_en": "English input"… See the full description on the dataset page: https://huggingface.co/datasets/RuterNorway/OpenOrcaNo-15k.texttext-classification10K<n<100K5 likes29 downloads3y agoHugging Face25shutajb /openorca-zht Dataset Card for "openorca-chinese-zhtw" Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope. The data is primarily used for training and evaluation in the field of… See the full description on the dataset page: https://huggingface.co/datasets/shutajb/openorca-zht.texttext-classification1M<n<10M0 likes25 downloads3mo agoHugging Face26WiNE-iNEFF /1M-OpenOrca_beEn/Be 🐋 The Belarusian OpenOrca Dataset! 🐋 Belarusian OpenOrca dataset - is rich collection of augmented FLAN data aligns, that translated in belarusian language. That dataset should help training LLM in belarusian language and should help on other NLP tasks. This dataset have 2 version: ~1M GPT-4 completions (Now translating) ~3.2M GPT-3.5 completions (Can be translated in future) Data Fields The fields are: 'id', a unique numbered identifier which includes one of 'niv'… See the full description on the dataset page: https://huggingface.co/datasets/WiNE-iNEFF/1M-OpenOrca_be.texttext-classification100K<n<1M0 likes24 downloads1y agoHugging Face27georgesung /OpenOrca_35k Dataset Card for "OpenOrca_35k" The first 35k examples from Open-Orca/OpenOrca texttext-classification10K<n<100K2 likes19 downloads3y agoHugging Face28villagerad /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/villagerad/OpenOrca.texttext-classification1M<n<10M0 likes15 downloads7mo agoHugging Face29schuler /open-orca-slimorca-deduped-cleaned-corrected-for-pascal-txtThis is a modified version of the slimorca-deduped-cleaned-corrected dataset. It contains English only characters. Open Orca Slim for Pascal Developers is a subset of the original Open Orca dataset . Open Orca Slim for Pascal Developers dataset was created with: from datasets import load_dataset # Coded by Gemini def biggest_char_code(input_string): """ Returns the largest character code in a string. """ if not input_string: return None # Handle empty string case largest_code… See the full description on the dataset page: https://huggingface.co/datasets/schuler/open-orca-slimorca-deduped-cleaned-corrected-for-pascal-txt.texttext-generation100K<n<1M0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.