CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Open-Orca /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.texttext-classification1M<n<10M1.6k likes21k downloads2y agoHugging Face02Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes19k downloads3y agoHugging Face03Open-Orca /SlimOrca Overview This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions. The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset. This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.texttext-classification100K<n<1M300 likes4.1k downloads3y agoHugging Face04superdrew100 /split_OpenOrca_1M-GPT4-Augmentedtext100K<n<1M0 likes1.7k downloads2y agoHugging Face05malhajar /OpenOrca-tr Dataset Card for "OpenOrca-tr" This Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish dataset collection to enhance the performance of LLM's Produced in the Turkish Language. malhajar/orca-tr is a translated version of the OpenOrca and is the first ever SFT dataset in the Turkish Language with more than 2M entries! Translated by: Mohamad Alhajar Dataset Summary The OpenOrca dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/OpenOrca-tr.texttext-classification1M<n<10M21 likes912 downloads2y agoHugging Face06Open-Orca /SlimOrca-Dedup Overview "SlimOrca Dedup" is a deduplicated, unfiltered subset of the SlimOrca dataset, excluding RLHF instances, resulting in 363k unique examples. Key Features Removal of RLHF instances. Deduplication using minhash and Jaccard similarity techniques. Demo Models Note: These models were trained on the full SlimOrca dataset, not the deduplicated, unfiltered version. * https://huggingface.co/openaccess-ai-collective/jackalope-7b *… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.texttext-classification100K<n<1M94 likes894 downloads1y agoHugging Face07lchakkei /OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源! 這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.texttext-classification1M<n<10M11 likes560 downloads3y agoHugging Face08yys /OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源! 这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.texttext-classification1M<n<10M104 likes445 downloads3y agoHugging Face09Korea-MES /Openorca-Herems2.5text10M<n<100M0 likes331 downloads1y agoHugging Face10lchakkei /OpenOrca-Traditional-Chinese-LLama2-Formattext1M<n<10M0 likes312 downloads3y agoHugging Face11d0rj /OpenOrca-ru OpenOrca-ru This is translated version of Open-Orca/OpenOrca into Russian. texttext-classification1M<n<10M16 likes289 downloads3y agoHugging Face12ctuning /MLPerf-OpenOrcatext10K<n<100K0 likes238 downloads2y agoHugging Face13erhwenkuo /openorca-chinese-zhtw Dataset Card for "openorca-chinese-zhtw" Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope. The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.texttext-classification1M<n<10M3 likes228 downloads3y agoHugging Face14xDAN-Engine /OpenOrca_full_prompttext1M<n<10M0 likes208 downloads2y agoHugging Face15tessiw /german_OpenOrca_Format2 Dataset Card for "german_OpenOrca_Format2" More Information needed text1M<n<10M1 likes205 downloads3y agoHugging Face16SKT27182 /Preprocessed_OpenOrca Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Languages Langugage of the dataset is mostly English. Dataset Structure Data Fields The fields are: 'id', a unique numbered identifier which includes one of 'niv', 't0', 'cot', or 'flan' to represent which source FLAN Collection submix the 'question' is sourced from. 'system_prompt'… See the full description on the dataset page: https://huggingface.co/datasets/SKT27182/Preprocessed_OpenOrca.texttext-classification1M<n<10M1 likes201 downloads3y agoHugging Face17ChuGyouk /OpenOrca_Solar_filtered TODO To be consistent, we need to change column name into ['instruction', 'input', 'output'], which is same as alpaca-gpt4. Dataset Summary This is a filtered version of OpenOrca dataset based on Solar 10.7B paper. In this version, of the 4.2M OpenOrca data, 113k data is removed. In more conservative version here, of the 4.2M OpenOrca data, 117k data is removed. Step 1 FLAN data link broken Based on DataProvenanceInitiative/flan2021_submix_original… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/OpenOrca_Solar_filtered.text1M<n<10M0 likes185 downloads3y agoHugging Face18mlfoundations-dev /open-orca_gpt-4o-mini_scale_x4text1M<n<10M0 likes180 downloads2y agoHugging Face19lchakkei /OpenOrca-Traditional-Chinese-ChatML-Formattext1M<n<10M1 likes161 downloads3y agoHugging Face20Sharathhebbar24 /Cleansed_OpenOrca Orca Cleansed Dataset This is a cleansed version of Open-Orca/OpenOrca Usage Using only Train Split from datasets import load_dataset dataset = load_dataset("Sharathhebbar24/Cleansed_OpenOrca", split="train") It has only train split texttext-generation1M<n<10M0 likes159 downloads3y agoHugging Face21SUSTech /OpenOrca Dataset Card for "OpenOrca" More Information needed text1M<n<10M1 likes158 downloads3y agoHugging Face22Open-Orca /1million-gpt-4text100K<n<1M46 likes138 downloads3y agoHugging Face23dim /OpenOrca-ru-gpt4text100K<n<1M1 likes130 downloads3y agoHugging Face24skymizer /open-orca-conversationstext1M<n<10M5 likes125 downloads2y agoHugging Face25Post-training-Data-Flywheel /OpenOrcatext1M<n<10M0 likes114 downloads2y agoHugging Face26Josephgflowers /OpenOrca-Step-by-step-reasoningThis work was performed to help models with reasoning. I developed it working on my Cinder model, a STEM q and a model. Modified OpenORCA Step-by-Step Reasoning Dataset Overview The Modified OpenORCA Step-by-Step Reasoning Dataset represents a groundbreaking resource in the field of artificial intelligence, specifically designed to enhance the reasoning capabilities of AI models. This unique dataset is the result of a meticulous process of sorting, selecting, and altering dialogues from the… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/OpenOrca-Step-by-step-reasoning.text10K<n<100K18 likes113 downloads3y agoHugging Face27kyujinpy /OpenOrca-KO OpenOrca-KO OpenOrca dataset 중 약 2만개를 sampling하여 번역한 데이터셋 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭 Dataset inf0 NIV // 1571개 FLAN // 9434개 T0 // 6351개 CoT // 2117개 KoCoT // 2159개 Translation Using DeepL Pro API. Thanks. Below is original dataset card 🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/OpenOrca-KO.texttext-classification10K<n<100K31 likes111 downloads3y agoHugging Face28vilm /OpenOrca-Viet 🇻🇳 Vietnamese OpenOrca is here 🐋 Dive into the Vietnamese linguistic landscape with OpenOrca, a cutting-edge dataset crafted through a pioneering partnership between Virtual Interactive and Alignment Lab AI. Drawing inspiration and methodology from the renowned Orca paper, we've expanded our horizons to distill knowledge from a more eclectic mix of leading LLMs including GPT-4, PaLM-2, and Claude. Our vision with this dataset is to fuel research and development that will… See the full description on the dataset page: https://huggingface.co/datasets/vilm/OpenOrca-Viet.text100K<n<1M16 likes107 downloads3y agoHugging Face29zhan1993 /openorca_task_id_cluster Dataset Card for "openorca_task_id_cluster" More Information needed text1M<n<10M0 likes90 downloads2y agoHugging Face30kyujinpy /KOR-OpenOrca-Platypus KOR-OpenOrca-Platypus OpenOrca-Ko + KOpen-platypus 데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭 KOpen-platpyus Repo: KOpen-platypus 고품질 한국어 데이터셋 코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정 1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존 단일 숫자와 영어는 본래의 결과 그대로 가져옴 DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음) DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정 번역하고자 하는 글자수가 1500자 이상일 경우, API로 변경해서 번역 고유명사는 최대한 유지함 Post-processing… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus.texttext-classification10K<n<100K7 likes81 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.