datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.FLAN🍮 The WHOLE FLAN Collection! 🍮
Overview
This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets.
Generated using the official seqio templating from the Google FLAN Collection GitHub repo.
The data is subject to all the same licensing of the component datasets.
To keep up with our continued work on OpenOrca and other exciting research, find our Discord here:
https://AlignmentLab.ai
Motivation
This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.SlimOrca
Overview
This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions.
The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset.
This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.split_OpenOrca_1M-GPT4-AugmentedOpenOrca-tr
Dataset Card for "OpenOrca-tr"
This Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish dataset collection to enhance the performance of LLM's Produced in the Turkish Language.
malhajar/orca-tr is a translated version of the OpenOrca and is the first ever SFT dataset in the Turkish Language with more than 2M entries!
Translated by: Mohamad Alhajar
Dataset Summary
The OpenOrca dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/OpenOrca-tr.SlimOrca-Dedup
Overview
"SlimOrca Dedup" is a deduplicated, unfiltered subset of the SlimOrca dataset, excluding RLHF instances, resulting in 363k unique examples.
Key Features
Removal of RLHF instances.
Deduplication using minhash and Jaccard similarity techniques.
Demo Models
Note: These models were trained on the full SlimOrca dataset, not the deduplicated, unfiltered version.
* https://huggingface.co/openaccess-ai-collective/jackalope-7b
*… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源!
這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源!
这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.Openorca-Herems2.5OpenOrca-Traditional-Chinese-LLama2-FormatOpenOrca-ru
OpenOrca-ru
This is translated version of Open-Orca/OpenOrca into Russian.
MLPerf-OpenOrcaopenorca-chinese-zhtw
Dataset Card for "openorca-chinese-zhtw"
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope.
The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.OpenOrca_full_promptgerman_OpenOrca_Format2
Dataset Card for "german_OpenOrca_Format2"
More Information needed
Preprocessed_OpenOrca
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Languages
Langugage of the dataset is mostly English.
Dataset Structure
Data Fields
The fields are:
'id', a unique numbered identifier which includes one of 'niv', 't0', 'cot', or 'flan' to represent which source FLAN Collection submix the 'question' is sourced from.
'system_prompt'… See the full description on the dataset page: https://huggingface.co/datasets/SKT27182/Preprocessed_OpenOrca.OpenOrca_Solar_filtered
TODO
To be consistent, we need to change column name into ['instruction', 'input', 'output'], which is same as alpaca-gpt4.
Dataset Summary
This is a filtered version of OpenOrca dataset based on Solar 10.7B paper.
In this version, of the 4.2M OpenOrca data, 113k data is removed.
In more conservative version here, of the 4.2M OpenOrca data, 117k data is removed.
Step 1
FLAN data link broken
Based on DataProvenanceInitiative/flan2021_submix_original… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/OpenOrca_Solar_filtered.open-orca_gpt-4o-mini_scale_x4OpenOrca-Traditional-Chinese-ChatML-FormatCleansed_OpenOrca
Orca Cleansed Dataset
This is a cleansed version of Open-Orca/OpenOrca
Usage
Using only Train Split
from datasets import load_dataset
dataset = load_dataset("Sharathhebbar24/Cleansed_OpenOrca", split="train")
It has only train split
OpenOrca
Dataset Card for "OpenOrca"
More Information needed
1million-gpt-4OpenOrca-ru-gpt4open-orca-conversationsOpenOrcaOpenOrca-Step-by-step-reasoningThis work was performed to help models with reasoning. I developed it working on my Cinder model, a STEM q and a model.
Modified OpenORCA Step-by-Step Reasoning Dataset Overview
The Modified OpenORCA Step-by-Step Reasoning Dataset represents a groundbreaking resource in the field of artificial intelligence, specifically designed to enhance the reasoning capabilities of AI models. This unique dataset is the result of a meticulous process of sorting, selecting, and altering dialogues from the… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/OpenOrca-Step-by-step-reasoning.OpenOrca-KO
OpenOrca-KO
OpenOrca dataset 중 약 2만개를 sampling하여 번역한 데이터셋
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭
Dataset inf0
NIV // 1571개
FLAN // 9434개
T0 // 6351개
CoT // 2117개
KoCoT // 2159개
Translation
Using DeepL Pro API. Thanks.
Below is original dataset card
🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/OpenOrca-KO.OpenOrca-Viet
🇻🇳 Vietnamese OpenOrca is here 🐋
Dive into the Vietnamese linguistic landscape with OpenOrca, a cutting-edge dataset crafted through a pioneering partnership between Virtual Interactive and Alignment Lab AI. Drawing inspiration and methodology from the renowned Orca paper, we've expanded our horizons to distill knowledge from a more eclectic mix of leading LLMs including GPT-4, PaLM-2, and Claude. Our vision with this dataset is to fuel research and development that will… See the full description on the dataset page: https://huggingface.co/datasets/vilm/OpenOrca-Viet.openorca_task_id_cluster
Dataset Card for "openorca_task_id_cluster"
More Information needed
KOR-OpenOrca-Platypus
KOR-OpenOrca-Platypus
OpenOrca-Ko + KOpen-platypus
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭
KOpen-platpyus
Repo: KOpen-platypus
고품질 한국어 데이터셋
코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정
1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존
단일 숫자와 영어는 본래의 결과 그대로 가져옴
DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음)
DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정
번역하고자 하는 글자수가 1500자 이상일 경우, API로 변경해서 번역
고유명사는 최대한 유지함
Post-processing… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus.
