datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Schalk-2009-EEGMotorMovementImagery
EEG Motor Movement/Imagery Dataset
This is an unofficial mirror of the PhysioNet EEG Motor Movement/Imagery
Dataset, version 1.0.0. It is not affiliated with or endorsed by the dataset
contributors, their institutions, or PhysioNet.
Source and documentation
Original dataset: PhysioNet EEGMMIDB v1.0.0
Source version: 1.0.0, published September 9, 2009
Original publication: Schalk et al., IEEE Transactions on Biomedical Engineering (2004)
BCI2000 project:… See the full description on the dataset page: https://huggingface.co/datasets/raei/Schalk-2009-EEGMotorMovementImagery.HAE_RAE_BENCH_1.1The HAE_RAE_BENCH 1.1 is an ongoing project to develop a suite of evaluation tasks designed to test the
understanding of models regarding Korean cultural and contextual nuances.
Currently, it comprises 13 distinct tasks, with a total of 4900 instances.
Please note that although this repository contains datasets from the original HAE-RAE BENCH paper,
the contents are not completely identical. Specifically, the reading comprehension subset from the original version has been removed due to… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.1.Jeong-2022-InternationalBCICompetition2020Review-track3
BCI Competition 2020 Track 3: imagined speech classification
This is an unofficial mirror of Track 3 only from the 2020 International BCI
Competition, described by Jeong et al. (2022). It is not affiliated with or
endorsed by the authors, their institutions, or OSF.
Source and attribution
Original data: 2020 International BCI Competition, OSF project pq7vb,
folder Track#3 Imagined speech classification.
Paper: Jeong et al., 2020 International brain–computer… See the full description on the dataset page: https://huggingface.co/datasets/raei/Jeong-2022-InternationalBCICompetition2020Review-track3.NIFTY
The News-Informed Financial Trend Yield (NIFTY) Dataset.
The News-Informed Financial Trend Yield (NIFTY) Dataset. Details of the dataset, including data procurement and filtering can be found in the paper here: https://arxiv.org/abs/2405.09747.
For the NIFTY-RL LLM alignment dataset please use nifty-rl.
📋 Table of Contents
🧩 NIFTY Dataset
📋 Table of Contents
📖 Usage
Downloading the dataset
Dataset structure
Large Language Models
✍️ Contributing
📝 Citing
🙏… See the full description on the dataset page: https://huggingface.co/datasets/raeidsaqur/NIFTY.HAE_RAE_BENCH_1.0The HAE_RAE_BENCH 1.0 is the original implementation of the dataset froom the paper: HAE-RAE BENCH paper.
The benchmark is a collection of 1,538 instances across 6 tasks: standard_nomenclature, loan_word, rare_word, general_knowledge, history and reading comprehension.
To replicate the studies from the paper, see below.
Dataset Overview
Task
Instances
Version
Explanation
standard_nomenclature
153
v1.0
Multiple-choice questions about Korean standard nomenclatures from… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.0.Nieto-2022-ThinkingOutLoudOpenAccessEEGBasedBCIDatasetInnerSpeech
Thinking out loud: an open-access EEG-based BCI dataset for inner speech recognition
This is an unofficial mirror of OpenNeuro dataset ds003626, version
2.1.2. It is not affiliated with or endorsed by the dataset authors, their
institutions, or OpenNeuro.
Source and documentation
Original dataset: OpenNeuro ds003626 v2.1.2
Article: Nieto et al., Scientific Data (2022)
Original dataset documentation: README
Official analysis code: N-Nieto/Inner_Speech_Dataset
The… See the full description on the dataset page: https://huggingface.co/datasets/raei/Nieto-2022-ThinkingOutLoudOpenAccessEEGBasedBCIDatasetInnerSpeech.nifty-rl
The News-Informed Financial Trend Yield (NIFTY) Dataset.
The News-Informed Financial Trend Yield (NIFTY) Dataset. Details of the dataset, including data procurement and filtering can be found in the paper here: https://arxiv.org/abs/2405.09747.
📋 Table of Contents
🧩 NIFTY Dataset
📋 Table of Contents
📖 Usage
Downloading the dataset
Dataset structure
Large Language Models
✍️ Contributing
📝 Citing
🙏 Acknowledgements
📖 Usage
Downloading and using… See the full description on the dataset page: https://huggingface.co/datasets/raeidsaqur/nifty-rl.Hansard
Pedagogical Machine Translation (Dialect) dataset: the filtered Canadian Hansard Dataset.
The Canadian [Hansard](https://www.ourcommons.ca/documentviewer/en/35-2/house/hansard-index) is an archive of parliamentary sessions in the two official languages in Canada - English and Franch.
📋 Table of Contents
🧩 Hansard Dataset
📋 Table of Contents
📖 Usage
Downloading the dataset
Dataset structure
Loading the dataset
Loading the dataset
The three partitions… See the full description on the dataset page: https://huggingface.co/datasets/raeidsaqur/Hansard.HAE-RAE-COT-1.5M
Dataset Card for "HAE-RAE-COT-1.5M"
HAE-RAE-COT-1.5M is a dataset encompassing 1,586,688 samples of questions paired with CoT (Chain of Thought) rationales.
The majority of this dataset is a translation of samples from the CoT-Collection, with a portion of samples derived from Korean datasets through the utilization of the gpt-3.5-turbo API. The translation of the CoT-Collection was carried out using the NLLB 600M model.
To the best of our knowledge, HAE-RAE-COT-1.5M represents the… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE-RAE-COT-1.5M.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
HAE_RAE_BENCH_2.0HAE_RAE_BENCH 2.0 is a miny implementation of Big-Bench consisted of 5 tasks: date_understanding, context_definition_alignment, proverb_unscrambling, 2_digit_multiply,
and 3_digit_subtract.
Paper Coming Soon (probably).
Satori_RL_data_with_RAEarticles_2024-10-19Brunner-2008-BCICompetition2008GrazDataSetA
BCI Competition IV Dataset 2a — Graz data set A
This is an unofficial mirror of the dataset described by C. Brunner,
R. Leeb, G. R. Müller-Putz, A. Schlögl, and G. Pfurtscheller (2008) at Graz
University of Technology. Originally released as BCI Competition IV Dataset
2a (BCICIV2A), it is distributed by BNCI Horizon 2020 as 001-2014.
This mirror is not affiliated with or endorsed by the dataset contributors,
their institutions, BNCI Horizon 2020, or the competition organizers.… See the full description on the dataset page: https://huggingface.co/datasets/raei/Brunner-2008-BCICompetition2008GrazDataSetA.articles_2024-10-21article_inferences_2024-10-19Rae-Taylor-Lora-Dataarticle_inferences_2024-10-20articles_2024-10-22articles_2024-10-23article_inferences_2024-10-17article_inferences_2024-10-24custom_llama_kor
Dataset Card for "custom_llama_kor"
More Information needed
article_inferences_2024-10-18articles_2024-10-24en_hi_translationarticles_inferredarticles_2024-10-16article_inferences_2024-10-16articles_2024-10-17
