datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wildjailbreak
WildJailbreak Dataset Card
WildJailbreak is an open-source synthetic safety-training dataset with 262K vanilla (direct harmful requests) and adversarial (complex adversarial jailbreaks) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreaks provides two contrastive types of queries: 1) harmful queries (both vanilla and adversarial) and 2) benign queries that resemble harmful queries in form but contain no harmful intent.
Vanilla Harmful: direct… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wildjailbreak.central-bank-exchange-rates
Central Bank Exchange Rates
Official exchange rates published by 103 central banks and 4 tax authorities, as one CSV per institution: 14,612,276 rows, the oldest series from 1914. Refreshed daily from the GitHub source repository.
Every row is the figure the institution itself published for that date: the ECB euro reference rate, the Federal Reserve H.10 table, the Bank of England spot rates, RBI reference rates, PBoC central parity, HMRC monthly rates for VAT, US Treasury… See the full description on the dataset page: https://huggingface.co/datasets/AllRates/central-bank-exchange-rates.discoverybenchData-driven Discovery Benchmark from the paper:
"DiscoveryBench: Towards Data-Driven Discovery with Large Language Models"
🔭 Overview
DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.platonic-all-experimentsPRISM
PRISM
[Paper] [arXiv] [Project Website]
Purpose-driven Robotic Interaction in Scene Manipulation (PRISM) is a large-scale synthetic dataset for Task-Oriented Grasping featuring cluttered environments and diverse, realistic task descriptions. We use 2365 object instances from ShapeNet-Sem along with stable grasps from ACRONYM to compose 10,000 unique and diverse scenes. Within each scene we capture 10 views, within which there are multiple tasks to be performed. This results in 379k… See the full description on the dataset page: https://huggingface.co/datasets/allenai/PRISM.BenchMIRT-item-statisticsPermitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/BenchMIRT-item-statistics.BenchMIRT-model-statisticsPermitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/BenchMIRT-model-statistics.Allsides_political_bias_properThis is a copy of the "valurank/PoliticalBias_AllSides_Txt" dataset. This is just private repo for my use. I just read all the text files and made them into a csv file. Nothing is different apart from that.
Dataset Card for news-12factor
Dataset Description
~20k articles labeled left, right, or center by the editors of allsides.com.
Languages
The text in the dataset is in English
Dataset Structure
3 folders, with many text files in each. Each text… See the full description on the dataset page: https://huggingface.co/datasets/Faith1712/Allsides_political_bias_proper.The-Bible-KJVklej-polemo2-in
klej-polemo2-in
Description
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains.
We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-in.klej-polemo2-out
klej-polemo2-out
Description
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains.
We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-out.klej-psc
klej-psc
Description
The Polish Summaries Corpus (PSC) is a dataset of summaries for 569 news articles. The human annotators created five extractive summaries for each article by choosing approximately 5% of the original text. A different annotator created each summary. The subset of 154 articles was also supplemented with additional five abstractive summaries each, i.e., not created from the fragments of the original article. In huggingface version of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-psc.klej-dyk
klej-dyk
Description
The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question.
Tasks (input, output, and metrics)
The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.zomato_delivery_EDA📹 Video walkthrough:
Zomato Delivery Operations — EDA & Dataset
Dataset Overview
Real-world delivery data from Zomato operations across multiple Indian cities,
covering courier attributes, weather conditions, traffic density, GPS coordinates,
and delivery outcomes.
Source
Kaggle — saurabhbadole/zomato-delivery-operations-analytics-dataset
Original size
45,584 rows × 20 columns
Final size
38,964 rows × 22 columns
Target variable
Time_taken (min)… See the full description on the dataset page: https://huggingface.co/datasets/allenborochin/zomato_delivery_EDA.passages_gutenberg_populartulu-3-harmbench-evalThis data comes from the HarmBench benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one.
passages_gutenberg_unpopularklej-nkjp-nerpassages_wikipediaall-nli-tr
Dataset Card for AllNLITR
This dataset is a formatted version of NLI-TR datasets, sharing the same licenses. The format is intended to be in line with AllNLI by Sentence Transformers for ease of training.
Despite originally being intended for Natural Language Inference (NLI), this dataset can be used for training/finetuning an embedding model for semantic textual similarity.
Dataset Subsets
pair-class subset
Columns: "premise", "hypothesis", "label"… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/all-nli-tr.klej-cdsc-e
klej-cdsc-e
Description
Polish CDSCorpus consists of 10K Polish sentence pairs which are human-annotated for semantic relatedness (CDSC-R) and entailment (CDSC-E). The dataset may be used to evaluate compositional distributional semantics models of Polish. The dataset was presented at ACL 2017.
Although the SICK corpus inspires the main design of the dataset, it differs in detail. As in SICK, the sentences come from image captions, but the set of chosen images is much… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-cdsc-e.allium-cepa-dataset
Allium cepa - Cell Segmentation and Calssification Dataset
Composed of:
INA : Custom samples collected at the National Water Institute (Argentina)
Onion Cell Merged : A dataset found online (we really want a citation for the people that did this)
Parts:
full_fov : Full field of view of microscope image.
cropped : Possible cell instances cropped from full-fov and normalized.
cell_detection : Dataset to train a cell instance detection supervised model. Test split was selected… See the full description on the dataset page: https://huggingface.co/datasets/GIAR-UTN/allium-cepa-dataset.testset_popqaRAG-Evaluation-Dataset-KO
Allganize RAG Leaderboard
Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다.
RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다.
RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다.
Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.testset_piqaor-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.biographies_yagoklej-allegro-reviewsstarcop_allbands_mini
MINI version of the STARCOP dataset
For full details please refer to https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1
dia-earning21-all
Earnings 21
The Earnings 21 dataset ( also referred to as earnings21 ) is a 39-hour corpus of earnings calls containing entity dense speech from nine different financial sectors. This corpus is intended to benchmark automatic speech recognition (ASR) systems in the wild with special attention towards named entity recognition (NER).
This work has been recently accepted to Interspeech 2021!
File Format Overview
In the following section, we provide an overview of the file… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-earning21-all.
