datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vindr-cxr-testsetDr.Sparse-OTF-test-set
Dr.Sparse OTF Test Set
100 sparse matrices from the SuiteSparse Matrix Collection,
converted to the flat binary format the Dr.Sparse
benchmark harness reads. This is the held-out evaluation set for LLM-generated
CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the
models were developed against.
Layout
Matrices are grouped into size tiers by row count, the convention Dr.Sparse task
discovery scans for:
tier
rows
matrices
size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.hate_speech_open_data_original_class_test_setVox2_testsetcommon_voice_16_1_spanish_test_set
Dataset Card for Common Voice Corpus 16 Spanish Dataset
Acknowledgement
The dataset belongs to COMMON VOICE MOZILLA FOUNDATION.
I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main)
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Languages
Spanish
How to use
The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.toxic-detection-testset-perturbations
Dataset Card for toxic-detection-testset-perturnations
Dataset Summary
This dataset a test set for toxic detection that contains both clean data and it's perturbed version with human-written perturbations online.
In addition, our dataset can be used to benchmark misspelling correctors as well.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
Data Instances
{
"clean_version": "this… See the full description on the dataset page: https://huggingface.co/datasets/yiran223/toxic-detection-testset-perturbations.HPA_Test_Sethans_reduced_testsetadaptation-test-setThis is the training data for the models of the project.
helpfulness_test_set
Dataset Card: Helpfulness Classification Based on ALERT Dataset
Dataset Description
This dataset is derived from the ALERT dataset and has been labeled to assess whether responses in question-answer pairs are helpful or not.
Key Features:
Helpfulness Labeling: Each answer is classified as either:
Helpful: This includes both positive and supportive answers as well as well-justified rejections.
Not Helpful: Answers that lack relevance, clarity, or a justified… See the full description on the dataset page: https://huggingface.co/datasets/julius8787/helpfulness_test_set.hans_testsetMNLI_Binary_Testsetsarcastic_test_setfake-news-testsetnli_binary_testsetTestSetchexpert-test-sethelpfulness_test_set
Dataset Card: Helpfulness Classification Based on ALERT Dataset
Dataset Description
This dataset is derived from the ALERT dataset and has been labeled to assess whether responses in question-answer pairs are helpful or not.
Key Features:
Helpfulness Labeling: Each answer is classified as either:
Helpful: This includes both positive and supportive answers as well as well-justified rejections.
Not Helpful: Answers that lack relevance, clarity, or a justified… See the full description on the dataset page: https://huggingface.co/datasets/juliushase/helpfulness_test_set.
