datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.vindr-cxr-testsettestset_popqatestset_piqatestset_mmlutestset_hellaswagtestset_winogrande-infilltestset_ellietestset_winogrande-mcqtestset_munchDr.Sparse-OTF-test-set
Dr.Sparse OTF Test Set
100 sparse matrices from the SuiteSparse Matrix Collection,
converted to the flat binary format the Dr.Sparse
benchmark harness reads. This is the held-out evaluation set for LLM-generated
CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the
models were developed against.
Layout
Matrices are grouped into size tiers by row count, the convention Dr.Sparse task
discovery scans for:
tier
rows
matrices
size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.Pharmacology-LLM-test-setPharmacology-LLM-test-set: A test set for a large language model focused on pharmacology tasks
1 Inroduction
Large language models (LLM), including ChatGPT, have fundamentally transformed the knowledge query schemes and methods in pharmacology for pharmacologists, drug researchers, clinical drug researchers, and artificial intelligence researchers in pharmacology. They can conduct multi-round consultations and query pharmacological issues in a question-and-answer format. However… See the full description on the dataset page: https://huggingface.co/datasets/zhangyingbo1984/Pharmacology-LLM-test-set.hate_speech_open_data_original_class_test_setVox2_testsettest_setcommon_voice_16_1_spanish_test_set
Dataset Card for Common Voice Corpus 16 Spanish Dataset
Acknowledgement
The dataset belongs to COMMON VOICE MOZILLA FOUNDATION.
I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main)
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Languages
Spanish
How to use
The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.uspto_1k_tpl_randomly_selected_10_classes_test_setovos-intents-ilenia-testset-caovos-intents-ilenia-testset-nlg_test_set_newtoxic-detection-testset-perturbations
Dataset Card for toxic-detection-testset-perturnations
Dataset Summary
This dataset a test set for toxic detection that contains both clean data and it's perturbed version with human-written perturbations online.
In addition, our dataset can be used to benchmark misspelling correctors as well.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
Data Instances
{
"clean_version": "this… See the full description on the dataset page: https://huggingface.co/datasets/yiran223/toxic-detection-testset-perturbations.translated_facts_test_setovos-intents-ilenia-testset-eswithin_family_test_setg_test_set_fewDTE-testsetaugmented_posts_test_setcb_test_setcustom_test_setHPA_Test_Set
