datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_domain_ai_human_text
multi_domain_ai_human_text — Datasheet
Balanced, multi-domain AI-vs-human text detection benchmark with dedicated
out-of-distribution and adversarial evaluation panels. Built by
scripts/build_paper_dataset.py from an 11-corpus unified aggregation.
Splits
Split
AI
Human
Total
Purpose
train
300,000
300,000
600,000
training (balanced, English, clean)
validation
2,996
2,999
5,995
model selection
test
4,991
4,999
9,990
in-distribution test… See the full description on the dataset page: https://huggingface.co/datasets/acmc/multi_domain_ai_human_text.multi-humanevalThis dataset contains a viewer-friendly version of the dataset at mxeval/multi-humaneval with language-specific stop tokens added in. It is made available separately for the convenience of the vllm-code-harness package.
human_multi_classifications_500human_multi_classificationsgpt_4o_mini_classifications_multi_humanspeaker_evaluation_multi_test_v0
Seamless Interaction Pairs
This dataset contains paired query and document audio clips for interaction-based
speaker evaluation. Each row describes a query clip and a related document clip,
with segment metadata and durations for analysis.
Data structure
The dataset uses a single split stored in data.parquet.
Audio files are stored under audio/ and referenced by relative paths in the
parquet file.
Columns
pair_id (string): Pair identifier.
interaction… See the full description on the dataset page: https://huggingface.co/datasets/humanify/speaker_evaluation_multi_test_v0.qwen2_72b_classifications_multi_humanllama_3_1_8b_classifications_multi_humangpt_4o_classifications_multi_humanmistral_large_classifications_multi_humangemini_1_5_pro_classifications_multi_humanllama_3_1_405b_classifications_multi_humancommand_r_plus_classifications_multi_humanllama_3_1_70b_classifications_multi_humanasia-humanitarian-needs-2022-multi-sectoral-needs-assessment-occ
occupied Palestinian territory (oPt) - 2022 Multi-Sectoral Needs Assessment
Publisher: REACH Initiative · Source: HDX · License: cc-by-igo · Updated: 2024-06-07
Abstract
The 2022 Multi-Sector Needs Assessment (MSNA), conducted by the REACH Initiative in close collaboration with the United Nations Office for the Coordination of Humanitarian Affairs (OCHA), aims to identify and assess multi-sectoral and sector-specific needs, circumstances, and vulnerabilities of… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-humanitarian-needs-2022-multi-sectoral-needs-assessment-occ.MULTI_VALUE_mnli_plural_to_singular_human
Dataset Card for "MULTI_VALUE_mnli_plural_to_singular_human"
More Information needed
MULTI_VALUE_cola_plural_to_singular_human
Dataset Card for "MULTI_VALUE_cola_plural_to_singular_human"
More Information needed
MULTI_VALUE_mrpc_plural_to_singular_human
Dataset Card for "MULTI_VALUE_mrpc_plural_to_singular_human"
More Information needed
MULTI_VALUE_wnli_plural_to_singular_human
Dataset Card for "MULTI_VALUE_wnli_plural_to_singular_human"
More Information needed
MULTI_VALUE_sst2_plural_to_singular_human
Dataset Card for "MULTI_VALUE_sst2_plural_to_singular_human"
More Information needed
MULTI_VALUE_qqp_plural_to_singular_human
Dataset Card for "MULTI_VALUE_qqp_plural_to_singular_human"
More Information needed
MULTI_VALUE_rte_plural_to_singular_human
Dataset Card for "MULTI_VALUE_rte_plural_to_singular_human"
More Information needed
MULTI_VALUE_stsb_plural_to_singular_human
Dataset Card for "MULTI_VALUE_stsb_plural_to_singular_human"
More Information needed
gemini_1_5_pro_classifications_multi_human
