multi-human
multilingual-ai-human-detector_xlm-roberta-base_Spydaz_Web_LCARS_Artificial_Human_R1_002-Multi-lingual-GGUFsplicebert-human.510_Spydaz_Web_LCARS_Artificial_Human_R1_002-Multi-lingual-Thinking-GGUFdeltasplice-humanborzoi-humancrypto-destroy-all-humans-multiple-modelsBERT_VDA_MULTIPLEYE_no_HumanRights_seed_45
MultiHumanCarRacing
MultiHumanCarRacing
Documentation: https://github.com/NaOH12/RacingSentimentAnalysis
(meta_data folder is not available since it is being flagged as unsafe. This can be generated/modified with sample_builder.py)
Dataset and code are free to use. Please consider reaching out if you are hiring! 😊
multi_domain_ai_human_text
multi_domain_ai_human_text — Datasheet
Balanced, multi-domain AI-vs-human text detection benchmark with dedicated
out-of-distribution and adversarial evaluation panels. Built by
scripts/build_paper_dataset.py from an 11-corpus unified aggregation.
Splits
Split
AI
Human
Total
Purpose
train
300,000
300,000
600,000
training (balanced, English, clean)
validation
2,996
2,999
5,995
model selection
test
4,991
4,999
9,990
in-distribution test… See the full description on the dataset page: https://huggingface.co/datasets/acmc/multi_domain_ai_human_text.RenderMatte-Human-Multi-Street-2K
RenderMatte Human Multi-Person Street 2K
A synthetic video-matting dataset for the multi-person case. Two or three rigged 3D human characters are
animated and rendered together in one Blender (Cycles) scene, then composited onto real street footage.
Every frame comes with its 16-bit alpha matte.
It is built on top of the single-person VideoMatting pipeline and its source-asset collection, and keeps the
same render and compositing settings wherever possible.
Research use only.… See the full description on the dataset page: https://huggingface.co/datasets/Deanshy1/RenderMatte-Human-Multi-Street-2K.multi-humanevalThis dataset contains a viewer-friendly version of the dataset at mxeval/multi-humaneval with language-specific stop tokens added in. It is made available separately for the convenience of the vllm-code-harness package.
humaneval_multi
humaneval_multi — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror (config humaneval_multi_gen). Source nuprl/MultiPL-E — license MIT, unchanged; all rights remain with the original authors.
Human-Multi-Reference-MT-Benchmark
Human Benchmark Version v3
Parallel sentence benchmarks for the following language pairs:
English–Hindi
English–Telugu
Hindi–Telugu
This repository includes multiple domains, reference translations, and
balanced splits for development and evaluation.
Authors
Vandan Mujadia
Dipti Misra Sharma
Acknowledgment
Developed as part of Himangy, LTRC IIIT H.
Source Data Layout
Raw data is organized by language pair and domain:
English-Hindi/… See the full description on the dataset page: https://huggingface.co/datasets/HimangY/Human-Multi-Reference-MT-Benchmark.
