datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.ccrfcd-mrms-hrrr-env-2021-2025
1H gauge accumulation + MRMS/HRRR zarr dataset for the Desert Southwest
50+ MRMS+HRRR variables; 220+ gauges; 400k samples
NOTE: work in-progress.
This is a dataset for training and evaluating synthetic quantitative precipiation estimation (QPE) models. Given some input context (e.g., radar fields, envionrmental parameters), predict how much rain fell at a rain gauge site over some period of time. Concretely, we've gather and QC'd data from 220 tipping bucket gauges through… See the full description on the dataset page: https://huggingface.co/datasets/leharris3/ccrfcd-mrms-hrrr-env-2021-2025.href_resultsBioWiCweibo-opinion-dynamic-single-dim
Weibo Sentiment Evolution Dataset
This dataset contains Weibo posts and their associated comment threads used for studying sentiment evolution and opinion dynamics in social media discussions.
The dataset is distributed as a single JSON Lines file:
weibo_dataset.jsonl
Each line is one Weibo post record. Comments for that post are embedded in the comments field.
Dataset Details
Number of post records: 1,379
Number of embedded comments: 93,569
Number of Weibo… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/weibo-opinion-dynamic-single-dim.TUC-HRI-CS
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the cross-subject validation dataset. For random validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI-CS.acl-arcTUC-HRI
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the random validation dataset. For cross-subject validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI.pt-br-moderation-eval
Dataset de Validação: Moderação
Este dataset contém 1680 exemplos de moderação traduzidos do EN-US para o PT-BR, com foco em manter a toxicidade e vulgaridade original sem qualquer suavização. Foi desenvolvido para treinar e validar sistemas de moderação que precisam lidar com gírias brasileiras e conteúdo altamente tóxico de forma precisa.
Categorias de Moderação (Labels):
sexual: Conteúdo sexual explícito ou serviços sexuais.
ódio: Conteúdo de ódio baseado em… See the full description on the dataset page: https://huggingface.co/datasets/HRB25/pt-br-moderation-eval.HR-Conflict-Dataset-V2
HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored)
Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license.
The generator is the product
This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/ConsumerDividends/HR-Conflict-Dataset-V2.HR-Conflict-Dataset-V2
HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored)
Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license.
The generator is the product
This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/CDividends/HR-Conflict-Dataset-V2.CUNI-MH-v2-csde-datahatemmahmoud__qwen2.5-1.5b-sft-raft-grpo-hra-doc-details
Dataset Card for Evaluation run of hatemmahmoud/qwen2.5-1.5b-sft-raft-grpo-hra-doc
Dataset automatically created during the evaluation run of model hatemmahmoud/qwen2.5-1.5b-sft-raft-grpo-hra-doc
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hatemmahmoud__qwen2.5-1.5b-sft-raft-grpo-hra-doc-details.CUNI-MH-v2-encs-dataweibo-opinion-dynamic-multi-dim
