datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HRSID
HRSID: High-Resolution SAR Images Dataset (Ship Detection)
Unofficial redistribution of the HRSID high-resolution SAR ship-detection dataset, reformatted into a standardized YOLO-compatible directory layout. License status is unclear -- see License before using this beyond research.
Disclaimer
This repository is not an official release of HRSID.
HRSID was created by Shunjun Wei, Xiangfeng Zeng, Qizhe Qu, Mou Wang, Hao Su, and Jun Shi and released via… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/HRSID.PerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.GTA-UAV-HR
GTA-UAV dataset
# Merge splited files
cat drone_part_* > drone.tar.gz
# Extract the archive
tar -xzvf drone.tar.gz
tar -xzvf satellite.tar.gz
For more information, please check our project page.
Sources
Repository: https://github.com/Yux1angJi/GTA-UAV
Paper: https://arxiv.org/abs/2409.16925
ccrfcd-mrms-hrrr-env-2021-2025
1H gauge accumulation + MRMS/HRRR zarr dataset for the Desert Southwest
50+ MRMS+HRRR variables; 220+ gauges; 400k samples
NOTE: work in-progress.
This is a dataset for training and evaluating synthetic quantitative precipiation estimation (QPE) models. Given some input context (e.g., radar fields, envionrmental parameters), predict how much rain fell at a rain gauge site over some period of time. Concretely, we've gather and QC'd data from 220 tipping bucket gauges through… See the full description on the dataset page: https://huggingface.co/datasets/leharris3/ccrfcd-mrms-hrrr-env-2021-2025.href_resultshr-policies-qa-dataset
📚 HR Policies Q&A Dataset
🔎 Overview
This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for:
🤖 LLM fine-tuning
💬 HR & compliance chatbots
🏢 Enterprise policy automation
By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can:
✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/hr-policies-qa-dataset.SWE-bench-plus
SWE-bench-Plus: Test Enhancer
SWE-bench-Plus is a coverage-guided test generation and evaluation layer built on top of the official SWE-bench harness. It automates iterative LLM-based test generation, avoids duplicates, targets uncovered code paths, and stops when coverage plateaus. It is designed for high-throughput, resume-friendly batch runs with robust logging and fault tolerance.
Key Features
Coverage-guided generation: After each iteration, the harness measures… See the full description on the dataset page: https://huggingface.co/datasets/hrtxsny/SWE-bench-plus.BioWiChr-interview-datasetsciercadaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.HRP4K
HRP4K: High-Resolution Road Pothole Detection Dataset
Unofficial redistribution of the HRP4K road pothole-detection dataset (V1.00, Zenodo), under the original CC BY 4.0 license, with a documented upstream train-split completeness gap.
Disclaimer
This repository is not an official release of the HRP4K dataset.
HRP4K was created by Hanshen Chen, Zhoulin Tu, Yu Zhao, and Jianfeng Ye, who retain all copyright and intellectual property rights (to the… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/HRP4K.weibo-opinion-dynamic-single-dim
Weibo Sentiment Evolution Dataset
This dataset contains Weibo posts and their associated comment threads used for studying sentiment evolution and opinion dynamics in social media discussions.
The dataset is distributed as a single JSON Lines file:
weibo_dataset.jsonl
Each line is one Weibo post record. Comments for that post are embedded in the comments field.
Dataset Details
Number of post records: 1,379
Number of embedded comments: 93,569
Number of Weibo… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/weibo-opinion-dynamic-single-dim.NOAA-HRRR-CONUS-ImageCaptionhr-ops-tools
HR-Ops: 8,621 rows of tool calling and cited policy for HR assistants
A training set for HR-operations assistants, built around one idea: make the HR task
objectively checkable. The headline shard is tool calling against authored HR-ops
function schemas, where a correct answer is exact JSON and a wrong one cannot hide behind
fluent prose. Built for the Adaption AutoScientist Challenge, Part 2 (HR).
What this dataset proves, and how you check it
rows
8… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/hr-ops-tools.hrvatski-dataset
Hrvatski dataset
Širok hrvatski korpus za jezičnu prilagodbu i fino ugađanje malih jezičnih
modela, osobito Gemma 3 1B, Gemma 3 4B te kompatibilnih Gemma 4 modela.
Skup nije samo zbirka kratkih uputa. Sastoji se od dva komplementarna dijela:
cpt: 22,5 milijuna riječi književnog, enciklopedijskog i autentičnog
govornog hrvatskog za continued pretraining
sft: 47.588 razgovora za praćenje uputa, prirodne odgovore, dulji tekst,
književni nastavak i razgovorne replike
Za najbolji… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset.TUC-HRI-CS
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the cross-subject validation dataset. For random validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI-CS.acl-arcTUC-HRI
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the random validation dataset. For cross-subject validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI.icml2026-hrtEiSftwk-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
hr-policies-qa-dataset
📚 HR Policies Q&A Dataset
🔎 Overview
This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for:
🤖 LLM fine-tuning
💬 HR & compliance chatbots
🏢 Enterprise policy automation
By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can:
✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/SaumyaGupta/hr-policies-qa-dataset.GuidedBench
GuidedBench
GuidedBench is a guideline-grounded benchmark for evaluating LLM jailbreak
methods. Each question is paired with verified, case-specific entity and
action guidelines describing the content a successful response should
contain.
Dataset structure
core: 180 cases from 15 topics that were consistently refused by the
victim-model families studied in the paper.
additional: 20 cases from five policy-dependent topics, reported
separately because vendor… See the full description on the dataset page: https://huggingface.co/datasets/HRXUST/GuidedBench.hr-practitioner-seed-v1
HR Practitioner Seed (v2)
A curated instruction-tuning seed for HR and recruiting assistants, built for the
Adaption AutoScientist Challenge
(Part 2, HR track).
3,059 training rows + 244 held-out evaluation rows.
What this is
Group
Rows
Notes
Grounded recruiting tasks
~2,692
real job adverts as context, 8 task types, 2 markets
Generic HR policy questions
512
questions only
Authored generalist HR questions
99
original, 12 practice areas
All… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/hr-practitioner-seed-v1.hr-jd-bias-audit
JD-BiasAudit
JD-BiasAudit is a provenance-tracked instruction-tuning dataset for HR teams, compliance reviewers, and model builders who need to neutralize coded language in job descriptions without deleting legitimate requirements. It derives structured, span-grounded audits from real postings in lang-uk/recruitment-dataset-job-descriptions-english. The upstream corpus provides job descriptions, not paired neutral rewrites, exact removed spans, protected-attribute proxy… See the full description on the dataset page: https://huggingface.co/datasets/0xkamal7/hr-jd-bias-audit.hr-policy-dataset
HR Policy Dataset
Description
A dataset of HR-related question-answer pairs used to fine-tune a Qwen2.5-based HR assistant.
Topics
Working hours
Leave policies
Employee benefits
Onboarding
Workplace conduct
Company procedures
Format
{
"instruction": "What are the official working hours?",
"response": "Employees work from 9:00 AM to 6:00 PM..."
}
pt-br-moderation-eval
Dataset de Validação: Moderação
Este dataset contém 1680 exemplos de moderação traduzidos do EN-US para o PT-BR, com foco em manter a toxicidade e vulgaridade original sem qualquer suavização. Foi desenvolvido para treinar e validar sistemas de moderação que precisam lidar com gírias brasileiras e conteúdo altamente tóxico de forma precisa.
Categorias de Moderação (Labels):
sexual: Conteúdo sexual explícito ou serviços sexuais.
ódio: Conteúdo de ódio baseado em… See the full description on the dataset page: https://huggingface.co/datasets/HRB25/pt-br-moderation-eval.HR-Conflict-Dataset-V2
HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored)
Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license.
The generator is the product
This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/ConsumerDividends/HR-Conflict-Dataset-V2.hr-practitioner-adapted-v1
HR Practitioner (Adaption-adapted) v1
The adapted dataset used to fine-tune our HR and people operations model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/hr-practitioner-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
3,059 rows. Adaption writes its output to enhanced_prompt / enhanced_completion
and leaves the uploaded prompt /… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/hr-practitioner-adapted-v1.LiFT-HRA
LiFT-HRA
Dataset Summary
This dataset is derived from LiFT-HRA-20K for our UnifiedReward-7B training.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/pdf/2503.05236
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/LiFT-HRA.Norwegian-Synthetic-HR-data-v-1
Synthetic norwegian public sector HR dataset
Dataset description
This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector.
The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use.
The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.
