datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HRM-He-corpus-objective
Hebrew reasoning traces
Generated Hebrew chain-of-thought over code, cybersecurity, agentic, math and
general-reasoning seeds. Built for a Hebrew/English code-specialised LM, where
off-the-shelf Hebrew reasoning data is effectively nonexistent.
What the default config contains
Every row the training corpus keeps -- not a filtered highlight reel. Two things
are disqualifying and are absent: a wrong final answer (answer_ok is False), and
Arabic drift. Everything… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/HRM-He-corpus-objective.HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.hebrew-hrm-corpus
Hebrew HRM-Text Corpus
Training corpus for a Hebrew Hierarchical Reasoning Model, replicating the
sapientinc/HRM-Text-1B recipe
(train-from-scratch, PrefixLM over {condition, instruction, response}, loss on response only).
Schema
Each line: {"condition": "<tags>", "instruction": "...", "response": "..."}.
Condition tags map to special tokens: direct→<|object_ref_start|>, cot→<|object_ref_end|>,
noisy→<|quad_start|>, synth→<|quad_end|> (composite tags… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-hrm-corpus.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.task1283_hrngo_quality_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1283_hrngo_quality_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1283_hrngo_quality_classification.hrwacThe Croatian web corpus hrWaC was built by crawling the .hr top-level domain in 2011 and again in 2014. The corpus was near-deduplicated on paragraph level, normalised via diacritic restoration, morphosyntactically annotated and lemmatised. The corpus is shuffled by paragraphs. Each paragraph contains metadata on the URL, domain and language identification (Croatian vs. Serbian).
Version 2.0 of this corpus is described in http://www.aclweb.org/anthology/W14-0405. Version 2.1 contains newer and better linguistic annotations.task1186_nne_hrngo_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1186_nne_hrngo_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1186_nne_hrngo_classification.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.weibo-opinion-dynamic-single-dim
Weibo Sentiment Evolution Dataset
This dataset contains Weibo posts and their associated comment threads used for studying sentiment evolution and opinion dynamics in social media discussions.
The dataset is distributed as a single JSON Lines file:
weibo_dataset.jsonl
Each line is one Weibo post record. Comments for that post are embedded in the comments field.
Dataset Details
Number of post records: 1,379
Number of embedded comments: 93,569
Number of Weibo… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/weibo-opinion-dynamic-single-dim.hr-ops-tools
HR-Ops: 8,621 rows of tool calling and cited policy for HR assistants
A training set for HR-operations assistants, built around one idea: make the HR task
objectively checkable. The headline shard is tool calling against authored HR-ops
function schemas, where a correct answer is exact JSON and a wrong one cannot hide behind
fluent prose. Built for the Adaption AutoScientist Challenge, Part 2 (HR).
What this dataset proves, and how you check it
rows
8… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/hr-ops-tools.hrvatski-dataset
Hrvatski dataset
Širok hrvatski korpus za jezičnu prilagodbu i fino ugađanje malih jezičnih
modela, osobito Gemma 3 1B, Gemma 3 4B te kompatibilnih Gemma 4 modela.
Skup nije samo zbirka kratkih uputa. Sastoji se od dva komplementarna dijela:
cpt: 22,5 milijuna riječi književnog, enciklopedijskog i autentičnog
govornog hrvatskog za continued pretraining
sft: 47.588 razgovora za praćenje uputa, prirodne odgovore, dulji tekst,
književni nastavak i razgovorne replike
Za najbolji… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset.service-hrlpt
传奇私服分布式路由与自动化接口索引库 - Batch 007
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 传奇sf一条龙搭建 - 传奇sf一条龙网站 —— 承载源站 🌐 agergames.com
👉 传奇sf发布站程序 - 传奇sf发布站搭建 —— 承载源站 🌐 alien.cneagames.com
👉 传奇sf服务端网盘 - 传奇sf补丁网盘 —— 承载源站 🌐 alien.dunia-games.com
👉 传奇sf发布站模板 - 传奇sf排行榜 —— 承载源站 🌐 alien.duyougame.com
👉 传奇sf素材 - 传奇sf素材网 —— 承载源站 🌐 alien.fysagame.com
👉 传奇sf版本下载 - 传奇sf商业版本 —— 承载源站 🌐 alien.gamerping.com
👉 传奇sf网盘下载 - 传奇sf版本网盘 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/kusesde/service-hrlpt.Bns_Law_Rag_DB
BNS Law RAG Dataset
Dataset: Hrutik2003/Bns_Law_Rag_DBPurpose: Text corpus for RAG systems on the new Indian criminal laws (BNS, BNSS, BSA).
Dataset Summary
This dataset contains cleaned and processed text extracted from the new Indian criminal laws introduced in 2023 and old IPC laws:
Bharatiya Nyaya Sanhita (BNS) 2023
Bharatiya Nagarik Suraksha Sanhita (BNSS) 2023
Bharatiya Sakshya Adhiniyam (BSA) 2023
The Code of Criminal Procedure (CrPC)
The Indian Penal Code (IPC)… See the full description on the dataset page: https://huggingface.co/datasets/Hrutik2003/Bns_Law_Rag_DB.href_preference
HREF: Human Reference-Guided Evaluation of Instruction Following in Language Models
📑 Paper | 🤗 Leaderboard | 📁 Codebase
HREF is evaluation benchmark that evaluates language models' capacity of following human instructions. This dataset contains the human agreement set of HREF, which contains 1,752 pairs of language model outputs along with the preference data from 4 human annotators for each model pairs. The dataset contains 438 instructions human-written instruction and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/href_preference.href
HREF: Human Reference-Guided Evaluation of Instruction Following in Language Models
📑 Paper | 🤗 Leaderboard | 📁 Codebase
HREF is evaluation benchmark that evaluates language models' capacity of following human instructions. This dataset contains the validation set of HREF, which contains 430 human-written instruction and response pairs from the test split of No Robots, covering 8 categories (removing Coding and Chat).
For each instruction, we generate a baseline model… See the full description on the dataset page: https://huggingface.co/datasets/allenai/href.hrm-text-code-tools-sft
HRM-Text Code & Tools SFT
Curated coding and fixed-runtime tool-use supervised fine-tuning data for
HRM-Text-1B. This sealed
research release is not a general chat corpus. It contains only canonical
v2 records—there are no legacy prefix-field files.
Contents
Stage
Rows
Train
Validation
Serialized cap
Train SHA-256
stage-a
1,157,064
1,133,817
23,247
4,096
b6e340ea64568570b24e2d5f08c506a6f3221482d9fca6778a49c5577832171c
stage-b
283,547
278,018
5,529
4… See the full description on the dataset page: https://huggingface.co/datasets/pzarzycki/hrm-text-code-tools-sft.Loom-HR
Loom-HR - Local-first, human-in-the-loop AI hiring assistant. The AI advises with reasons; a human makes the call.
Mirrored in two places, same version everywhere:
Hugging Face (you are here) · GitLab
Get the files - GitLab is the most reliable plain-git route:
git clone https://gitlab.com/LucioLiu/Loom-HR.git
# or from this page
hf download LucioLiu/Loom-HR --repo-type dataset --local-dir ./Loom-HR
Licence: PolyForm Noncommercial 1.0.0 - the LICENSE file in this repo is authoritative.… See the full description on the dataset page: https://huggingface.co/datasets/LucioLiu/Loom-HR.hrvatski-dataset-v2
Hrvatski Dataset v2
Comprehensive Croatian (hrvatski) language dataset with 455 examples, optimized for
chatbot training and article generation. Every example is a question/answer (or
instruction/output) pair written in natural Croatian.
This repo is a single-source-of-truth catalog: all format files are generated from one
canonical file (source/hrvatski_dataset_v2.jsonl), so every format is guaranteed to
contain the exact same 455 examples. If a format is ever out of sync, that… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset-v2.hrm-text-opus46-math-coding
YL95/hrm-text-opus46-math-coding
This dataset keeps only Opus 4.6 math, coding, and nearby technical reasoning tasks from the requested source datasets.
Contents
prompt_completion/train: the main training split for base-model fine-tuning
prompt_completion/over_4096_tokens: rows longer than the token limit
chat/train: a message-form version of the same kept rows
chat/over_4096_tokens: the message-form over-limit subset
Notes
HRM-Text-1B is a base… See the full description on the dataset page: https://huggingface.co/datasets/YL95/hrm-text-opus46-math-coding.hr-practitioner-seed-v1
HR Practitioner Seed (v2)
A curated instruction-tuning seed for HR and recruiting assistants, built for the
Adaption AutoScientist Challenge
(Part 2, HR track).
3,059 training rows + 244 held-out evaluation rows.
What this is
Group
Rows
Notes
Grounded recruiting tasks
~2,692
real job adverts as context, 8 task types, 2 markets
Generic HR policy questions
512
questions only
Authored generalist HR questions
99
original, 12 practice areas
All… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/hr-practitioner-seed-v1.hr-jd-bias-audit
JD-BiasAudit
JD-BiasAudit is a provenance-tracked instruction-tuning dataset for HR teams, compliance reviewers, and model builders who need to neutralize coded language in job descriptions without deleting legitimate requirements. It derives structured, span-grounded audits from real postings in lang-uk/recruitment-dataset-job-descriptions-english. The upstream corpus provides job descriptions, not paired neutral rewrites, exact removed spans, protected-attribute proxy… See the full description on the dataset page: https://huggingface.co/datasets/0xkamal7/hr-jd-bias-audit.hr-policy-dataset
HR Policy Dataset
Description
A dataset of HR-related question-answer pairs used to fine-tune a Qwen2.5-based HR assistant.
Topics
Working hours
Leave policies
Employee benefits
Onboarding
Workplace conduct
Company procedures
Format
{
"instruction": "What are the official working hours?",
"response": "Employees work from 9:00 AM to 6:00 PM..."
}
HR-Conflict-Dataset-V2
HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored)
Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license.
The generator is the product
This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/ConsumerDividends/HR-Conflict-Dataset-V2.hr-practitioner-adapted-v1
HR Practitioner (Adaption-adapted) v1
The adapted dataset used to fine-tune our HR and people operations model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/hr-practitioner-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
3,059 rows. Adaption writes its output to enhanced_prompt / enhanced_completion
and leaves the uploaded prompt /… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/hr-practitioner-adapted-v1.Norwegian-Synthetic-HR-data-v-1
Synthetic norwegian public sector HR dataset
Dataset description
This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector.
The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use.
The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.ICLEC
Indian Supreme Court Legal Corpus – CaseBridge Edition
This dataset presents a structured, passage-level corpus of Indian Supreme Court legal judgments. It is specifically curated and annotated for machine learning, information retrieval, and exploratory data analysis (EDA) tasks, and is a core resource in the CaseBridge project.
Dataset Overview
This corpus enables high-quality benchmarking and experimentation in:
Information Retrieval: Passage and document ranking… See the full description on the dataset page: https://huggingface.co/datasets/HRITHIKRAJ2537H/ICLEC.HRM-Text-data-io-cleaned-20260515-copyPre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/huankguan2/HRM-Text-data-io-cleaned-20260515-copy.swe-bench-arkts
SWE-bench ArkTS
SWE-bench-style tasks mined from ArkTS pull requests and direct commits in HarmonyOS/OpenHarmony projects. The dataset has two evaluator-specific splits and 21 tasks in total.
Split
Tasks
Oracle
failed2pass
14
Hvigor case-level unit tests: 50 FAIL_TO_PASS, 8,653 PASS_TO_PASS, no regressions
compile2pass
7
ArkTS compiler/API transitions: 5 diagnostic-removal tasks and 2 SDK/toolchain migrations
Loading
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/swe-bench-arkts.task1284_hrngo_informativeness_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1284_hrngo_informativeness_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1284_hrngo_informativeness_classification.HR-Conflict-Dataset-V2
HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored)
Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license.
The generator is the product
This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/CDividends/HR-Conflict-Dataset-V2.
