CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01guychuk /HRM-He-corpus-objective Hebrew reasoning traces Generated Hebrew chain-of-thought over code, cybersecurity, agentic, math and general-reasoning seeds. Built for a Hebrew/English code-specialised LM, where off-the-shelf Hebrew reasoning data is effectively nonexistent. What the default config contains Every row the training corpus keeps -- not a filtered highlight reel. Two things are disqualifying and are absent: a wrong final answer (answer_ok is False), and Arabic drift. Everything… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/HRM-He-corpus-objective.tabulartext-generation100K<n<1M0 likes4.6k downloads11d agoHugging Face02sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.5k downloads4mo agoHugging Face03guychuk /hebrew-hrm-corpus Hebrew HRM-Text Corpus Training corpus for a Hebrew Hierarchical Reasoning Model, replicating the sapientinc/HRM-Text-1B recipe (train-from-scratch, PrefixLM over {condition, instruction, response}, loss on response only). Schema Each line: {"condition": "<tags>", "instruction": "...", "response": "..."}. Condition tags map to special tokens: direct→<|object_ref_start|>, cot→<|object_ref_end|>, noisy→<|quad_start|>, synth→<|quad_end|> (composite tags… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-hrm-corpus.text-generation0 likes2.8k downloads3mo agoHugging Face04abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes708 downloads3mo agoHugging Face05Lots-of-LoRAs /task1283_hrngo_quality_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1283_hrngo_quality_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1283_hrngo_quality_classification.texttext-generation1K<n<10K0 likes192 downloads2y agoHugging Face06community-datasets /hrwacThe Croatian web corpus hrWaC was built by crawling the .hr top-level domain in 2011 and again in 2014. The corpus was near-deduplicated on paragraph level, normalised via diacritic restoration, morphosyntactically annotated and lemmatised. The corpus is shuffled by paragraphs. Each paragraph contains metadata on the URL, domain and language identification (Croatian vs. Serbian). Version 2.0 of this corpus is described in http://www.aclweb.org/anthology/W14-0405. Version 2.1 contains newer and better linguistic annotations.text-generation1B<n<10B0 likes154 downloads3y agoHugging Face07Lots-of-LoRAs /task1186_nne_hrngo_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1186_nne_hrngo_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1186_nne_hrngo_classification.texttext-generation1K<n<10K0 likes112 downloads2y agoHugging Face08miscusi /adaption-hr-advisory-onet HR Advisory Instruction Dataset (O*NET-grounded) Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record. Built for the Adaption Labs AutoScientist Challenge Part 2, HR track. What is in it Rows 5,415 (4,836 train / 579 eval) Task families 19 Occupations covered 907 of 923 available Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.texttext-generation1K<n<10K0 likes73 downloads2mo agoHugging Face09hreyulog /weibo-opinion-dynamic-single-dim Weibo Sentiment Evolution Dataset This dataset contains Weibo posts and their associated comment threads used for studying sentiment evolution and opinion dynamics in social media discussions. The dataset is distributed as a single JSON Lines file: weibo_dataset.jsonl Each line is one Weibo post record. Comments for that post are embedded in the comments field. Dataset Details Number of post records: 1,379 Number of embedded comments: 93,569 Number of Weibo… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/weibo-opinion-dynamic-single-dim.tabulartext-classification1K<n<10K0 likes66 downloads3mo agoHugging Face10Jainamshahhh /hr-ops-tools HR-Ops: 8,621 rows of tool calling and cited policy for HR assistants A training set for HR-operations assistants, built around one idea: make the HR task objectively checkable. The headline shard is tool calling against authored HR-ops function schemas, where a correct answer is exact JSON and a wrong one cannot hide behind fluent prose. Built for the Adaption AutoScientist Challenge, Part 2 (HR). What this dataset proves, and how you check it rows 8… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/hr-ops-tools.texttext-generation1K<n<10K0 likes61 downloads2mo agoHugging Face11administraktor /hrvatski-dataset Hrvatski dataset Širok hrvatski korpus za jezičnu prilagodbu i fino ugađanje malih jezičnih modela, osobito Gemma 3 1B, Gemma 3 4B te kompatibilnih Gemma 4 modela. Skup nije samo zbirka kratkih uputa. Sastoji se od dva komplementarna dijela: cpt: 22,5 milijuna riječi književnog, enciklopedijskog i autentičnog govornog hrvatskog za continued pretraining sft: 47.588 razgovora za praćenje uputa, prirodne odgovore, dulji tekst, književni nastavak i razgovorne replike Za najbolji… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset.texttext-generation100K<n<1M0 likes50 downloads2mo agoHugging Face12kusesde /service-hrlpt 传奇私服分布式路由与自动化接口索引库 - Batch 007 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 传奇sf一条龙搭建 - 传奇sf一条龙网站 —— 承载源站 🌐 agergames.com 👉 传奇sf发布站程序 - 传奇sf发布站搭建 —— 承载源站 🌐 alien.cneagames.com 👉 传奇sf服务端网盘 - 传奇sf补丁网盘 —— 承载源站 🌐 alien.dunia-games.com 👉 传奇sf发布站模板 - 传奇sf排行榜 —— 承载源站 🌐 alien.duyougame.com 👉 传奇sf素材 - 传奇sf素材网 —— 承载源站 🌐 alien.fysagame.com 👉 传奇sf版本下载 - 传奇sf商业版本 —— 承载源站 🌐 alien.gamerping.com 👉 传奇sf网盘下载 - 传奇sf版本网盘 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/kusesde/service-hrlpt.text-generation0 likes49 downloads2mo agoHugging Face13Hrutik2003 /Bns_Law_Rag_DB BNS Law RAG Dataset Dataset: Hrutik2003/Bns_Law_Rag_DBPurpose: Text corpus for RAG systems on the new Indian criminal laws (BNS, BNSS, BSA). Dataset Summary This dataset contains cleaned and processed text extracted from the new Indian criminal laws introduced in 2023 and old IPC laws: Bharatiya Nyaya Sanhita (BNS) 2023 Bharatiya Nagarik Suraksha Sanhita (BNSS) 2023 Bharatiya Sakshya Adhiniyam (BSA) 2023 The Code of Criminal Procedure (CrPC) The Indian Penal Code (IPC)… See the full description on the dataset page: https://huggingface.co/datasets/Hrutik2003/Bns_Law_Rag_DB.text-generation2 likes44 downloads11mo agoHugging Face14allenai /href_preference HREF: Human Reference-Guided Evaluation of Instruction Following in Language Models 📑 Paper | 🤗 Leaderboard | 📁 Codebase HREF is evaluation benchmark that evaluates language models' capacity of following human instructions. This dataset contains the human agreement set of HREF, which contains 1,752 pairs of language model outputs along with the preference data from 4 human annotators for each model pairs. The dataset contains 438 instructions human-written instruction and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/href_preference.texttext-generation1K<n<10K1 likes43 downloads1y agoHugging Face15allenai /href HREF: Human Reference-Guided Evaluation of Instruction Following in Language Models 📑 Paper | 🤗 Leaderboard | 📁 Codebase HREF is evaluation benchmark that evaluates language models' capacity of following human instructions. This dataset contains the validation set of HREF, which contains 430 human-written instruction and response pairs from the test split of No Robots, covering 8 categories (removing Coding and Chat). For each instruction, we generate a baseline model… See the full description on the dataset page: https://huggingface.co/datasets/allenai/href.texttext-generationn<1K0 likes43 downloads1y agoHugging Face16pzarzycki /hrm-text-code-tools-sft HRM-Text Code & Tools SFT Curated coding and fixed-runtime tool-use supervised fine-tuning data for HRM-Text-1B. This sealed research release is not a general chat corpus. It contains only canonical v2 records—there are no legacy prefix-field files. Contents Stage Rows Train Validation Serialized cap Train SHA-256 stage-a 1,157,064 1,133,817 23,247 4,096 b6e340ea64568570b24e2d5f08c506a6f3221482d9fca6778a49c5577832171c stage-b 283,547 278,018 5,529 4… See the full description on the dataset page: https://huggingface.co/datasets/pzarzycki/hrm-text-code-tools-sft.text-generation0 likes42 downloads2mo agoHugging Face17LucioLiu /Loom-HR Loom-HR - Local-first, human-in-the-loop AI hiring assistant. The AI advises with reasons; a human makes the call. Mirrored in two places, same version everywhere: Hugging Face (you are here) · GitLab Get the files - GitLab is the most reliable plain-git route: git clone https://gitlab.com/LucioLiu/Loom-HR.git # or from this page hf download LucioLiu/Loom-HR --repo-type dataset --local-dir ./Loom-HR Licence: PolyForm Noncommercial 1.0.0 - the LICENSE file in this repo is authoritative.… See the full description on the dataset page: https://huggingface.co/datasets/LucioLiu/Loom-HR.text-generationn<1K0 likes42 downloads1mo agoHugging Face18administraktor /hrvatski-dataset-v2 Hrvatski Dataset v2 Comprehensive Croatian (hrvatski) language dataset with 455 examples, optimized for chatbot training and article generation. Every example is a question/answer (or instruction/output) pair written in natural Croatian. This repo is a single-source-of-truth catalog: all format files are generated from one canonical file (source/hrvatski_dataset_v2.jsonl), so every format is guaranteed to contain the exact same 455 examples. If a format is ever out of sync, that… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset-v2.texttext-generationn<1K0 likes39 downloads1mo agoHugging Face19YL95 /hrm-text-opus46-math-coding YL95/hrm-text-opus46-math-coding This dataset keeps only Opus 4.6 math, coding, and nearby technical reasoning tasks from the requested source datasets. Contents prompt_completion/train: the main training split for base-model fine-tuning prompt_completion/over_4096_tokens: rows longer than the token limit chat/train: a message-form version of the same kept rows chat/over_4096_tokens: the message-form over-limit subset Notes HRM-Text-1B is a base… See the full description on the dataset page: https://huggingface.co/datasets/YL95/hrm-text-opus46-math-coding.texttext-generation10K<n<100K0 likes37 downloads4mo agoHugging Face2015juneee /hr-practitioner-seed-v1 HR Practitioner Seed (v2) A curated instruction-tuning seed for HR and recruiting assistants, built for the Adaption AutoScientist Challenge (Part 2, HR track). 3,059 training rows + 244 held-out evaluation rows. What this is Group Rows Notes Grounded recruiting tasks ~2,692 real job adverts as context, 8 task types, 2 markets Generic HR policy questions 512 questions only Authored generalist HR questions 99 original, 12 practice areas All… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/hr-practitioner-seed-v1.textquestion-answering1K<n<10K0 likes30 downloads2mo agoHugging Face210xkamal7 /hr-jd-bias-audit JD-BiasAudit JD-BiasAudit is a provenance-tracked instruction-tuning dataset for HR teams, compliance reviewers, and model builders who need to neutralize coded language in job descriptions without deleting legitimate requirements. It derives structured, span-grounded audits from real postings in lang-uk/recruitment-dataset-job-descriptions-english. The upstream corpus provides job descriptions, not paired neutral rewrites, exact removed spans, protected-attribute proxy… See the full description on the dataset page: https://huggingface.co/datasets/0xkamal7/hr-jd-bias-audit.texttext-generation1K<n<10K0 likes30 downloads2mo agoHugging Face22adarsh-08 /hr-policy-dataset HR Policy Dataset Description A dataset of HR-related question-answer pairs used to fine-tune a Qwen2.5-based HR assistant. Topics Working hours Leave policies Employee benefits Onboarding Workplace conduct Company procedures Format { "instruction": "What are the official working hours?", "response": "Employees work from 9:00 AM to 6:00 PM..." } textquestion-answeringn<1K0 likes28 downloads2mo agoHugging Face23ConsumerDividends /HR-Conflict-Dataset-V2 HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored) Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license. The generator is the product This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/ConsumerDividends/HR-Conflict-Dataset-V2.tabulartext-classificationn<1K0 likes27 downloads2mo agoHugging Face2415juneee /hr-practitioner-adapted-v1 HR Practitioner (Adaption-adapted) v1 The adapted dataset used to fine-tune our HR and people operations model for the Adaption AutoScientist Challenge. Produced by running 15juneee/hr-practitioner-seed-v1 through Adaption's datasets.run. The seed carries the prompts and the curation; this carries the completions the model was actually trained on. Rows 3,059 rows. Adaption writes its output to enhanced_prompt / enhanced_completion and leaves the uploaded prompt /… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/hr-practitioner-adapted-v1.textquestion-answering1K<n<10K0 likes27 downloads2mo agoHugging Face25Hebbelille /Norwegian-Synthetic-HR-data-v-1 Synthetic norwegian public sector HR dataset Dataset description This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector. The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use. The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face26HRITHIKRAJ2537H /ICLEC Indian Supreme Court Legal Corpus – CaseBridge Edition This dataset presents a structured, passage-level corpus of Indian Supreme Court legal judgments. It is specifically curated and annotated for machine learning, information retrieval, and exploratory data analysis (EDA) tasks, and is a core resource in the CaseBridge project. Dataset Overview This corpus enables high-quality benchmarking and experimentation in: Information Retrieval: Passage and document ranking… See the full description on the dataset page: https://huggingface.co/datasets/HRITHIKRAJ2537H/ICLEC.text-classification10K<n<100K0 likes23 downloads1y agoHugging Face27huankguan2 /HRM-Text-data-io-cleaned-20260515-copyPre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/huankguan2/HRM-Text-data-io-cleaned-20260515-copy.texttext-generation100M<n<1B0 likes23 downloads3mo agoHugging Face28hreyulog /swe-bench-arktsgated SWE-bench ArkTS SWE-bench-style tasks mined from ArkTS pull requests and direct commits in HarmonyOS/OpenHarmony projects. The dataset has two evaluator-specific splits and 21 tasks in total. Split Tasks Oracle failed2pass 14 Hvigor case-level unit tests: 50 FAIL_TO_PASS, 8,653 PASS_TO_PASS, no regressions compile2pass 7 ArkTS compiler/API transitions: 5 diagnostic-removal tasks and 2 SDK/toolchain migrations Loading from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/swe-bench-arkts.tabulartext-generationn<1K3 likes23 downloads1mo agoHugging Face29Lots-of-LoRAs /task1284_hrngo_informativeness_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1284_hrngo_informativeness_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1284_hrngo_informativeness_classification.texttext-generation1K<n<10K0 likes22 downloads2y agoHugging Face30CDividends /HR-Conflict-Dataset-V2 HR Conflict Resolution Dataset - 500 (EEOC/BLS-Anchored) Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license. The generator is the product This sample was produced by our synthetic HR-conflict dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to EEOC FY2024 charge patterns and BLS wage data, with a cleaner and refiner pipeline built in. Real employee-dispute dialogue can't be… See the full description on the dataset page: https://huggingface.co/datasets/CDividends/HR-Conflict-Dataset-V2.tabulartext-classificationn<1K0 likes22 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.