datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
occamy-data-1.0isco_esco_occupations_taxonomy
Dataset Card for {{ pretty_name | default("Dataset Name", true) }}
{{ dataset_summary | default("", true) }}
Dataset Details
Dataset Description
{{ dataset_description | default("", true) }}
Curated by: {{ curators | default("[More Information Needed]", true)}}
Funded by [optional]: {{ funded_by | default("[More Information Needed]", true)}}
Shared by [optional]: {{ shared_by | default("[More Information Needed]", true)}}
Language(s) (NLP): {{ language |… See the full description on the dataset page: https://huggingface.co/datasets/ICILS/isco_esco_occupations_taxonomy.OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
Dataset Description
OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation.
Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.ud-campus-parking-occupancy-synthetic
UD Parking Occupancy Classifier (Tiny Demo)
Very small scikit-learn RandomForest classifier trained on the synthetic dataset:
BuildingTHEITGUY/ud-campus-parking-occupancy-synthetic
This is a teaching / portfolio model, not a production campus system.
What it predicts
Class label: open | busy | full
Features
capacity, occupied, free, occupancy_ratio, hour_local, weekday
Files
parking_occupancy_rf.joblib — model artifact
metrics.json —… See the full description on the dataset page: https://huggingface.co/datasets/BuildingTHEITGUY/ud-campus-parking-occupancy-synthetic.medical-certificate-validity-by-occupation
How long medical certificates last for pilots, commercial drivers and merchant mariners — by class, age and service type
Canonical, always-current version: https://referencesource.org/medical-certificate-validity-by-occupation/
Machine-readable: https://referencesource.org/medical-certificate-validity-by-occupation/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2027-08-11 (past this date, prefer the canonical copy —
it re-verifies on a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/medical-certificate-validity-by-occupation.occultexperttaktkrone-occ-corpus
TAKTKRONE OCC Dialogue Corpus
Dataset Summary
The TAKTKRONE OCC Dialogue Corpus is a specialized dataset for training language models to assist in metro operations control center (OCC) scenarios. It contains realistic dialogue samples between operators and control center staff during various transit incidents and operational situations.
Dataset Details
Created by: Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov
Language: English
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/olaflaitinen/taktkrone-occ-corpus.high_priest_occult_50k
High Priest Mindset Training Dataset
Dataset Description
This dataset contains 25,000 high-quality training examples designed to train language models in the mindset, reasoning patterns, and operational framework of a high priest across various spiritual and occult traditions.
Dataset Summary
The dataset draws from ancient texts including:
The Sixth and Seventh Books of Moses
The Egyptian Book of the Dead
The Corpus Hermeticum
The Key of Solomon… See the full description on the dataset page: https://huggingface.co/datasets/11-47/high_priest_occult_50k.glm-base-ood-repair-mix-10k
GLM base OOD repair mix 10k
Bucket-targeted BFCL-style tool-calling repair dataset for GLM native tool-call finetuning.
Built from public OOD tool-call datasets and filtered against BFCL single-call eval prompts.
Primary file: train.jsonl
Rows: 9788 after dropping exact BFCL eval prompt overlaps.
Format: messages, tools, target_call. Training should use GLM native target formatting from target_call, not the legacy target_text_cohere field.
Audit files included:… See the full description on the dataset page: https://huggingface.co/datasets/Occupying-Mars/glm-base-ood-repair-mix-10k.MA_Occupational_Safety_Reports
Massachusetts Occupational Safety and Health Statistics Dataset
Dataset Description
This dataset contains workplace safety information extracted from the Massachusetts Occupational Safety and Health Statistics Program between 2017 and 2022, including injuries by industry, occupation, and demographic data. It provides structured, machine-readable data converted from PDF reports that offer insights into workplace safety trends across Massachusetts.
Overview
The… See the full description on the dataset page: https://huggingface.co/datasets/evijit/MA_Occupational_Safety_Reports.Occupational_Classification_Code_of_PRC_2022occiglot__occiglot-7b-es-en-instruct-details
Dataset Card for Evaluation run of occiglot/occiglot-7b-es-en-instruct
Dataset automatically created during the evaluation run of model occiglot/occiglot-7b-es-en-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/occiglot__occiglot-7b-es-en-instruct-details.
