datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
Dataset Description
OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation.
Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.taktkrone-occ-corpus
TAKTKRONE OCC Dialogue Corpus
Dataset Summary
The TAKTKRONE OCC Dialogue Corpus is a specialized dataset for training language models to assist in metro operations control center (OCC) scenarios. It contains realistic dialogue samples between operators and control center staff during various transit incidents and operational situations.
Dataset Details
Created by: Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov
Language: English
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/olaflaitinen/taktkrone-occ-corpus.glm-base-ood-repair-mix-10k
GLM base OOD repair mix 10k
Bucket-targeted BFCL-style tool-calling repair dataset for GLM native tool-call finetuning.
Built from public OOD tool-call datasets and filtered against BFCL single-call eval prompts.
Primary file: train.jsonl
Rows: 9788 after dropping exact BFCL eval prompt overlaps.
Format: messages, tools, target_call. Training should use GLM native target formatting from target_call, not the legacy target_text_cohere field.
Audit files included:… See the full description on the dataset page: https://huggingface.co/datasets/Occupying-Mars/glm-base-ood-repair-mix-10k.occiglot-fineweb-v0.5
Occiglot Fineweb v0.5
We present a preliminary version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 230M heavily cleaned documents from 10 languages.
Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data.
Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and globally depuplicated.
We are actively working on extending this… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v0.5.
