datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-itsm-dataset
Arabic ITSM Dataset
A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release.
Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.teletype
Dataset Card for Teletype
This dataset is a scrape of all articles published on teletype, a popular platform for publishing articles, especially in Telegram. The dataset includes the original article HTML, as well as text extracted using the trafilatura library with favor_recall=True and other metadata provided by teletype.
Additionally, language identification was applied using the lingua-py library and the identification results are available in the lang column.
Curated by: its5Q
EgoSieve-Eval
EgoSieve-Eval
EgoSieve-Eval is the metadata-only, source-grouped evidence index used for
EgoSieve-S v0.1. It contains 1407 labeled window rows across
992 train, 219 validation, and
196 test examples. Source and generated videos are deliberately
not redistributed.
What the labels mean
Readiness and boundary targets are derived from HoloAssist v1_1 fine-action
intervals using a published fixed-grid occupancy rule. The test set contains
0 direct-human and 142… See the full description on the dataset page: https://huggingface.co/datasets/itspublu/EgoSieve-Eval.trm-mastering-chess
