datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-auxiliary-train-auto-labelled
Dataset Card for MMLU Auxiliary Trained Set Labelled by e5-mistral-7b-instruct
Dataset Description
Dataset Summary
This dataset, named "MMLU Auxiliary Trained Set Labelled by e5-mistral-7b-instruct," consists of 99,842 examples spanning various subjects. Each instance includes a question, multiple choice options, a subject category, and an answer. The unique aspect of this dataset is the task label for each question, generated by a zero-shot classifier… See the full description on the dataset page: https://huggingface.co/datasets/kz919/mmlu-auxiliary-train-auto-labelled.jobseek-postings-labelled
jobseek-postings-labelled
Gold-standard labelled job postings sampled daily from public company
career pages. Produced by a Codex-first agent pipeline with task-specific
subagents for HTML normalization, section splitting, and structured
extraction. The dataset is the substrate for training an improved
structured-information extractor for jseek.co.
Current row counts by date: 2026-09-22: 10 · 2026-09-21: 10 · 2026-09-20: 10 · 2026-09-14: 10 · 2026-09-13: 10 · 2026-08-25: 10 ·… See the full description on the dataset page: https://huggingface.co/datasets/viktor-shcherb/jobseek-postings-labelled.labelled_regex
Labelled Regex
This dataset consists of Regexes and their descriptive labels. As far as I am aware, this is the largest, cleanly labelled regex dataset on this platform.
I constructed this dataset by taking innovatorved/regex_dataset and using gemma-3-27b-it LLM to generate a concise and suitable title for each regex.
For each regex that was larger than 100 characters, I used a slightly different prompt to generate an even more detailed description.
