datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-auxiliary-train-auto-labelled
Dataset Card for MMLU Auxiliary Trained Set Labelled by e5-mistral-7b-instruct
Dataset Description
Dataset Summary
This dataset, named "MMLU Auxiliary Trained Set Labelled by e5-mistral-7b-instruct," consists of 99,842 examples spanning various subjects. Each instance includes a question, multiple choice options, a subject category, and an answer. The unique aspect of this dataset is the task label for each question, generated by a zero-shot classifier… See the full description on the dataset page: https://huggingface.co/datasets/kz919/mmlu-auxiliary-train-auto-labelled.labelled_regex
Labelled Regex
This dataset consists of Regexes and their descriptive labels. As far as I am aware, this is the largest, cleanly labelled regex dataset on this platform.
I constructed this dataset by taking innovatorved/regex_dataset and using gemma-3-27b-it LLM to generate a concise and suitable title for each regex.
For each regex that was larger than 100 characters, I used a slightly different prompt to generate an even more detailed description.
