datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLU_ExpertPrompt_RAGThis dataset contains a copy of the cais/mmlu HF dataset but without the auxiliary_train split that takes a long time to generate again each time when loading multiple subsets of the dataset.
Please visit https://huggingface.co/datasets/cais/mmlu for more information on the MMLU dataset.
markdown-table-expert
Markdown Table Expert
A large-scale dataset for teaching language models to read, understand, and reason over markdown tables. Contains 44,000 samples (40,000 train + 4,000 validation) spanning 35 real-world domains with detailed step-by-step reasoning traces.
Why This Dataset
Markdown tables are everywhere — in documentation, reports, READMEs, financial statements, and web content. Yet most LLMs struggle with structured tabular data, especially when asked to perform… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-expert.
