datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.oak
NEWS:
A new version of the dataset with 120,000,000 more tokens is upload: OAK v1.1
Open Artificial Knowledge (OAK) Dataset
Overview
The Open Artificial Knowledge (OAK) dataset is a large-scale resource of over 650 Millions tokens designed to address the challenges of acquiring high-quality, diverse, and ethically sourced training data for Large Language Models (LLMs). OAK leverages an ensemble of state-of-the-art LLMs to generate high-quality text… See the full description on the dataset page: https://huggingface.co/datasets/tabularisai/oak.SVAMP_de
SVAMP_de: A Translated German Math Dataset
This is a high-quality German translation of the SVAMP dataset.
We employed a State-of-the-Art (SOTA) LLM within a strict validation pipeline to ensure 100% numerical consistency and logical fidelity.
Math word problems rely on precise numbers and logic. Standard translations often hallucinate numbers or mangle units.
SVAMP_de was generated with a "Strict Logic" pipeline:
Translation: Using SOTA LLMs.
Validation: Every single row was… See the full description on the dataset page: https://huggingface.co/datasets/tabularisai/SVAMP_de.tabular-synthetic-data
