datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MetaMathQA-R1
oumi-ai/MetaMathQA-R1
MetaMathQA-R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning.
Prompts were augmented from GSM8K and MATH training sets with responses directly from DeepSeek-R1.
MetaMathQA-R1 was used to train MiniMath-R1-1.5B, which achieves 44.4% accuracy on MMLU-Pro-Math, the highest of any model with <=1.5B parameters.
Curated by: Oumi AI using Oumi inference on Parasail
Language(s) (NLP): English
License:… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/MetaMathQA-R1.lmsys_chat_1m_clean_R1
oumi-ai/lmsys_chat_1m_clean_R1
lmsys_chat_1m_clean_R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning.
Prompts were pulled from LMSYS and filtered to lmsys_chat_1m_clean, and responses were taken from DeepSeek-R1 without additional filters present.
We release lmsys_chat_1m_clean_R1 to help enable the community to develop the best fully open reasoning model!
lmsys_chat_1m_clean queries with responses generated from… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/lmsys_chat_1m_clean_R1.oumi-letter-countbanking77-oumi-quickstartMM-MathInstruct-to-r1-format-filtered
MM-MathInstruct-to-r1-format-filtered
MM-MathInstruct dataset transformed to R1 format and filtered by token length and image quality
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized input… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/MM-MathInstruct-to-r1-format-filtered.oumi-ai_lmsys_chat_1m_clean_R1-1k-think-1k-response-ShareGPTgrpo-oumi-synthetic-document-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.walton-multimodal-cold-start-r1-format
walton-multimodal-cold-start-r1-format
WaltonFuture/Multimodal-Cold-Start converted to multimodal-open-r1-8k-verified format with filtering
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/walton-multimodal-cold-start-r1-format.multimodal-open-r1-8192-filtered-mid-ic
multimodal-open-r1-8192-filtered-mid-ic
Original dataset structure preserved, filtered by token length and image quality
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized input sequences… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/multimodal-open-r1-8192-filtered-mid-ic.oumi_rag_grpo_dataexamplesoumi-synthetic-document-claims
oumi-ai/oumi-synthetic-document-claims
oumi-synthetic-document-claims is a text dataset designed to fine-tune language models for Claim Verification.
Prompts and responses were produced synthetically from Llama-3.1-405B-Instruct.
oumi-synthetic-document-claims was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc.
Curated by: Oumi AI using Oumi inference
Language(s) (NLP): English
License: Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims.berrybench-v0.1.1s1-vis-mid-resize
s1-vis-mid-resize
Original dataset structure preserved, filtered by token length and image quality
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized input sequences
attention_mask: Attention… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/s1-vis-mid-resize.oumi-groundedness-benchmark
oumi-ai/oumi-groundedness-benchmark
oumi-groundedness-benchmark is a text dataset designed to evaluate language models for Claim Verification / Hallucination Detection.
Prompts and responses were produced synthetically from Llama-3.1-405B-Instruct.
oumi-groundedness-benchmark was used to properly evaluate HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc.
Curated by: Oumi AI using Oumi inference
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-groundedness-benchmark.oumi_math_1k_3koumi-synthetic-claims
oumi-ai/oumi-synthetic-claims
oumi-synthetic-claims is a text dataset designed to fine-tune language models for Claim Verification.
Prompts and responses were produced synthetically from Llama-3.1-405B-Instruct.
oumi-synthetic-claims was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc.
Curated by: Oumi AI using Oumi inference
Language(s) (NLP): English
License: Llama 3.1 Community License… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims.grpo-oumi-c2d-d2c-subset
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the oumi-ai/oumi-c2d-d2c-subset dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-c2d-d2c-subset
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a single data instance with… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-c2d-d2c-subset.oumi-c2d-d2c-subset
oumi-ai/oumi-c2d-d2c-subset
oumi-c2d-d2c-subset is a text dataset designed to fine-tune language models for Claim Verification.
Prompts were pulled from C2D-and-D2C-MiniCheck training sets with responses created from Llama-3.1-405B-Instruct.
oumi-c2d-d2c-subset was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc.
Curated by: Oumi AI using Oumi inference
Language(s) (NLP): English
License: Llama… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-c2d-d2c-subset.oumi-letter-count-cleanoumi-walton-exclude-geometry-biologygrpo-oumi-anli-subset
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the oumi-ai/oumi-anli-subset dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-anli-subset
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a single data instance with a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-anli-subset.grpo-oumi-synthetic-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the TEEN-D/grpo-oumi-anli-subset dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a single data instance… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-claims.oumi-walton-exclude-geometry-biology-statisticsoumi-walton-exclude-geometry-biology-statistics-none-of-aboveoumi_gsmoumi_math_hardoumi-web-agentoumi-anli-subset
oumi-ai/oumi-anli-subset
oumi-anli-subset is a text dataset designed to fine-tune language models for Claim Verification.
Prompts were pulled from ANLI training sets with responses created from Llama-3.1-405B-Instruct.
oumi-anli-subset was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc.
Curated by: Oumi AI using Oumi inference
Language(s) (NLP): English
License: CC-BY-NC-4.0, Llama 3.1 Community… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-anli-subset.s1.1-VL
