CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oumi-ai /MetaMathQA-R1 oumi-ai/MetaMathQA-R1 MetaMathQA-R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning. Prompts were augmented from GSM8K and MATH training sets with responses directly from DeepSeek-R1. MetaMathQA-R1 was used to train MiniMath-R1-1.5B, which achieves 44.4% accuracy on MMLU-Pro-Math, the highest of any model with <=1.5B parameters. Curated by: Oumi AI using Oumi inference on Parasail Language(s) (NLP): English License:… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/MetaMathQA-R1.texttext-generation100K<n<1M7 likes663 downloads2y agoHugging Face02oumi-ai /lmsys_chat_1m_clean_R1 oumi-ai/lmsys_chat_1m_clean_R1 lmsys_chat_1m_clean_R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning. Prompts were pulled from LMSYS and filtered to lmsys_chat_1m_clean, and responses were taken from DeepSeek-R1 without additional filters present. We release lmsys_chat_1m_clean_R1 to help enable the community to develop the best fully open reasoning model! lmsys_chat_1m_clean queries with responses generated from… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/lmsys_chat_1m_clean_R1.texttext-generation100K<n<1M9 likes225 downloads2y agoHugging Face03oumi-ai /oumi-letter-counttext100K<n<1M0 likes68 downloads2y agoHugging Face04oumi-ai /banking77-oumi-quickstarttext10K<n<100K0 likes54 downloads6mo agoHugging Face05oumi-ai /MM-MathInstruct-to-r1-format-filtered MM-MathInstruct-to-r1-format-filtered MM-MathInstruct dataset transformed to R1 format and filtered by token length and image quality Dataset Description This dataset was processed using the data-preproc package for vision-language model training. Processing Configuration Base Model: Qwen/Qwen2.5-7B-Instruct Tokenizer: Qwen/Qwen2.5-7B-Instruct Sequence Length: 16384 Processing Type: Vision Language (VL) Dataset Features input_ids: Tokenized input… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/MM-MathInstruct-to-r1-format-filtered.text10K<n<100K0 likes53 downloads1y agoHugging Face06PJMixers-Dev /oumi-ai_lmsys_chat_1m_clean_R1-1k-think-1k-response-ShareGPTtabular100K<n<1M0 likes44 downloads2y agoHugging Face07Teen-Different /grpo-oumi-synthetic-document-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.texttext-generation1K<n<10K0 likes26 downloads1y agoHugging Face08oumi-ai /walton-multimodal-cold-start-r1-format walton-multimodal-cold-start-r1-format WaltonFuture/Multimodal-Cold-Start converted to multimodal-open-r1-8k-verified format with filtering Dataset Description This dataset was processed using the data-preproc package for vision-language model training. Processing Configuration Base Model: Qwen/Qwen2.5-7B-Instruct Tokenizer: Qwen/Qwen2.5-7B-Instruct Sequence Length: 16384 Processing Type: Vision Language (VL) Dataset Features input_ids: Tokenized… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/walton-multimodal-cold-start-r1-format.image10K<n<100K1 likes26 downloads1y agoHugging Face09oumi-ai /multimodal-open-r1-8192-filtered-mid-ic multimodal-open-r1-8192-filtered-mid-ic Original dataset structure preserved, filtered by token length and image quality Dataset Description This dataset was processed using the data-preproc package for vision-language model training. Processing Configuration Base Model: Qwen/Qwen2.5-7B-Instruct Tokenizer: Qwen/Qwen2.5-7B-Instruct Sequence Length: 16384 Processing Type: Vision Language (VL) Dataset Features input_ids: Tokenized input sequences… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/multimodal-open-r1-8192-filtered-mid-ic.image1K<n<10K0 likes24 downloads1y agoHugging Face10shanghong /oumi_rag_grpo_datatext1K<n<10K0 likes24 downloads1y agoHugging Face11oumi-ai /examplestextn<1K0 likes24 downloads5mo agoHugging Face12oumi-ai /oumi-synthetic-document-claims oumi-ai/oumi-synthetic-document-claims oumi-synthetic-document-claims is a text dataset designed to fine-tune language models for Claim Verification. Prompts and responses were produced synthetically from Llama-3.1-405B-Instruct. oumi-synthetic-document-claims was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc. Curated by: Oumi AI using Oumi inference Language(s) (NLP): English License: Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims.text1K<n<10K6 likes23 downloads1y agoHugging Face13oumi-ai /berrybench-v0.1.1text10K<n<100K0 likes23 downloads2y agoHugging Face14oumi-ai /s1-vis-mid-resize s1-vis-mid-resize Original dataset structure preserved, filtered by token length and image quality Dataset Description This dataset was processed using the data-preproc package for vision-language model training. Processing Configuration Base Model: Qwen/Qwen2.5-7B-Instruct Tokenizer: Qwen/Qwen2.5-7B-Instruct Sequence Length: 16384 Processing Type: Vision Language (VL) Dataset Features input_ids: Tokenized input sequences attention_mask: Attention… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/s1-vis-mid-resize.imagen<1K0 likes23 downloads1y agoHugging Face15oumi-ai /oumi-groundedness-benchmark oumi-ai/oumi-groundedness-benchmark oumi-groundedness-benchmark is a text dataset designed to evaluate language models for Claim Verification / Hallucination Detection. Prompts and responses were produced synthetically from Llama-3.1-405B-Instruct. oumi-groundedness-benchmark was used to properly evaluate HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc. Curated by: Oumi AI using Oumi inference Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-groundedness-benchmark.text1K<n<10K6 likes22 downloads1y agoHugging Face16Dicksonycx /oumi_math_1k_3ktext10K<n<100K0 likes22 downloads4mo agoHugging Face17oumi-ai /oumi-synthetic-claims oumi-ai/oumi-synthetic-claims oumi-synthetic-claims is a text dataset designed to fine-tune language models for Claim Verification. Prompts and responses were produced synthetically from Llama-3.1-405B-Instruct. oumi-synthetic-claims was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc. Curated by: Oumi AI using Oumi inference Language(s) (NLP): English License: Llama 3.1 Community License… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims.text10K<n<100K6 likes21 downloads1y agoHugging Face18Teen-Different /grpo-oumi-c2d-d2c-subset Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-c2d-d2c-subset dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-c2d-d2c-subset Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a single data instance with… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-c2d-d2c-subset.texttext-generation10K<n<100K0 likes21 downloads1y agoHugging Face19oumi-ai /oumi-c2d-d2c-subset oumi-ai/oumi-c2d-d2c-subset oumi-c2d-d2c-subset is a text dataset designed to fine-tune language models for Claim Verification. Prompts were pulled from C2D-and-D2C-MiniCheck training sets with responses created from Llama-3.1-405B-Instruct. oumi-c2d-d2c-subset was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc. Curated by: Oumi AI using Oumi inference Language(s) (NLP): English License: Llama… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-c2d-d2c-subset.text10K<n<100K4 likes18 downloads1y agoHugging Face20oumi-ai /oumi-letter-count-cleantext100K<n<1M0 likes17 downloads1y agoHugging Face21yosubshin /oumi-walton-exclude-geometry-biologyimage1K<n<10K0 likes17 downloads10mo agoHugging Face22Teen-Different /grpo-oumi-anli-subset Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-anli-subset dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-anli-subset Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a single data instance with a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-anli-subset.texttext-generation10K<n<100K0 likes15 downloads1y agoHugging Face23Teen-Different /grpo-oumi-synthetic-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the TEEN-D/grpo-oumi-anli-subset dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a single data instance… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-claims.texttext-generation10K<n<100K0 likes15 downloads1y agoHugging Face24yosubshin /oumi-walton-exclude-geometry-biology-statisticsimage1K<n<10K1 likes14 downloads10mo agoHugging Face25yosubshin /oumi-walton-exclude-geometry-biology-statistics-none-of-aboveimage1K<n<10K0 likes13 downloads10mo agoHugging Face26Dicksonycx /oumi_gsmtext100K<n<1M0 likes13 downloads5mo agoHugging Face27Dicksonycx /oumi_math_hardtext10K<n<100K0 likes12 downloads5mo agoHugging Face28shanghong /oumi-web-agentimage1K<n<10K0 likes11 downloads1y agoHugging Face29oumi-ai /oumi-anli-subset oumi-ai/oumi-anli-subset oumi-anli-subset is a text dataset designed to fine-tune language models for Claim Verification. Prompts were pulled from ANLI training sets with responses created from Llama-3.1-405B-Instruct. oumi-anli-subset was used to train HallOumi-8B, which achieves 77.2% Macro F1, outperforming SOTA models such as Claude Sonnet 3.5, OpenAI o1, etc. Curated by: Oumi AI using Oumi inference Language(s) (NLP): English License: CC-BY-NC-4.0, Llama 3.1 Community… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/oumi-anli-subset.text10K<n<100K5 likes8 downloads1y agoHugging Face30oumi-ai /s1.1-VLimage1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.