datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu
Dataset Card for MMLU
Dataset Summary
Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021).
This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks… See the full description on the dataset page: https://huggingface.co/datasets/flunardelli/mmlu.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.FluxHands-FingerCount
Random Sample
Citation
@misc{FluxHandsFingerCount,
title = {FluxHands-FingerCount Dataset},
author = {Taesiri, Mohammad Reza and Ghotbizadeh, Marjan and Mirabolghasemi, Pejman},
year = {2025},
howpublished = {Hugging Face Datasets},
url = {https://huggingface.co/datasets/taesiri/FluxHands-FingerCount},
}
reasoning-1-1k
Reasoning-1 1K
Short about
This dataset will help in SFT training of LLM on the Alpaca format.
The goal of the dataset: to teach LLM to reason and analyze its mistakes using SFT training.
The size of 1.15K is quite small, so for effective training on SFTTrainer set 4-6 epochs instead of 1-3.
Made by Fluently Team (@ehristoforu) using distilabel with love🥰
Dataset structure
This subset can be loaded as:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/reasoning-1-1k.ultraset
Ultraset - all-in-one dataset for SFT training in Alpaca format
About the dataset
This dataset is designed to facilitate training and retraining of LLM models using the SFT method in the Alpaca format.
Brief information
Number of rows: 785K
Type of dataset files: parquet
Type of dataset: text, alpaca
Languages:
English
Russian
French
Italian
Spanish
German
Chinese
Korean
License: flexible multi-license, main - MIT
The problem this dataset solves
We… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultraset.fluid-knowledge-validation
Fluid Knowledge
Public synthetic release-validation fixtures. These repeated arithmetic items test artifact publication and verification only; they were not authored or blindly reviewed by frontier models and are not a usable benchmark.
Each immutable epochs/<id>/manifest.json binds its published artifacts. commitment.json reveals the nonce for verification. Protocol and source attribution accompany each epoch. Pin the returned Hugging Face commit SHA for reproduction. Public… See the full description on the dataset page: https://huggingface.co/datasets/actuallymentor/fluid-knowledge-validation.V-FLUTE
Description
Large Vision-Language models (VLMs) have demonstrated strong reasoning capabilities in tasks requiring a fine-grained understanding of literal images and text, such as visual question-answering or visual entailment. However, there has been little exploration of these models' capabilities when presented with images and captions containing figurative phenomena such as metaphors or humor, the meaning of which is often implicit. To close this gap, we propose a new task and… See the full description on the dataset page: https://huggingface.co/datasets/ColumbiaNLP/V-FLUTE.ultrathink
Ultrathink - reasoning-thinking-data dataset for SFT training in Alpaca format
About the dataset
This dataset is designed for universal SFT-training of LLM to think, reason, analyze a problem, solve a problem step by step, and break it down into subtasks.
Brief information
Number of rows: 391K
Type of dataset files: parquet
Type of dataset: text, alpaca
Language: English
License: flexible multi-license, main - MIT
The problem this dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultrathink.persian-natural-fluently
Persian scientific dataset
I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology.
The content of the dataset has been generated by : Human, Grok3, DeepSeek R1.
License
This dataset is licensed under apache-2.0.
MATH-500-Overall
MATH-500-Overall
About the dataset
This dataset of only 500 examples combines mathematics, physics and logic in English with reasoning and step-by-step problem solving, the dataset was created synthetically, CoT of Qwen2.5-72B-Instruct and Llama3.3-70B-Instruct.
Brief information
Number of rows: 500
Type of dataset files: parquet
Type of dataset: text, alpaca with system prompts
Language: English
License: MIT
Structure:
math¯¯¯¯¯⌉
school-level (100 rows)… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/MATH-500-Overall.FluxPromptingFluid-Mechanics-CoT
🌊 Engineering Fluid Mechanics CoT Dataset (工程流体力学思维链数据集)
📖 Dataset Description (数据集简介)
This dataset focuses on Engineering Fluid Mechanics, specifically designed to enhance Large Language Models' (LLMs) reasoning capabilities in complex physics problems.
Unlike standard QA datasets, this dataset provides Chain-of-Thought (CoT) annotations, breaking down the problem-solving process into:
Analysis & Reasoning: Strategy selection and physical law identification.… See the full description on the dataset page: https://huggingface.co/datasets/zshiyi/Fluid-Mechanics-CoT.code-domaine-public-fluvial-navigation-interieure
Code du domaine public fluvial et de la navigation intérieure, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-public-fluvial-navigation-interieure.meno-rag-dataset
🪷 Meno-RAG Dataset
Curated educational snippets + JSONL supervised fine-tuning pairs for a menopause guidance assistant.
⚠️ Disclaimer: Educational use only. Not medical advice. Consult a licensed clinician for personal health concerns.
📂 Contents
• snippets/ → plain-language educational notes on:
• hot_flashes.txt
• sleep_disturbance.txt
• mood_regulation.txt
• standard_test_questions.txt
• data/menopause_sft.jsonl → structured fine-tuning conversations with a 4-part… See the full description on the dataset page: https://huggingface.co/datasets/fluentnsunshine/meno-rag-dataset.
