datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEP-3MMephisto-Knowledge_538k
Mephisto-Knowledge_538k
538,861 English knowledge SFT examples generated by
Qwen/Qwen3.5-4B in non-thinking
(Instruct) mode on the Knowledge prompts of
openbmb/UltraData-SFT-2605.
Responses contain no chain-of-thought — thinking was disabled at generation
time, so each assistant turn is a direct answer, usually with a short
justification.
Companion dataset: Mephisto-IF_172k
(instruction-following, same teacher and pipeline).
Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.mep-plant-room-pipe-routing
MEP Plant Room Pipe Routing Dataset
This repository provides case-level data for the paper:
Automated Multi-pipe Routing in MEP Plant Rooms: A Preference-driven Multi-objective Coordination Framework for Design Prototyping
DOI: 10.1016/j.autcon.2026.107201
Link: https://doi.org/10.1016/j.autcon.2026.107201
Dataset Contents
Each case folder contains:
input_case.json
Scene input, including room boundary, maintenance areas, obstacles, pipe terminal points, pipe… See the full description on the dataset page: https://huggingface.co/datasets/IHNF/mep-plant-room-pipe-routing.Mephisto-MathCode_2M
Mephisto-MathCode_2M
2,000,000 non-thinking SFT examples — an even 1M/1M split of math and code
— drawn from
openbmb/UltraData-SFT-2605,
filtered to English and globally shuffled.
This is a curation pass, not a generation one: no model produced these
answers for this dataset. All credit for the content belongs to OpenBMB. What
is added here is language filtering, an exact 1M/1M balance, and a global
shuffle so the file can be streamed without a shuffle buffer.
domain
rows… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-MathCode_2M.Mephisto-IF_172k
Mephisto-IF_172k
172,761 English instruction-following SFT examples, generated by
Qwen/Qwen3.5-4B in non-thinking
(Instruct) mode on the instruction-following prompts of
openbmb/UltraData-SFT-2605.
Responses contain no chain-of-thought — thinking was disabled at generation
time, so every assistant turn is a direct answer.
Format
One JSON object per line:
{
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-IF_172k.MEPC
Multi-level Product Category Recognition Image Dataset
Summary
Wordcloud
Introduce
MEPC - 1000 Dataset:
Classes: 1000
Images: 164,117
Train: 131,293
Val: 32824
MEPC - 10 Dataset:
Classes: 10
Images: 2,192
Train: 1,753
Val: 439
Statistics
Statistics of the number of multi-level categories in the two datasets MEPC-10 and MEPC-1000
Label-only embeddings visualizing label connections… See the full description on the dataset page: https://huggingface.co/datasets/sherlockvn/MEPC.MePO
MePO Prompt Optimization Dataset
This dataset is designed for research in prompt optimization, particularly for training and evaluating MePO — a lightweight, locally deployable prompt optimization model.
📂 File: MePO.jsonl (40,151 entries)
Each JSONL record includes:
rejectedThe original prompt from BPO or Alpaca, used as the rejected example.
chosenThe optimized prompt generated by MePO, used as the chosen example.
sliver_responseThe response produced from the… See the full description on the dataset page: https://huggingface.co/datasets/zixiaozhu/MePO.Diabetic-trackerMePO_BPO
MePO Prompt Optimization Dataset (BPO version)
This dataset is designed for research in prompt optimization, particularly for training and evaluating MePO — a lightweight, locally deployable prompt optimization model.
Each JSONL record includes:
rejectedThe original prompt from BP, used as the rejected example.
chosenThe optimized prompt generated by MePO, used as the chosen example.
sliver_responseThe response produced from the BPO prompt (baseline response).
golden_responseThe… See the full description on the dataset page: https://huggingface.co/datasets/zixiaozhu/MePO_BPO.Aksharantar
Dataset Card for Aksharantar
Dataset Summary
Aksharantar is the largest publicly available transliteration dataset for 20 Indic languages. The corpus has 26M Indic language-English transliteration pairs.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Assamese (asm)
Hindi (hin)
Maithili (mai)
Marathi (mar)
Punjabi (pan)
Tamil (tam)
Bengali (ben)
Kannada (kan)
Malayalam (mal)
Nepali (nep)
Sanskrit (san)
Telugu… See the full description on the dataset page: https://huggingface.co/datasets/mephistosir329/Aksharantar.Public_Call_Transcripts_Data_KaggleMePO_Alpaca
📦 MePO Prompt Optimization Dataset (Alpaca Version)
The MePO Prompt Optimization Dataset is designed to support research in prompt optimization, especially for training and evaluating MePO — a lightweight and locally deployable prompt optimization model.
📁 Dataset Structure
Each .jsonl record contains the following fields:
rejectedThe original prompt from the Alpaca dataset, serving as the rejected example.
chosenThe optimized prompt generated by MePO, serving as the… See the full description on the dataset page: https://huggingface.co/datasets/zixiaozhu/MePO_Alpaca.mePics;oertjh
hvac-mep-services-pakistanAgentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/mepartha/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.outlineopenscad-vision-sftmeps_speeches_with_translation.csvMep3mmeps_speechesThis dataset contains nearly 18,000 European Member of Parliament (meps) speeches beween 2019 and 2023.
The speeches are from Italian, German, French and Belgium meps.
All the speeches were gently scraped for the european parliament website using this code: https://github.com/misclassified/meps-text-mining
list_mepMinerU2.5-2509-1.2B_MEP_pythonmisclassified-meps_speechesCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/misclassified/meps_speeches.
mep3m-crossvit-debertamepepperaisensor-pod-historyMePhimVietTHVLit_mepsMEPs27N_lhjBkQGh64mepetss
