datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.tatoeba_mt
Dataset Card for [Dataset Name]
Dataset Summary
The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development data sorted by language pair. It includes test sets for hundreds of language pairs and is continuously updated. Please, check the version number tag to refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt.MT-Nemotron-CC
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT Nemotron-CC subset of MultiSynt is made of automatic translations into multiple languages from a subset of approximately 100B tokens from the high-quality split of the English Nemotron-CC dataset.
This subset is made available using different translation models:
Unbabel/Tower-Plus-9B (translations into 16 languages)
Unbabel/Tower-Plus-72B (translations into 5 languages)
Opus-MT and HPLT-MT (translations… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Nemotron-CC.apex-r1-real-world-documents
Apex-R1 Real-World Benchmark Documents
This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation.
The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks.
Contents
benchmark_documents/
EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.Magpie-Llama-3.1-Pro-MT-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.Funcdex-MT-Function-Calling
Funcdex-MT-Function-Calling Dataset
Funcdex-MT-Function-Calling is a multi-turn function calling dataset designed for training language models to interact with real-world tools and APIs. The dataset contains 1,787 conversations covering 10 individual toolkits and 5 multi-toolkit bundles, with comprehensive system prompts and realistic multi-turn interactions.The code used to generate the dataset can be found here.
Models trained on this dataset have excellent… See the full description on the dataset page: https://huggingface.co/datasets/prem-research/Funcdex-MT-Function-Calling.MT-Reasoning
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses.
lang
rows
prompt_tokens
reasoning_tokens
response_tokens
total_tokens
deu_Latn
17_354_716
1_873_153_732
26_010_932_738
14_862_651_336
42_746_737_806
fra_Latn
17_354_716
1_802_885_115
25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.Semantic-Flow-Dynamics-SFD
Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus
TL;DR: 614 Chinese-language formalized social-science concepts across 25
papers, UUID-linked with typed derivation relations (derives_from,
leads_to, falsified_by, …) — usable for knowledge-graph construction,
RAG over structured theory, or as a Chinese formal-reasoning corpus.
Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com
What This Dataset Is
This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.us-caselaw-mt
Montana Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of
2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-mt.tatoeba_mt_parquet
Dataset Card for DigitalLearningGmbH/tatoeba_mt_parquet
This is a mirror of Helsinki-NLP/tatoeba_mt, converted to parquet for compatibility with newer huggingface requirements.
Original dataset card follows.
Dataset Summary
The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/tatoeba_mt_parquet.midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full
piece (no time-windowing); training crops sequences from packed token bins.
The source column is the original piece metadata as JSON so a row can be
traced back to its EPR Labs source dataset.
Based on MIDI datasets gathered by EPR Labs.
Codec
name: dyadic
tokenizer vocab size: 512
max_time_step: 1.0
n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.Neo-GATE
Dataset card for Neo-GATE
Homepage: https://mt.fbk.eu/neo-gate/
Dataset summary
Neo-GATE is a bilingual corpus designed to benchmark the ability of machine translation (MT) systems to translate from English into Italian using gender-inclusive neomorphemes.
It is built upon GATE (Rarrick et al., 2023), a benchmark for the evaluation of gender rewriters and gender bias in MT.
Neo-GATE includes 841 test entries (Neo-GATE.tsv) and 100 dev entries (Neo-GATE-dev.tsv).
Each… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Neo-GATE.finbenchv2-opengpt-x_truthfulqax-fi-mtThis is an archived version of LumiOpen/opengpt-x_truthfulqax used in Finbench version 2, as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models.
Code: https://github.com/LumiOpen/lm-evaluation-harness
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-opengpt-x_truthfulqax-fi-mt.hplt-greek-ge8-no-mt-clean60-wave4
HPLT Greek GE8 No-MT Clean60 Wave4
A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass.
Snapshot
Rows: 48728774
Data parquet files: 250
Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60
Quality bins: 8, 9, 10
MT/register filtering: applied before this release
Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.fineweb-dutch-edu-mt
FineWeb-Edu Dutch Machine Translated
Machine-translated Dutch text dataset derived from the FineWeb-Edu corpus.
Dataset Details
Source: HuggingFaceFW/fineweb-edu (sample-10BT subset)
Translation: English → Dutch using Unbabel/Tower-Plus-9B
Size: Up to 1.5M samples
Format: Translated text with original metadata
Schema
text: Machine-translated Dutch text
id: Original sample identifier from FineWeb-Edu
url: Source URL
Quality Notice
⚠️ This… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/fineweb-dutch-edu-mt.Electrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
MTAC-IFBench
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
🌟 Overview
MTAC-IFBench benchmarks instruction following in multi-turn agentic coding.
Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.bigbenchhard-mt-pt
BBH-PT (Big-Bench Hard)
Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks.
Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations.
Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese.
Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard
Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.MTS-All
MTS-All
MTS-All is the data release for the EMNLP 2026 accepted paper
Reactivating Test-Time Scaling for Plane Geometry Problem Solving
(PDF).
Training and evaluation code is available in the
ReTTS-PGPS repository.
The dataset contains multi-trace supervised fine-tuning data and test files for three plane geometry benchmarks:
PGPS9K-All
Geometry3K-All
GeoQA-All
Each training problem is represented with four reasoning traces:
Program: symbolic geometry program.
COT-program:… See the full description on the dataset page: https://huggingface.co/datasets/kxiaoqiangrexian/MTS-All.finbenchv2-squad-strip-fi-mt
finbenchv2-squad-strip-fi-mt
This dataset is a subset of our SQuAD v2 HF dataset with
unanswerable questions removed, to be used within the FIN-bench-v2 benchmark suite. An additional feature of this
dataset is that the text in the title fields have been machine-translated to Finnish.
Paper: https://huggingface.co/papers/2512.13330
Code: https://github.com/LumiOpen/lm-evaluation-harness
Considerations for Using the Data
Due to DeepL terms and conditions, this… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-squad-strip-fi-mt.apigen-mt-5k-parsed
[PARSED] APIGen-MT-5k
The data in this dataset is a full of the original Salesforce/APIGen-MT-5k
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
apigen-mt-5k
yes
no
yes
complex
5k
This is a re-parsing formatting dataset for the APIGen-MT-5k official dataset.
Load the dataset
from datasets import load_dataset
ds = load_dataset("minpeter/apigen-mt-5k-parsed")
print(ds)
# DatasetDict({
# train: Dataset({
#… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/apigen-mt-5k-parsed.thomas-yanxin-MT-SFT-ShareGPT
thomas-yanxin/MT-SFT-ShareGPT
This is the complete thomas-yanxin/MT-SFT-ShareGPT dataset,
with duplicates removed and the entire dataset shuffled. Sensitive data has been redacted.
For practical work, consider using agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample
which is smaller and split by language.
Magpie-Air-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Air-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Air-MT-300K-v0.1-ko.thomas-yanxin-MT-SFT-ShareGPT-sample
MT-SFT-ShareGPT Sample Dataset
This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets.
Dataset Contents
train.jsonl: Contains 1/10 of the original data, shuffled
EN.jsonl: English conversations from train.jsonl
ZH.jsonl: Chinese conversations from train.jsonl
Each row represents a conversation with an optional system message, followed by human and GPT turns.
Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.ifeval_mt
IFEval Multilingual
These are machine-translated versions of Instruction Following Evaluation (IFEval). We will do our best to correct the translations. Translations were done using DeepL and the translations were reviewed and corrected by native speakers. We use this dataset in our fork of LM Eval Harness that supports multilingual ifeval.
Supported languages
Finnish: machine-translated manually corrected
Swedish: machine-translated but not corrected
mtob
MTOB (Machine Translation from One Book)
Last updated: Wednesday, July 9, 2025
Machine Translation from One Book evaluates a language model's ability to translate sentences from English to Kalamang (a low-resource language) and from Kalamang to English.
As of July 2, 2025, additional tasks for this groq-bench implementation include:
Kalamang-to-English translation
adding the option to perform long-context evaluation where the Kalamang corpus is used as input to the model
adding… See the full description on the dataset page: https://huggingface.co/datasets/Groq/mtob.ea-mt-benchmark
Dataset Card for EA-MT
EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate.
Here is an example of a simple sentence with a challenging entity mention:
English: "What is the plot of The Catcher in the Rye?"
Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.MultiTurn-Chat-MT-Bench
SEA-MTBench
SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use gpt-4-1106-preview as the judge model and compare against gpt-3.5-turbo-0125 as the baseline model. It is based on MT-Bench and was manually translated by native speakers for Indonesian (id), Javanese (jv), Sundanese (su), and Vietnamese (vi). The Thai split of this dataset uses MT-Bench Thai from the ThaiLLM leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench.MT-HPLT2c
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT-HPLT2c subset of MultiSynt is a large-scale, machine translated variant of the HPLT v2 English dataset to study LLM training on translated data.
From the English source, we offer translations for the following 4 target languages:
deu_Latn, fin_Latn, spa_Latn, swe_Latn.
For each language, we provide 3 splits:
all: The entire data.
parallel: A subset of 115,082,738 aligned documents, such that a document… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-HPLT2c.Magpie-Pro-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Pro-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Pro-MT-300K-v0.1-ko.
