datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magpie-Llama-3.1-Pro-MT-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.MT-Reasoning
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses.
lang
rows
prompt_tokens
reasoning_tokens
response_tokens
total_tokens
deu_Latn
17_354_716
1_873_153_732
26_010_932_738
14_862_651_336
42_746_737_806
fra_Latn
17_354_716
1_802_885_115
25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full
piece (no time-windowing); training crops sequences from packed token bins.
The source column is the original piece metadata as JSON so a row can be
traced back to its EPR Labs source dataset.
Based on MIDI datasets gathered by EPR Labs.
Codec
name: dyadic
tokenizer vocab size: 512
max_time_step: 1.0
n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.hplt-greek-ge8-no-mt-clean60-wave4
HPLT Greek GE8 No-MT Clean60 Wave4
A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass.
Snapshot
Rows: 48728774
Data parquet files: 250
Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60
Quality bins: 8, 9, 10
MT/register filtering: applied before this release
Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.thomas-yanxin-MT-SFT-ShareGPT-sample
MT-SFT-ShareGPT Sample Dataset
This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets.
Dataset Contents
train.jsonl: Contains 1/10 of the original data, shuffled
EN.jsonl: English conversations from train.jsonl
ZH.jsonl: Chinese conversations from train.jsonl
Each row represents a conversation with an optional system message, followed by human and GPT turns.
Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.thomas-yanxin-MT-SFT-ShareGPT
thomas-yanxin/MT-SFT-ShareGPT
This is the complete thomas-yanxin/MT-SFT-ShareGPT dataset,
with duplicates removed and the entire dataset shuffled. Sensitive data has been redacted.
For practical work, consider using agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample
which is smaller and split by language.
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full
piece (no time-windowing); training crops sequences from packed token bins.
The source column is the original piece metadata as JSON so a row can be
traced back to Maestro, GiantMIDI, ATEPP, or MusicNet.
Based on MIDI datasets gathered by EPR Labs.
Codec
name: dyadic
tokenizer vocab size: 512
max_time_step: 1.0
n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.mt-dialogues-100k-v2
Multi-turn medical dialogues — V2 (pass@k)
39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup
as V1, but
generated with pass@k=4: a case is regenerated from scratch, graded with
Inspect's model_graded_fact, until the conclusion is graded correct or 4
attempts are spent. 30,234 dialogues (76%) end on a graded-correct
conclusion, roughly 2 attempts per case on average thanks to stopping as soon
as one succeeds.
Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.spai-ss6-corpus-scb-mt-en-th
SPAI SS6 SCB MT EN-TH Corpus Index
Index repo for the SCB MT English-Thai corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: scb_mt_en_th_2020_mt_opus
Rows in canonical config: 4,319,905
Parquet size in canonical config: 0.45 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-scb-mt-en-th.gender-bias-PE
Dataset Card for gender-bias-PE data
Dataset Description
The gender-bias-PE dataset contains the post-edits and associated behavioural data of the human-centered experiments presented in the paper:
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study accepted at EMNLP 2024.
The dataset allows to study the impact of gender bias in Machine Translation (MT) via human-centered measures like post-editing effort (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/gender-bias-PE.hermes_fc_call_mt_R1
hermes_fc_call_mt_R1
繁體中文(台灣用語)函式呼叫(function calling)監督式微調(SFT)資料集。以 NousResearch/hermes-function-calling-v1 為基礎,經過 DeepSeek-R1 推理蒸餾 與 ACE-2 繁體中文翻譯/正規化 而成,最終以 Hermes <tool_call> 格式提供,並保留思考鏈(<think>)。
資料集概述
來源:衍生自 NousResearch/hermes-function-calling-v1(其中含 Glaive 衍生資料)。
產製方式:以 DeepSeek-R1 對原始題目蒸餾出推理與答案,再以 ACE-2 翻譯/正規化為繁體中文(台灣用語)。
用途:SFT,訓練模型以 Hermes <tool_call> 格式進行 function calling,並具備繁體中文思考與回答能力。
主要輸出:data/datasets.jsonl(10,428 筆)。… See the full description on the dataset page: https://huggingface.co/datasets/minyichen/hermes_fc_call_mt_R1.
