CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Magpie-Align /Magpie-Llama-3.1-Pro-MT-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.2k downloads2y agoHugging Face02MultiSynt /MT-Reasoning MultiSynt MultiSynt is an open multilingual synthetic dataset. The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses. lang rows prompt_tokens reasoning_tokens response_tokens total_tokens deu_Latn 17_354_716 1_873_153_732 26_010_932_738 14_862_651_336 42_746_737_806 fra_Latn 17_354_716 1_802_885_115 25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.tabulartext-generation100M<n<1B0 likes512 downloads7mo agoHugging Face03wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to its EPR Labs source dataset. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.tabulartext-generation1M<n<10M0 likes229 downloads19d agoHugging Face04fffoivos /hplt-greek-ge8-no-mt-clean60-wave4 HPLT Greek GE8 No-MT Clean60 Wave4 A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass. Snapshot Rows: 48728774 Data parquet files: 250 Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60 Quality bins: 8, 9, 10 MT/register filtering: applied before this release Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.tabulartext-generation10M<n<100M0 likes179 downloads4mo agoHugging Face05agentlans /thomas-yanxin-MT-SFT-ShareGPT-sample MT-SFT-ShareGPT Sample Dataset This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets. Dataset Contents train.jsonl: Contains 1/10 of the original data, shuffled EN.jsonl: English conversations from train.jsonl ZH.jsonl: Chinese conversations from train.jsonl Each row represents a conversation with an optional system message, followed by human and GPT turns. Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.tabulartext-generation1M<n<10M0 likes109 downloads10mo agoHugging Face06agentlans /thomas-yanxin-MT-SFT-ShareGPT thomas-yanxin/MT-SFT-ShareGPT This is the complete thomas-yanxin/MT-SFT-ShareGPT dataset, with duplicates removed and the entire dataset shuffled. Sensitive data has been redacted. For practical work, consider using agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample which is smaller and split by language. tabulartext-generation1M<n<10M1 likes109 downloads10mo agoHugging Face07wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.tabulartext-generation10K<n<100K0 likes70 downloads1mo agoHugging Face08zacbrld /mt-dialogues-100k-v2 Multi-turn medical dialogues — V2 (pass@k) 39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup as V1, but generated with pass@k=4: a case is regenerated from scratch, graded with Inspect's model_graded_fact, until the conclusion is graded correct or 4 attempts are spent. 30,234 dialogues (76%) end on a graded-correct conclusion, roughly 2 attempts per case on average thanks to stopping as soon as one succeeds. Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.tabulartext-generation10K<n<100K0 likes24 downloads1mo agoHugging Face09SPAISS6F1 /spai-ss6-corpus-scb-mt-en-th SPAI SS6 SCB MT EN-TH Corpus Index Index repo for the SCB MT English-Thai corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: scb_mt_en_th_2020_mt_opus Rows in canonical config: 4,319,905 Parquet size in canonical config: 0.45 GB Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-scb-mt-en-th.tabulartext-generationn<1K0 likes14 downloads4mo agoHugging Face10FBK-MT /gender-bias-PEgated Dataset Card for gender-bias-PE data Dataset Description The gender-bias-PE dataset contains the post-edits and associated behavioural data of the human-centered experiments presented in the paper: What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study accepted at EMNLP 2024. The dataset allows to study the impact of gender bias in Machine Translation (MT) via human-centered measures like post-editing effort (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/gender-bias-PE.tabulartranslation1K<n<10K3 likes11 downloads2y agoHugging Face11minyichen /hermes_fc_call_mt_R1gated hermes_fc_call_mt_R1 繁體中文(台灣用語)函式呼叫(function calling)監督式微調(SFT)資料集。以 NousResearch/hermes-function-calling-v1 為基礎,經過 DeepSeek-R1 推理蒸餾 與 ACE-2 繁體中文翻譯/正規化 而成,最終以 Hermes <tool_call> 格式提供,並保留思考鏈(<think>)。 資料集概述 來源:衍生自 NousResearch/hermes-function-calling-v1(其中含 Glaive 衍生資料)。 產製方式:以 DeepSeek-R1 對原始題目蒸餾出推理與答案,再以 ACE-2 翻譯/正規化為繁體中文(台灣用語)。 用途:SFT,訓練模型以 Hermes <tool_call> 格式進行 function calling,並具備繁體中文思考與回答能力。 主要輸出:data/datasets.jsonl(10,428 筆)。… See the full description on the dataset page: https://huggingface.co/datasets/minyichen/hermes_fc_call_mt_R1.tabulartext-generation10K<n<100K0 likes7 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.