CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesforce /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.textquestion-answering1K<n<10K115 likes5k downloads1y agoHugging Face02Helsinki-NLP /tatoeba_mtgated Dataset Card for [Dataset Name] Dataset Summary The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development data sorted by language pair. It includes test sets for hundreds of language pairs and is continuously updated. Please, check the version number tag to refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt.texttext-generation1M<n<10M64 likes4k downloads5d agoHugging Face03MultiSynt /MT-Nemotron-CC MultiSynt MultiSynt is an open multilingual synthetic dataset. The MT Nemotron-CC subset of MultiSynt is made of automatic translations into multiple languages from a subset of approximately 100B tokens from the high-quality split of the English Nemotron-CC dataset. This subset is made available using different translation models: Unbabel/Tower-Plus-9B (translations into 16 languages) Unbabel/Tower-Plus-72B (translations into 5 languages) Opus-MT and HPLT-MT (translations… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Nemotron-CC.texttext-generation10B<n<100B13 likes2.1k downloads8mo agoHugging Face04mtybilly /apex-r1-real-world-documents Apex-R1 Real-World Benchmark Documents This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation. The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks. Contents benchmark_documents/ EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.documentdocument-question-answeringn<1K0 likes1.4k downloads3mo agoHugging Face05Magpie-Align /Magpie-Llama-3.1-Pro-MT-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.3k downloads2y agoHugging Face06prem-research /Funcdex-MT-Function-Calling Funcdex-MT-Function-Calling Dataset Funcdex-MT-Function-Calling is a multi-turn function calling dataset designed for training language models to interact with real-world tools and APIs. The dataset contains 1,787 conversations covering 10 individual toolkits and 5 multi-toolkit bundles, with comprehensive system prompts and realistic multi-turn interactions.The code used to generate the dataset can be found here. Models trained on this dataset have excellent… See the full description on the dataset page: https://huggingface.co/datasets/prem-research/Funcdex-MT-Function-Calling.imagetext-generation1K<n<10K3 likes521 downloads11mo agoHugging Face07MultiSynt /MT-Reasoning MultiSynt MultiSynt is an open multilingual synthetic dataset. The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses. lang rows prompt_tokens reasoning_tokens response_tokens total_tokens deu_Latn 17_354_716 1_873_153_732 26_010_932_738 14_862_651_336 42_746_737_806 fra_Latn 17_354_716 1_802_885_115 25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.tabulartext-generation100M<n<1B0 likes519 downloads7mo agoHugging Face08mthreetw /Semantic-Flow-Dynamics-SFD Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus TL;DR: 614 Chinese-language formalized social-science concepts across 25 papers, UUID-linked with typed derivation relations (derives_from, leads_to, falsified_by, …) — usable for knowledge-graph construction, RAG over structured theory, or as a Chinese formal-reasoning corpus. Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com What This Dataset Is This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.textgraph-ml1K<n<10K0 likes510 downloads14d agoHugging Face09docketx /us-caselaw-mt Montana Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Source & credit — Free Law Project / CourtListener Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of 2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-mt.texttext-retrieval10K<n<100K0 likes502 downloads5d agoHugging Face10DigitalLearningGmbH /tatoeba_mt_parquet Dataset Card for DigitalLearningGmbH/tatoeba_mt_parquet This is a mirror of Helsinki-NLP/tatoeba_mt, converted to parquet for compatibility with newer huggingface requirements. Original dataset card follows. Dataset Summary The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/tatoeba_mt_parquet.texttext-generation1M<n<10M1 likes406 downloads5mo agoHugging Face11wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to its EPR Labs source dataset. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.tabulartext-generation1M<n<10M0 likes232 downloads19d agoHugging Face12FBK-MT /Neo-GATE Dataset card for Neo-GATE Homepage: https://mt.fbk.eu/neo-gate/ Dataset summary Neo-GATE is a bilingual corpus designed to benchmark the ability of machine translation (MT) systems to translate from English into Italian using gender-inclusive neomorphemes. It is built upon GATE (Rarrick et al., 2023), a benchmark for the evaluation of gender rewriters and gender bias in MT. Neo-GATE includes 841 test entries (Neo-GATE.tsv) and 100 dev entries (Neo-GATE-dev.tsv). Each… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Neo-GATE.texttranslation1K<n<10K11 likes212 downloads2y agoHugging Face13TurkuNLP /finbenchv2-opengpt-x_truthfulqax-fi-mtThis is an archived version of LumiOpen/opengpt-x_truthfulqax used in Finbench version 2, as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models. Code: https://github.com/LumiOpen/lm-evaluation-harness Citation Information If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from: @misc{thellmann2024crosslingual, title={Towards Cross-Lingual LLM… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-opengpt-x_truthfulqax-fi-mt.texttext-classification1K<n<10K0 likes199 downloads9mo agoHugging Face14fffoivos /hplt-greek-ge8-no-mt-clean60-wave4 HPLT Greek GE8 No-MT Clean60 Wave4 A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass. Snapshot Rows: 48728774 Data parquet files: 250 Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60 Quality bins: 8, 9, 10 MT/register filtering: applied before this release Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.tabulartext-generation10M<n<100M0 likes196 downloads4mo agoHugging Face15pdelobelle /fineweb-dutch-edu-mt FineWeb-Edu Dutch Machine Translated Machine-translated Dutch text dataset derived from the FineWeb-Edu corpus. Dataset Details Source: HuggingFaceFW/fineweb-edu (sample-10BT subset) Translation: English → Dutch using Unbabel/Tower-Plus-9B Size: Up to 1.5M samples Format: Translated text with original metadata Schema text: Machine-translated Dutch text id: Original sample identifier from FineWeb-Edu url: Source URL Quality Notice ⚠️ This… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/fineweb-dutch-edu-mt.texttext-generation1M<n<10M1 likes195 downloads1y agoHugging Face16STEM-AI-mtl /Electrical-engineering To the electrical engineering community This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes. Authors STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec textquestion-answering1K<n<10K61 likes193 downloads2y agoHugging Face17thu-coai /MTAC-IFBench MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding 🌟 Overview MTAC-IFBench benchmarks instruction following in multi-turn agentic coding. Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.texttext-generationn<1K1 likes190 downloads12d agoHugging Face18amalia-llm /bigbenchhard-mt-pt BBH-PT (Big-Bench Hard) Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks. Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations. Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese. Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.textquestion-answering1K<n<10K0 likes175 downloads3mo agoHugging Face19kxiaoqiangrexian /MTS-All MTS-All MTS-All is the data release for the EMNLP 2026 accepted paper Reactivating Test-Time Scaling for Plane Geometry Problem Solving (PDF). Training and evaluation code is available in the ReTTS-PGPS repository. The dataset contains multi-trace supervised fine-tuning data and test files for three plane geometry benchmarks: PGPS9K-All Geometry3K-All GeoQA-All Each training problem is represented with four reasoning traces: Program: symbolic geometry program. COT-program:… See the full description on the dataset page: https://huggingface.co/datasets/kxiaoqiangrexian/MTS-All.imagevisual-question-answering10K<n<100K0 likes170 downloads25d agoHugging Face20TurkuNLP /finbenchv2-squad-strip-fi-mt finbenchv2-squad-strip-fi-mt This dataset is a subset of our SQuAD v2 HF dataset with unanswerable questions removed, to be used within the FIN-bench-v2 benchmark suite. An additional feature of this dataset is that the text in the title fields have been machine-translated to Finnish. Paper: https://huggingface.co/papers/2512.13330 Code: https://github.com/LumiOpen/lm-evaluation-harness Considerations for Using the Data Due to DeepL terms and conditions, this… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-squad-strip-fi-mt.textquestion-answering10K<n<100K0 likes169 downloads9mo agoHugging Face21minpeter /apigen-mt-5k-parsed [PARSED] APIGen-MT-5k The data in this dataset is a full of the original Salesforce/APIGen-MT-5k Subset name multi-turn parallel multiple definition Last turn type number of dataset apigen-mt-5k yes no yes complex 5k This is a re-parsing formatting dataset for the APIGen-MT-5k official dataset. Load the dataset from datasets import load_dataset ds = load_dataset("minpeter/apigen-mt-5k-parsed") print(ds) # DatasetDict({ # train: Dataset({ #… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/apigen-mt-5k-parsed.textquestion-answering1K<n<10K0 likes133 downloads1y agoHugging Face22agentlans /thomas-yanxin-MT-SFT-ShareGPT thomas-yanxin/MT-SFT-ShareGPT This is the complete thomas-yanxin/MT-SFT-ShareGPT dataset, with duplicates removed and the entire dataset shuffled. Sensitive data has been redacted. For practical work, consider using agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample which is smaller and split by language. tabulartext-generation1M<n<10M1 likes117 downloads10mo agoHugging Face23nayohan /Magpie-Air-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Air-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b. This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered. @misc{xu2024magpie, title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing}, author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Air-MT-300K-v0.1-ko.texttext-generation100K<n<1M0 likes112 downloads2y agoHugging Face24agentlans /thomas-yanxin-MT-SFT-ShareGPT-sample MT-SFT-ShareGPT Sample Dataset This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets. Dataset Contents train.jsonl: Contains 1/10 of the original data, shuffled EN.jsonl: English conversations from train.jsonl ZH.jsonl: Chinese conversations from train.jsonl Each row represents a conversation with an optional system message, followed by human and GPT turns. Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.tabulartext-generation1M<n<10M0 likes112 downloads10mo agoHugging Face25LumiOpen /ifeval_mt IFEval Multilingual These are machine-translated versions of Instruction Following Evaluation (IFEval). We will do our best to correct the translations. Translations were done using DeepL and the translations were reviewed and corrected by native speakers. We use this dataset in our fork of LM Eval Harness that supports multilingual ifeval. Supported languages Finnish: machine-translated manually corrected Swedish: machine-translated but not corrected texttext-generation1K<n<10K2 likes111 downloads1y agoHugging Face26Groq /mtob MTOB (Machine Translation from One Book) Last updated: Wednesday, July 9, 2025 Machine Translation from One Book evaluates a language model's ability to translate sentences from English to Kalamang (a low-resource language) and from Kalamang to English. As of July 2, 2025, additional tasks for this groq-bench implementation include: Kalamang-to-English translation adding the option to perform long-context evaluation where the Kalamang corpus is used as input to the model adding… See the full description on the dataset page: https://huggingface.co/datasets/Groq/mtob.texttranslationn<1K0 likes106 downloads1y agoHugging Face27sapienzanlp /ea-mt-benchmark Dataset Card for EA-MT EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate. Here is an example of a simple sentence with a challenging entity mention: English: "What is the plot of The Catcher in the Rye?" Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.texttext-generation10K<n<100K6 likes100 downloads2y agoHugging Face28aisingapore /MultiTurn-Chat-MT-Benchgated SEA-MTBench SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use gpt-4-1106-preview as the judge model and compare against gpt-3.5-turbo-0125 as the baseline model. It is based on MT-Bench and was manually translated by native speakers for Indonesian (id), Javanese (jv), Sundanese (su), and Vietnamese (vi). The Thai split of this dataset uses MT-Bench Thai from the ThaiLLM leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench.texttext-generationn<1K0 likes94 downloads9mo agoHugging Face29MultiSynt /MT-HPLT2c MultiSynt MultiSynt is an open multilingual synthetic dataset. The MT-HPLT2c subset of MultiSynt is a large-scale, machine translated variant of the HPLT v2 English dataset to study LLM training on translated data. From the English source, we offer translations for the following 4 target languages: deu_Latn, fin_Latn, spa_Latn, swe_Latn. For each language, we provide 3 splits: all: The entire data. parallel: A subset of 115,082,738 aligned documents, such that a document… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-HPLT2c.texttext-generation1B<n<10B0 likes88 downloads8mo agoHugging Face30nayohan /Magpie-Pro-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Pro-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b. This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered. @misc{xu2024magpie, title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing}, author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Pro-MT-300K-v0.1-ko.texttext-generation100K<n<1M24 likes83 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.