datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mTSBench
mTSBench
mTSBench is a collection of 344 multivariate time series from 19 datasets commonly used in anomaly detection research. Each folder corresponds to one dataset and contains *_train.csv, *_test.csv, and *_val.csv files. See data_summary.csv for per-file statistics.
How to download
This repository uses Git LFS for the CSV files.
git lfs install
git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench
Load with Hugging Face
Select one of the… See the full description on the dataset page: https://huggingface.co/datasets/PLAN-Lab/mTSBench.MTS_Dialogue-Clinical_Note
MTS Dialogue (Clinical Note Summarisation)
Main Dataset
The MTS-Dialog dataset is a new collection of 1.7k short doctor-patient conversations and corresponding summaries (section headers and contents).
The training set consists of 1,201 pairs of conversations and associated summaries.
The validation set consists of 100 pairs of conversations and their summaries.
The "dialogue" column contain Doctor-Patient conversation. The "section_text" column contains the Clinical Note of the… See the full description on the dataset page: https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note.AMPBench-MT
AMPBench-MT
AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08.
Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT
The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.MT-intents-dataset-pt-PTsaas-vendor-outage-duration-incident-resolution-time-mttr
How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily
As of 2026-09-24 12:28 UTC. For every incident a vendor posted on its own public status
page with BOTH an opened time and a resolved time, this dataset computes
duration_minutes = resolved_at - started_at
and rolls it up per vendor. It is derived, every day, from the incident table in
saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot
disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.MTBLS289
MTBLS289
A dataset of ~110 paired Whole Slide Images (WSI) and Mass Spectrometry Images (MSI).
Publication: Gerbig, S., Golf, O., Balog, J. et al. Analysis of colorectal adenocarcinoma tissue by desorption electrospray ionization mass spectrometric imaging. Anal Bioanal Chem 403, 2315–2325 (2012).
mtsamplesscreen-benchmarksNeo-GATE
Dataset card for Neo-GATE
Homepage: https://mt.fbk.eu/neo-gate/
Dataset summary
Neo-GATE is a bilingual corpus designed to benchmark the ability of machine translation (MT) systems to translate from English into Italian using gender-inclusive neomorphemes.
It is built upon GATE (Rarrick et al., 2023), a benchmark for the evaluation of gender rewriters and gender bias in MT.
Neo-GATE includes 841 test entries (Neo-GATE.tsv) and 100 dev entries (Neo-GATE-dev.tsv).
Each… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Neo-GATE.fama-data
Dataset Description, Collection, and Source
The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST)
used to train the FAMA models family.
The ASR section of FAMA is derived from the MOSEL data collection, including the automatic
transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset.
The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.stsb-mt-turkish
STSb Turkish
Semantic textual similarity dataset for the Turkish language. It is a machine translation (Azure) of the STSb English dataset. This dataset is not reviewed by expert human translators.
Uploaded from this repository.
Citing & Authors
@misc{celik2020stsbtr,
author = {Emrecan Çelik},
title = {STSB-MT-Turkish},
howpublished = {Hugging Face dataset repository},
url = {https://huggingface.co/datasets/emrecan/stsb-mt-turkish}… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/stsb-mt-turkish.extrinsic_mt_evalmteb-example-submissionmt-en-vi
Dataset Card for Machine Translation Paired English-Vietnamese Sentences
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The language of the dataset text sentence is English ('en') and Vietnamese (vi).
Dataset Structure
Data Instances
An instance example:
{
'en': 'And what I think the world needs now is more connections.',
'vi': 'Và tôi nghĩ điều thế giới đang… See the full description on the dataset page: https://huggingface.co/datasets/ncduy/mt-en-vi.mturk_scoresibom-mtolivers-mtor-atlas
Oliver's mTOR Atlas
The mTOR pathway, mapped by what the evidence can actually carry. This dataset is the curated corpus behind mtor-atlas.org: 414 hand-selected studies on mTOR (mechanistic target of rapamycin) signalling, each labelled by the kind of study behind it, and a list of 149 pathway entities (genes and proteins, complexes, drugs, interventions, biological processes, diseases, outcomes, organelles, nutrients and conditions) that the studies refer to.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/pampalini1/olivers-mtor-atlas.turkish-legal-rag
Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı
Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı.
🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.GeNTE
🚨 GeNTE has been superseded by mGeNTE, a new multilingual release of the corpus with additional annotations.
Dataset Card for GeNTE
Homepage: https://mt.fbk.eu/gente/
Dataset Summary
GeNTE (Gender-Neutral Translation Evaluation) is a natural, bilingual corpus designed to benchmark the ability of machine translation systems to generate gender-neutral translations.
Built from European Parliament speeches, GeNTE comprises a subset of the English-Italian portion… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/GeNTE.MTEB_leaks_and_duplications
LLE MTEB
This dataset lists the presence or absence of leaks and duplicate data in the datasets constituting the MTEB leaderboard (EN & FR).
For more information concerning the methodology and find out what the column names correspond to, please consult the following blog post.To keep things simple, we invite the reader to read the percentages indicated in the text_and_label_test_biased column, which correspond to the proportion of biased data in the test split of the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/MTEB_leaks_and_duplications.formosan-mt
FormosanBank Machine Translation
Public parallel corpora for 15 Indigenous Formosan languages aligned with English and Mandarin Chinese. This release uses canonical MT-standardized Formosan text, excludes Formosan-Taiwan-Bible-Society-Bibles, and keeps DeepL pivot translations in training only.
Commercial AI use is prohibited without prior written permission. See the FormosanBank Terms of Use.
Release Summary
Config
Rows
Train
Validate
Test
Synthetic train… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/formosan-mt.mizo-mteb-retrievaltatoeba-mt-all-in-one
Dataset Card for The Tatoeba Translation Challenge | All In One
~7.3M entries.
Just more user-friendly version that combines all of the entries of original dataset in a single file:
https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt
mlcd-mteb-cifar-eval
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation
Evaluation results accompanying the MTEB integration of two MLCD image encoders
(PR #5406, resolving
issue #2571).
Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the
reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100
image-classification tasks alongside size-matched OpenAI CLIP baselines.
What was measured
Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.MTSBerquadMTSBerquad is a cleaned and enriched dataset SberQuAD transferred to the Generative QA task. All entities were truecased, refactored by hand to improve readability and consistency. Answers have been expanded and rearranged from MLM QA task to Generative/Long Form QA task. MTSBerquad presented in PyCon 2024 by MTS AI Search Group.
Developed by MTS AI Search Group (Krayko Nikita, Laputin Fedor, Sidorov Ivan)
MAGNETbenchmark4CALAMITA24
OPEN evaluation sets for the MAGNET Challenge @ CALAMITA 2024
Last update: 16 Sept 2024
Overview
This dataset represents the OPEN portion of the benchmark of the MAGNET Challenge @CALAMITA2024.
It consists of two Italian/English parallel sets, namely:
dev and devtest taken from the FLORES+ collection (https://github.com/openlanguagedata/flores)
The other three files of the benchmark are not publicly distributed.
Please contact the organizers for information… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MAGNETbenchmark4CALAMITA24.africa-synth-maternal-health-hepatitis-b-mtct-dataset-all
African Hepatitis B MTCT Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-maternal-health-hepatitis-b-mtct-dataset-all.mt5_eurusdmtg-links
mtg-links
This dataset contains 9,879,486 URLs pointing to Magic: The Gathering (MTG) content. These links come from the Internet Archive's Wayback Machine. They cover major MTG community sites, strategy blogs, and official news outlets.
This dataset is the index for a project to build a complete text corpus of Magic: The Gathering strategy, lore, and history.
Content breakdown
Site
Domain
Link Count
mtgsalvation
www.mtgsalvation.com
5,960,584
mtg_wiki… See the full description on the dataset page: https://huggingface.co/datasets/MattDTO/mtg-links.MtG-json-to-ForgeScript
