CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PLAN-Lab /mTSBench mTSBench mTSBench is a collection of 344 multivariate time series from 19 datasets commonly used in anomaly detection research. Each folder corresponds to one dataset and contains *_train.csv, *_test.csv, and *_val.csv files. See data_summary.csv for per-file statistics. How to download This repository uses Git LFS for the CSV files. git lfs install git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench Load with Hugging Face Select one of the… See the full description on the dataset page: https://huggingface.co/datasets/PLAN-Lab/mTSBench.tabular10M<n<100M3 likes993 downloads9d agoHugging Face02har1 /MTS_Dialogue-Clinical_Note MTS Dialogue (Clinical Note Summarisation) Main Dataset The MTS-Dialog dataset is a new collection of 1.7k short doctor-patient conversations and corresponding summaries (section headers and contents). The training set consists of 1,201 pairs of conversations and associated summaries. The validation set consists of 100 pairs of conversations and their summaries. The "dialogue" column contain Doctor-Patient conversation. The "section_text" column contains the Clinical Note of the… See the full description on the dataset page: https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note.textfeature-extraction1K<n<10K13 likes565 downloads2y agoHugging Face03ZihengZhou06 /AMPBench-MT AMPBench-MT AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08. Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.tabulartabular-classification100K<n<1M0 likes430 downloads2mo agoHugging Face04OpenVoiceOS /MT-intents-dataset-pt-PTtext10K<n<100K0 likes340 downloads1y agoHugging Face05APProjects /saas-vendor-outage-duration-incident-resolution-time-mttr How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily As of 2026-09-24 12:28 UTC. For every incident a vendor posted on its own public status page with BOTH an opened time and a resolved time, this dataset computes duration_minutes = resolved_at - started_at and rolls it up per vendor. It is derived, every day, from the incident table in saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.tabulartabular-regression10K<n<100K0 likes336 downloads2d agoHugging Face06aryan-f /MTBLS289 MTBLS289 A dataset of ~110 paired Whole Slide Images (WSI) and Mass Spectrometry Images (MSI). Publication: Gerbig, S., Golf, O., Balog, J. et al. Analysis of colorectal adenocarcinoma tissue by desorption electrospray ionization mass spectrometric imaging. Anal Bioanal Chem 403, 2315–2325 (2012). imageimage-to-imagen<1K0 likes299 downloads2mo agoHugging Face07harishnair04 /mtsamplestext1K<n<10K4 likes284 downloads2y agoHugging Face08mtapiapacheco /screen-benchmarkstext1M<n<10M0 likes259 downloads2mo agoHugging Face09FBK-MT /Neo-GATE Dataset card for Neo-GATE Homepage: https://mt.fbk.eu/neo-gate/ Dataset summary Neo-GATE is a bilingual corpus designed to benchmark the ability of machine translation (MT) systems to translate from English into Italian using gender-inclusive neomorphemes. It is built upon GATE (Rarrick et al., 2023), a benchmark for the evaluation of gender rewriters and gender bias in MT. Neo-GATE includes 841 test entries (Neo-GATE.tsv) and 100 dev entries (Neo-GATE-dev.tsv). Each… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Neo-GATE.texttranslation1K<n<10K11 likes219 downloads2y agoHugging Face10FBK-MT /fama-data Dataset Description, Collection, and Source The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST) used to train the FAMA models family. The ASR section of FAMA is derived from the MOSEL data collection, including the automatic transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset. The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.tabulartranslation1M<n<10M2 likes197 downloads1y agoHugging Face11emrecan /stsb-mt-turkish STSb Turkish Semantic textual similarity dataset for the Turkish language. It is a machine translation (Azure) of the STSb English dataset. This dataset is not reviewed by expert human translators. Uploaded from this repository. Citing & Authors @misc{celik2020stsbtr, author = {Emrecan Çelik}, title = {STSB-MT-Turkish}, howpublished = {Hugging Face dataset repository}, url = {https://huggingface.co/datasets/emrecan/stsb-mt-turkish}… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/stsb-mt-turkish.texttext-classification1K<n<10K8 likes196 downloads21d agoHugging Face12uoe-nlp /extrinsic_mt_evaltabular10K<n<100K0 likes177 downloads3y agoHugging Face13mteb /mteb-example-submissiontextn<1K0 likes175 downloads4y agoHugging Face14ncduy /mt-en-vi Dataset Card for Machine Translation Paired English-Vietnamese Sentences Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages The language of the dataset text sentence is English ('en') and Vietnamese (vi). Dataset Structure Data Instances An instance example: { 'en': 'And what I think the world needs now is more connections.', 'vi': 'Và tôi nghĩ điều thế giới đang… See the full description on the dataset page: https://huggingface.co/datasets/ncduy/mt-en-vi.text1M<n<10M12 likes147 downloads4y agoHugging Face15mdcgp /mturk_scorestabular1K<n<10K1 likes129 downloads2y agoHugging Face16howard-nlp /ibom-mttexttranslation10K<n<100K0 likes84 downloads1y agoHugging Face17pampalini1 /olivers-mtor-atlas Oliver's mTOR Atlas The mTOR pathway, mapped by what the evidence can actually carry. This dataset is the curated corpus behind mtor-atlas.org: 414 hand-selected studies on mTOR (mechanistic target of rapamycin) signalling, each labelled by the kind of study behind it, and a list of 149 pathway entities (genes and proteins, complexes, drugs, interventions, biological processes, diseases, outcomes, organelles, nutrients and conditions) that the studies refer to. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/pampalini1/olivers-mtor-atlas.tabulartext-classificationn<1K0 likes80 downloads9h agoHugging Face18mtntasci /turkish-legal-rag Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı. 🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.tabulartext-retrieval1K<n<10K2 likes78 downloads4mo agoHugging Face19FBK-MT /GeNTE 🚨 GeNTE has been superseded by mGeNTE, a new multilingual release of the corpus with additional annotations. Dataset Card for GeNTE Homepage: https://mt.fbk.eu/gente/ Dataset Summary GeNTE (Gender-Neutral Translation Evaluation) is a natural, bilingual corpus designed to benchmark the ability of machine translation systems to generate gender-neutral translations. Built from European Parliament speeches, GeNTE comprises a subset of the English-Italian portion… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/GeNTE.texttranslation1K<n<10K13 likes77 downloads2y agoHugging Face20lbourdois /MTEB_leaks_and_duplications LLE MTEB This dataset lists the presence or absence of leaks and duplicate data in the datasets constituting the MTEB leaderboard (EN & FR). For more information concerning the methodology and find out what the column names correspond to, please consult the following blog post.To keep things simple, we invite the reader to read the percentages indicated in the text_and_label_test_biased column, which correspond to the proportion of biased data in the test split of the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/MTEB_leaks_and_duplications.textn<1K0 likes76 downloads2y agoHugging Face21FormosanBank /formosan-mt FormosanBank Machine Translation Public parallel corpora for 15 Indigenous Formosan languages aligned with English and Mandarin Chinese. This release uses canonical MT-standardized Formosan text, excludes Formosan-Taiwan-Bible-Society-Bibles, and keeps DeepL pivot translations in training only. Commercial AI use is prohibited without prior written permission. See the FormosanBank Terms of Use. Release Summary Config Rows Train Validate Test Synthetic train… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/formosan-mt.texttranslation100K<n<1M0 likes69 downloads1mo agoHugging Face22robzchhangte /mizo-mteb-retrievaltext100K<n<1M0 likes69 downloads2mo agoHugging Face230x22almostEvil /tatoeba-mt-all-in-one Dataset Card for The Tatoeba Translation Challenge | All In One ~7.3M entries. Just more user-friendly version that combines all of the entries of original dataset in a single file: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt text1M<n<10M0 likes64 downloads3y agoHugging Face24b4ph /mlcd-mteb-cifar-eval MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation Evaluation results accompanying the MTEB integration of two MLCD image encoders (PR #5406, resolving issue #2571). Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100 image-classification tasks alongside size-matched OpenAI CLIP baselines. What was measured Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.tabularimage-classificationn<1K0 likes62 downloads18d agoHugging Face25MTS-AI-SearchSkill /MTSBerquadMTSBerquad is a cleaned and enriched dataset SberQuAD transferred to the Generative QA task. All entities were truecased, refactored by hand to improve readability and consistency. Answers have been expanded and rearranged from MLM QA task to Generative/Long Form QA task. MTSBerquad presented in PyCon 2024 by MTS AI Search Group. Developed by MTS AI Search Group (Krayko Nikita, Laputin Fedor, Sidorov Ivan) textquestion-answering10K<n<100K8 likes54 downloads2y agoHugging Face26FBK-MT /MAGNETbenchmark4CALAMITA24 OPEN evaluation sets for the MAGNET Challenge @ CALAMITA 2024 Last update: 16 Sept 2024 Overview This dataset represents the OPEN portion of the benchmark of the MAGNET Challenge @CALAMITA2024. It consists of two Italian/English parallel sets, namely: dev and devtest taken from the FLORES+ collection (https://github.com/openlanguagedata/flores) The other three files of the benchmark are not publicly distributed. Please contact the organizers for information… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MAGNETbenchmark4CALAMITA24.texttranslation1K<n<10K1 likes54 downloads2y agoHugging Face27electricsheepafrica /africa-synth-maternal-health-hepatitis-b-mtct-dataset-all African Hepatitis B MTCT Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-maternal-health-hepatitis-b-mtct-dataset-all.tabulartabular-classification1K<n<10K1 likes54 downloads2mo agoHugging Face28jukik45 /mt5_eurusdtabular100K<n<1M0 likes53 downloads3d agoHugging Face29MattDTO /mtg-links mtg-links This dataset contains 9,879,486 URLs pointing to Magic: The Gathering (MTG) content. These links come from the Internet Archive's Wayback Machine. They cover major MTG community sites, strategy blogs, and official news outlets. This dataset is the index for a project to build a complete text corpus of Magic: The Gathering strategy, lore, and history. Content breakdown Site Domain Link Count mtgsalvation www.mtgsalvation.com 5,960,584 mtg_wiki… See the full description on the dataset page: https://huggingface.co/datasets/MattDTO/mtg-links.texttext-retrieval1M<n<10M0 likes49 downloads4mo agoHugging Face30404NotF0und /MtG-json-to-ForgeScripttext10K<n<100K0 likes48 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.