CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jpwahle /machine-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools. It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses). The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.texttext-classification100K<n<1M7 likes130 downloads1y agoHugging Face02Silly-Machine /TuPyE-Dataset Portuguese Hate Speech Expanded Dataset (TuPyE) TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically selected for the purpose of hate speech detection within diverse social network contexts. This augmented dataset integrates supplementary annotations and amalgamates with datasets sourced from Fortuna et al. (2019), Leite et al. (2020), and Vargas et al. (2022), complemented by an infusion of 10,000 original… See the full description on the dataset page: https://huggingface.co/datasets/Silly-Machine/TuPyE-Dataset.tabulartext-classification10K<n<100K5 likes100 downloads3y agoHugging Face03jaio98 /basque_dialect_machine_translationtext100K<n<1M0 likes86 downloads3mo agoHugging Face04shwetha729 /quantum-machine-learninga continuous data scrape of arxiv and google scholar papers of quantum machine learning papers particularly regarding climate. tabularn<1K1 likes77 downloads4y agoHugging Face05HR-machine /QM9-Datasettabular100K<n<1M2 likes61 downloads2y agoHugging Face06gospelgit /Nigeria_Machinery_Dataset Nigeria Machinery Usage and Failures Dataset A structured numeric dataset covering machinery usage rates, equipment failures, capacity utilization, maintenance costs, and operational downtime across Nigeria's industrial manufacturing and oil & gas sectors, 2006–2025. It ships with a companion chain-of-thought reasoning layer derived directly from the records, for fine-tuning and evaluating LLMs on domain-grounded numeric tasks. This dataset addresses a real gap: machine-level… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/Nigeria_Machinery_Dataset.tabulartabular-classificationn<1K0 likes57 downloads3mo agoHugging Face07Tran1312 /MachineTranslation_en_viDữ liệu được thu thập từ nhiều nguồn: CCMatrix: https://opus.nlpl.eu/CCMatrix/en&vi/v1/CCMatrix OpenSubtitles: https://opus.nlpl.eu/OpenSubtitles/en&vi/v2024/OpenSubtitles MultiHPLT: https://opus.nlpl.eu/MultiHPLT/en&vi/v2/MultiHPLT CCAligned: https://opus.nlpl.eu/CCAligned/en&vi/v1/CCAligned ParaCrawl: https://opus.nlpl.eu/ParaCrawl-Bonus/en&vi/v9/ParaCrawl-Bonus PhoMT: https://huggingface.co/datasets/ura-hcmut/PhoMTVietAI: https://huggingface.co/datasets/wanhin/VietAI_MTet Dữ liệu đã trải… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/MachineTranslation_en_vi.texttranslation10M<n<100M1 likes53 downloads8mo agoHugging Face08alphuuuu02 /Machine-Learning-Credit-Card-Fraud-Detection-Projecttabular100K<n<1M2 likes53 downloads3mo agoHugging Face09kaenakiakona /machinetranslationspanishtext100K<n<1M1 likes45 downloads3y agoHugging Face10HR-machine /ESolDataset Card for ESol (Estimated Solubility) Dataset Dataset Summary The ESOL dataset is designed for estimating the aqueous solubility of chemical compounds directly from their molecular structure. This dataset includes 2,874 experimentally measured solubility values. The most significant features for predicting solubility include calculated octanol-water partition coefficient (logP), molecular weight, the proportion of heavy atoms in aromatic systems, and the number of rotatable bonds. tabular1K<n<10K0 likes34 downloads2y agoHugging Face11Silly-Machine /TuPy-Dataset Portuguese Hate Speech Dataset (TuPy) The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is comprised of 10,000 (ten thousand) unpublished, annotated, and anonymized documents collected on Twitter (currently known as X) in 2023. This repository is organized as follows: root. ├── binary : binary… See the full description on the dataset page: https://huggingface.co/datasets/Silly-Machine/TuPy-Dataset.tabulartext-classification10K<n<100K1 likes33 downloads3y agoHugging Face12dmatekenya /chichewa-machine-translationtexttranslation10K<n<100K3 likes33 downloads2y agoHugging Face13cuonguyenphu /Titanic-Machine-Learning-from-Disaster-0.77751tabularn<1K0 likes30 downloads15d agoHugging Face14HR-machine /TOX21tabular1K<n<10K0 likes26 downloads2y agoHugging Face15newadays /menyo_20k_a_multi_domain_english_yoruba_corpus_for_machine_translationtext10K<n<100K1 likes24 downloads1y agoHugging Face16Circularmachines /Batch_indexing_machine_tokenstabular1M<n<10M0 likes22 downloads3y agoHugging Face17prsdm /Machine-Learning-QA-datasettextn<1K11 likes20 downloads3y agoHugging Face18AyonRoy29 /informal_bn-en_machine_translation_datasettext10K<n<100K0 likes18 downloads1y agoHugging Face19electricsheepafrica /Africa-Automated-Teller-Machines-ATMs-per-100000-adults Africa Automated Teller Machines ATMs per 100000 adults | Africa (World Bank) Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Automated-Teller-Machines-ATMs-per-100000-adults.tabulartabular-classificationn<1K0 likes17 downloads1mo agoHugging Face20HomieSun /machine_mindset_mbti_infp_sfttext10K<n<100K0 likes15 downloads2y agoHugging Face21HR-machine /ClinToxtabular1K<n<10K1 likes15 downloads2y agoHugging Face22Kenneth12 /MachineLearning_EmojiDataset_Nov17text1K<n<10K1 likes13 downloads3y agoHugging Face23HR-machine /QM7-Datasettext1K<n<10K0 likes13 downloads2y agoHugging Face24nasserCha /machine-anomaly-detectiontabular100K<n<1M1 likes13 downloads10mo agoHugging Face25introvoyz041 /Biogas-Production-Machine-Learning-Analysistabularn<1K0 likes13 downloads4mo agoHugging Face26readerbench /ro-human-machine-60kThe corpus for this study consists of multiple datasets of comparable text lengths, both machine-generated and human-written. 1401 books: 841 manually written abstracts provided by the Central University Library of Bucharest, representing descriptions of Romanian old documents (literary magazines and books dated between the 19th century and the present), 560 books descriptions (cartigratis.com, accessed 8 January 2024); 4320 news articles crawled from DigiNews (digi24.ro, accessed 8… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-human-machine-60k.texttext-generation1K<n<10K2 likes12 downloads3y agoHugging Face27ReySajju742 /machine-anntext10M<n<100M0 likes10 downloads2y agoHugging Face28Mridul-Dixit /Machine-Learning-QA-Datasettextn<1K1 likes8 downloads2y agoHugging Face29whiteOUO /Ladder-machine-learning-MCQstextn<1K0 likes8 downloads2y agoHugging Face30HomieSun /machine_mindset_sft_estjtext10K<n<100K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.