datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
machine-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools.
It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses).
The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.TuPyE-Dataset
Portuguese Hate Speech Expanded Dataset (TuPyE)
TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically
selected for the purpose of hate speech detection within diverse social network contexts.
This augmented dataset integrates supplementary annotations and amalgamates with datasets sourced from
Fortuna et al. (2019),
Leite et al. (2020),
and Vargas et al. (2022),
complemented by an infusion of 10,000 original… See the full description on the dataset page: https://huggingface.co/datasets/Silly-Machine/TuPyE-Dataset.basque_dialect_machine_translationquantum-machine-learninga continuous data scrape of arxiv and google scholar papers of quantum machine learning papers particularly regarding climate.
QM9-DatasetNigeria_Machinery_Dataset
Nigeria Machinery Usage and Failures Dataset
A structured numeric dataset covering machinery usage rates, equipment failures,
capacity utilization, maintenance costs, and operational downtime across Nigeria's
industrial manufacturing and oil & gas sectors, 2006–2025. It ships
with a companion chain-of-thought reasoning layer derived directly from the
records, for fine-tuning and evaluating LLMs on domain-grounded numeric tasks.
This dataset addresses a real gap: machine-level… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/Nigeria_Machinery_Dataset.MachineTranslation_en_viDữ liệu được thu thập từ nhiều nguồn:
CCMatrix: https://opus.nlpl.eu/CCMatrix/en&vi/v1/CCMatrix
OpenSubtitles: https://opus.nlpl.eu/OpenSubtitles/en&vi/v2024/OpenSubtitles
MultiHPLT: https://opus.nlpl.eu/MultiHPLT/en&vi/v2/MultiHPLT
CCAligned: https://opus.nlpl.eu/CCAligned/en&vi/v1/CCAligned
ParaCrawl: https://opus.nlpl.eu/ParaCrawl-Bonus/en&vi/v9/ParaCrawl-Bonus
PhoMT: https://huggingface.co/datasets/ura-hcmut/PhoMTVietAI: https://huggingface.co/datasets/wanhin/VietAI_MTet
Dữ liệu đã trải… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/MachineTranslation_en_vi.Machine-Learning-Credit-Card-Fraud-Detection-ProjectmachinetranslationspanishESolDataset Card for ESol (Estimated Solubility) Dataset
Dataset Summary
The ESOL dataset is designed for estimating the aqueous solubility
of chemical compounds directly from their molecular structure.
This dataset includes 2,874 experimentally measured solubility
values. The most significant features for predicting solubility
include calculated octanol-water partition coefficient (logP),
molecular weight, the proportion of heavy atoms in aromatic systems,
and the number of rotatable bonds.
TuPy-Dataset
Portuguese Hate Speech Dataset (TuPy)
The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML)
and natural language processing (NLP) techniques. TuPy is comprised of 10,000 (ten thousand) unpublished, annotated, and anonymized documents collected
on Twitter (currently known as X) in 2023.
This repository is organized as follows:
root.
├── binary : binary… See the full description on the dataset page: https://huggingface.co/datasets/Silly-Machine/TuPy-Dataset.chichewa-machine-translationTitanic-Machine-Learning-from-Disaster-0.77751TOX21menyo_20k_a_multi_domain_english_yoruba_corpus_for_machine_translationBatch_indexing_machine_tokensMachine-Learning-QA-datasetinformal_bn-en_machine_translation_datasetAfrica-Automated-Teller-Machines-ATMs-per-100000-adults
Africa Automated Teller Machines ATMs per 100000 adults | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Automated-Teller-Machines-ATMs-per-100000-adults.machine_mindset_mbti_infp_sftClinToxMachineLearning_EmojiDataset_Nov17QM7-Datasetmachine-anomaly-detectionBiogas-Production-Machine-Learning-Analysisro-human-machine-60kThe corpus for this study consists of multiple datasets of comparable text lengths, both machine-generated and human-written.
1401 books:
841 manually written abstracts provided by the Central University Library of Bucharest, representing descriptions of Romanian old documents (literary magazines and books dated between the 19th century and the present),
560 books descriptions (cartigratis.com, accessed 8 January 2024);
4320 news articles crawled from DigiNews (digi24.ro, accessed 8… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-human-machine-60k.machine-annMachine-Learning-QA-DatasetLadder-machine-learning-MCQsmachine_mindset_sft_estj
