CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aisingapore /NLG-Machine-Translationgated SEA Machine Translation SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.texttext-generation10K<n<100K2 likes2.9k downloads9mo agoHugging Face02MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes901 downloads10mo agoHugging Face03jpwahle /machine-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools. It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses). The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.texttext-classification100K<n<1M7 likes116 downloads1y agoHugging Face04Neura-parse /quantum-machine-learning-models Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.tabulartext-generation100K<n<1M1 likes64 downloads3mo agoHugging Face05gospelgit /Nigeria_Machinery_Dataset Nigeria Machinery Usage and Failures Dataset A structured numeric dataset covering machinery usage rates, equipment failures, capacity utilization, maintenance costs, and operational downtime across Nigeria's industrial manufacturing and oil & gas sectors, 2006–2025. It ships with a companion chain-of-thought reasoning layer derived directly from the records, for fine-tuning and evaluating LLMs on domain-grounded numeric tasks. This dataset addresses a real gap: machine-level… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/Nigeria_Machinery_Dataset.tabulartabular-classificationn<1K0 likes60 downloads3mo agoHugging Face06Neura-parse /quantum-machine-learning-theory Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data A research-depth, proof-oriented vertical on the learning theory of quantum models and quantum data. Covers why parameterized quantum circuits train or don't (barren plateaus), what they can represent, when they generalize or provably beat classical models, and — for quantum data — how to predict properties of unknown states/channels with few measurements (classical… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-theory.tabulartext-generation100K<n<1M1 likes60 downloads3mo agoHugging Face07machinelearnear /multiturn_chat_milei_gpt Milei-GPT Dataset Che y si queremos hacer un LLM que hable de la misma forma que un famoso ... como hacemos? Este repo es una excusa para aprender a preparar un dataset para fine-tunear algún LLM, aprender como evaluarlo, como tokenizarlo, como extenderlo de formar sintética, y tantas otras cosas. Al final, si todo sale bien, vamos a tener un modelo que va a hablar como la persona que elegimos, y le podemos poner un RAG (retrieval augmented generation) encima para que nos traiga un… See the full description on the dataset page: https://huggingface.co/datasets/machinelearnear/multiturn_chat_milei_gpt.tabularquestion-answeringn<1K7 likes21 downloads2y agoHugging Face08Lots-of-LoRAs /task718_mmmlu_answer_generation_machine_learning Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task718_mmmlu_answer_generation_machine_learning Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task718_mmmlu_answer_generation_machine_learning.texttext-generationn<1K0 likes21 downloads2y agoHugging Face09BearNetworkChain /Machine-Checkable-Blockchain-Execution-Specification 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Machine-Checkable-Blockchain-Execution-Specification.textquestion-answeringn<1K1 likes20 downloads4mo agoHugging Face10BNES-BRNKC /Machine-Checkable-Blockchain-Execution-Specification 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Machine-Checkable-Blockchain-Execution-Specification.textquestion-answeringn<1K0 likes20 downloads4mo agoHugging Face11rubenbalbastre /machine-unlearning-holdout-evals rubenbalbastre/machine-unlearning-holdout-evals Hold-out completions and LLM-judge rubric labels for targeted machine-unlearning experiments. The dataset contains 17,280 prompt/completion evaluations from 288 model runs across 10 target entities. Each row includes the model size, reward function, training variant, and six independent boolean rubric labels. Rubrics lexical_leakage: the completion mentions the target or a surface variant. semantic_leakage: the… See the full description on the dataset page: https://huggingface.co/datasets/rubenbalbastre/machine-unlearning-holdout-evals.texttext-generation10K<n<100K0 likes16 downloads1d agoHugging Face12readerbench /ro-human-machine-60kThe corpus for this study consists of multiple datasets of comparable text lengths, both machine-generated and human-written. 1401 books: 841 manually written abstracts provided by the Central University Library of Bucharest, representing descriptions of Romanian old documents (literary magazines and books dated between the 19th century and the present), 560 books descriptions (cartigratis.com, accessed 8 January 2024); 4320 news articles crawled from DigiNews (digi24.ro, accessed 8… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-human-machine-60k.texttext-generation1K<n<10K2 likes11 downloads3y agoHugging Face13NRC-CNRC /Machine-Generated-Reviews-0.1 Machine Generated Reviews This dataset contains the machine generated peer reviews used in the study of machine generated text (MGT) output syntactic homogenization in "Emphasizing the Commendable": A Study of Homogenized Transitive Verb Constructions in Machine Generated Peer Reviews. The corresponding academic research papers and official reviews are available on OpenReview. The machine generated peer reviews are produced by three LLMs with a diverse background. The prompts and… See the full description on the dataset page: https://huggingface.co/datasets/NRC-CNRC/Machine-Generated-Reviews-0.1.textother100K<n<1M0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.