CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01unitedideas /pending-medicare-provider-enrollment-data Pending Medicare Provider Enrollment Data This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication. Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.tabularn<1K0 likes1.7k downloads2mo agoHugging Face02Anthropic /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/enabling-independent-research.tabular1K<n<10K36 likes1.5k downloads28d agoHugging Face03gplsi /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes1.4k downloads9mo agoHugging Face04beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.3k downloads7mo agoHugging Face05Intelligent-Internet /wikipedia_en wikipedia_en This is a curated Wikipedia English dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.tabularfeature-extraction10M<n<100M2 likes1.2k downloads1y agoHugging Face06ento3686 /TCGA_OncoTree_pt2 TCGA_OncoTree_pt2 1. Tổng quan [CẦN ĐIỀN THỦ CÔNG: mục đích, ngữ cảnh tạo dataset] Tổng số bản ghi (cộng tất cả manifest phát hiện được): 23984 Số manifest phát hiện được trong bộ nhớ: 3 (df, labels_df, progress) Repo HuggingFace chính: ento3686/TCGA_OncoTree_pt2 ⚠️ Dataset được lưu trên 2 repo/tài khoản HuggingFace khác nhau: ento3686/TCGA_OncoTree_pt2 (biến: REPO_ID_2, UPLOAD_REPO_ID, CENTRAL_PROGRESS_REPO_ID, _repo_id_var) tuna2004/TCGA_OncoTree (biến:… See the full description on the dataset page: https://huggingface.co/datasets/ento3686/TCGA_OncoTree_pt2.image10K<n<100K0 likes1.1k downloads21d agoHugging Face07electricsheepafrica /Environment-and-Natural-Resources-Indicators-For-African-Countries Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K0 likes1.1k downloads1mo agoHugging Face08yipyany /ted-translation-decisions-en-zh TED Translation Decision Dataset (EN–ZH 英-简中) 🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁 🧩 Searchable Keywords translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual, semantic nuance, translation rationale dataset, Chinese translation, English translation dataset, word-level translation, interpretability, translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.tabulartranslationn<1K1 likes1.1k downloads3h agoHugging Face09blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes990 downloads2y agoHugging Face10enviroscientist /EnviroExam Dataset Summary EnviroExam focuses on 42 core courses from the environmental science curriculum at Harbin Institute of Technology, after excluding general, duplicate, and practical courses from a total of 141 courses across undergraduate, master's, and doctoral programs. For these 42 courses, initial draft questions were generated using GPT-4 and Claude, combined with customized prompts. These drafts were then refined and proofread manually, resulting in a total of 1,290… See the full description on the dataset page: https://huggingface.co/datasets/enviroscientist/EnviroExam.texttext-classificationn<1K3 likes926 downloads2y agoHugging Face11npvinHnivqn /EnglishDictionarytexttoken-classification100K<n<1M5 likes856 downloads3y agoHugging Face12electricsheepafrica /Energy-Indicators-For-African-Countries Energy Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Energy-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K1 likes801 downloads1mo agoHugging Face13blo05 /cleaned_wiki_en_60-80text1M<n<10M1 likes557 downloads4y agoHugging Face14vitaliy-sharandin /energy-consumption-hourly-spaintabular10K<n<100K2 likes548 downloads3y agoHugging Face15engels /spotify-tracks-lite Context This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance. All data taken from Spotify API and is open source. This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic. Column Description danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.tabular10K<n<100K3 likes532 downloads2y agoHugging Face16serenityyyyy /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes525 downloads6mo agoHugging Face17blo05 /cleaned_wiki_en_40-60text1M<n<10M1 likes503 downloads4y agoHugging Face18marksverdhei /wordnet-definitions-en-2021 Wordnet definitions for English Dataset by Princeton WordNet and the Open English WordNet team https://github.com/globalwordnet/english-wordnet This dataset contains every entry in wordnet that has a definition and an example. Be aware that the word "null" can be misinterpreted as a null value if loading it in with e.g. pandas text10K<n<100K11 likes498 downloads1y agoHugging Face19blo05 /cleaned_wiki_en_80-100text1M<n<10M0 likes484 downloads4y agoHugging Face20aurelium /github-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table. It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and totals of the amounts of various events on those repositories in that time period. This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing, so I figured I would save the next person to try this some effort. tabular100M<n<1B7 likes450 downloads2y agoHugging Face21Alvaro8gb /enfermedades-wiki-marzo-2024 English Version This dataset contains detailed information on a total of 945 diseases, extracted from Wikipedia in Spanish (https://es.wikipedia.org/) in March 2024. The main purpose of this dataset is to serve as a comprehensive resource for training Large Language Models (LLMs) in Spanish, specifically for instruction tuning, pre-training, and other natural language processing (NLP) tasks. This dataset promises to be a valuable tool for research and development in Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/Alvaro8gb/enfermedades-wiki-marzo-2024.texttext-generationn<1K1 likes447 downloads3y agoHugging Face22blo05 /cleaned_wiki_enCleaned wikipedia dataset text1M<n<10M4 likes416 downloads4y agoHugging Face23electricsheepafrica /africa-synth-energy-oilgas-gas-infrastructure-nigeria Africa Synth Energy Oilgas Gas Infrastructure Nigeria | Africa (Electric Sheep Africa metadata inventory) Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-infrastructure-nigeria.texttabular-classificationn<1K0 likes387 downloads1mo agoHugging Face24BeIR /dbpedia-entity-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/dbpedia-entity-qrels.texttext-retrieval10K<n<100K0 likes329 downloads4y agoHugging Face25uneiaparjour /base-en Base uneIAparjour.fr — English — Generative AI Apps Open dataset listing one generative AI tool per day since February 16, 2023, featured on uneiaparjour.fr and translated to English. Description Every day, a new free or freemium generative AI tool is tested, described and categorized — originally in French, then translated to English by a dedicated pipeline once available. This dataset is the English mirror of uneIAparjour/base, the original French dataset.… See the full description on the dataset page: https://huggingface.co/datasets/uneiaparjour/base-en.texttext-classification1K<n<10K0 likes285 downloads3h agoHugging Face26ML-Owl /faang-engineered-time-series-features-2013-2025 FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025) Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!) DOCUMENT NAVIGATION GUIDE (ToC) 1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset 3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.imagetabular-classification10K<n<100K2 likes283 downloads7mo agoHugging Face27blo05 /cleaned_wiki_en_20-40text1M<n<10M1 likes275 downloads4y agoHugging Face28enlatics /Enlatics_benchmarking GAIA-style Evaluation Results (Public) This dataset contains GAIA-inspired benchmark question results for LLM evaluation. What is inside grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status. Notes These tasks are designed in a GAIA-style (multi-hop, web-grounded questions). Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.textn<1K0 likes270 downloads8mo agoHugging Face29injilashah /eng_kash_sentence_pairtext100K<n<1M0 likes263 downloads2y agoHugging Face30Roudranil /shakespearean-and-modern-english-conversational-dataset Dataset Card for Shakespearean and Modern English Conversational Dataset Dataset Summary This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details. text1K<n<10K4 likes248 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.