datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pending-medicare-provider-enrollment-data
Pending Medicare Provider Enrollment Data
This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication.
Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.enabling-independent-research
Overview
This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude".
Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/enabling-independent-research.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.wikipedia_en
wikipedia_en
This is a curated Wikipedia English dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.TCGA_OncoTree_pt2
TCGA_OncoTree_pt2
1. Tổng quan
[CẦN ĐIỀN THỦ CÔNG: mục đích, ngữ cảnh tạo dataset]
Tổng số bản ghi (cộng tất cả manifest phát hiện được): 23984
Số manifest phát hiện được trong bộ nhớ: 3 (df, labels_df, progress)
Repo HuggingFace chính: ento3686/TCGA_OncoTree_pt2
⚠️ Dataset được lưu trên 2 repo/tài khoản HuggingFace khác nhau:
ento3686/TCGA_OncoTree_pt2 (biến: REPO_ID_2, UPLOAD_REPO_ID, CENTRAL_PROGRESS_REPO_ID, _repo_id_var)
tuna2004/TCGA_OncoTree (biến:… See the full description on the dataset page: https://huggingface.co/datasets/ento3686/TCGA_OncoTree_pt2.Environment-and-Natural-Resources-Indicators-For-African-Countries
Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.ted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.EnviroExam
Dataset Summary
EnviroExam focuses on 42 core courses from the environmental science curriculum at Harbin Institute of Technology, after excluding general, duplicate, and practical courses from a total of 141 courses across undergraduate, master's, and doctoral programs.
For these 42 courses, initial draft questions were generated using GPT-4 and Claude, combined with customized prompts. These drafts were then refined and proofread manually, resulting in a total of 1,290… See the full description on the dataset page: https://huggingface.co/datasets/enviroscientist/EnviroExam.EnglishDictionaryEnergy-Indicators-For-African-Countries
Energy Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Energy-Indicators-For-African-Countries.cleaned_wiki_en_60-80energy-consumption-hourly-spainspotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.cleaned_wiki_en_40-60wordnet-definitions-en-2021
Wordnet definitions for English
Dataset by Princeton WordNet and the Open English WordNet team
https://github.com/globalwordnet/english-wordnet
This dataset contains every entry in wordnet that has a definition and an example.
Be aware that the word "null" can be misinterpreted as a null value if loading it in with e.g. pandas
cleaned_wiki_en_80-100github-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table.
It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and
totals of the amounts of various events on those repositories in that time period.
This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing,
so I figured I would save the next person to try this some effort.
enfermedades-wiki-marzo-2024
English Version
This dataset contains detailed information on a total of 945 diseases, extracted from Wikipedia in Spanish (https://es.wikipedia.org/) in March 2024. The main purpose of this dataset is to serve as a comprehensive resource for training Large Language Models (LLMs) in Spanish, specifically for instruction tuning, pre-training, and other natural language processing (NLP) tasks. This dataset promises to be a valuable tool for research and development in Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/Alvaro8gb/enfermedades-wiki-marzo-2024.cleaned_wiki_enCleaned wikipedia dataset
africa-synth-energy-oilgas-gas-infrastructure-nigeria
Africa Synth Energy Oilgas Gas Infrastructure Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-infrastructure-nigeria.dbpedia-entity-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/dbpedia-entity-qrels.base-en
Base uneIAparjour.fr — English — Generative AI Apps
Open dataset listing one generative AI tool per day since February 16, 2023, featured on uneiaparjour.fr and translated to English.
Description
Every day, a new free or freemium generative AI tool is tested, described and categorized — originally in French, then translated to English by a dedicated pipeline once available. This dataset is the English mirror of uneIAparjour/base, the original French dataset.… See the full description on the dataset page: https://huggingface.co/datasets/uneiaparjour/base-en.faang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.cleaned_wiki_en_20-40Enlatics_benchmarking
GAIA-style Evaluation Results (Public)
This dataset contains GAIA-inspired benchmark question results for LLM evaluation.
What is inside
grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status.
Notes
These tasks are designed in a GAIA-style (multi-hop, web-grounded questions).
Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.eng_kash_sentence_pairshakespearean-and-modern-english-conversational-dataset
Dataset Card for Shakespearean and Modern English Conversational Dataset
Dataset Summary
This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details.
