datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Algorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code
This dataset provides different algorithms and their corresponding source code in Python.
credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face.
search-source-audit
Sources of Truth — AI Search Citations for Mental Health Queries
Which external sources do consumer AI search products actually cite when people ask about mental
health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform,
Multilingual Audit of Citations in AI Mental Health Information Queries."
Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.PoliticalBias_Sources
Dataset Card for PoliticalBias_Sources
Dataset Description
908 rows of data containing source name of an article, the source bias and the type of source
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of three columns namely Source Name, Source Bias and Source Typ
Source Data
The dataset is scrapped from https://www.allsides.com/media-bias
im3_open_source_data_center_atlas_v2026.02.09
IM3 Open Source Data Center Atlas v2026.02.09 — refined database
This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and
adds a source-enriched, audited 43-column power-source table for all 1,479
source geometry records (1,474 unique IM3 IDs). The publication retains the exact
13 upstream columns plus 30 stable label, interpretation, and evidence fields.
Duplicate geometry records are intentionally retained.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.python-algorithm-sourcecode
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset provides algorithms and corresponding Python source code which can be leveraged for any type of code conversion applications.
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/annawleo/python-algorithm-sourcecode.chinese-official-source-reachability
Reachability of Chinese official verification portals from outside China
If you are doing due diligence on a Chinese supplier, the advice is always "check the official registry". This dataset measures whether you can.
Eight official Chinese verification sources were measured from public vantage points outside mainland China across several rounds between August and September 2026, plus two controls (www.gov.cn and www.baidu.com).
The result is not "Chinese government sites are… See the full description on the dataset page: https://huggingface.co/datasets/derrick459/chinese-official-source-reachability.fetch_huggingface_terminal_9123_89cwuq_source_reviews
Product Reviews Corpus
Raw product review text collected from public e-commerce listings for customer experience research.
Dataset ID: SRC-REV-101
Catalog: NovaInsights
Records: 12,480
Origin: Public web scrape of product review pages (2023-2024)
License
This dataset is released under the MIT License. See https://opensource.org/licenses/MIT for the full license text.
Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.github_fetch_huggingface_terminal_9091_n3v8x2_source_alpha
Alpha Commerce Reviews
Raw review text collected from public e-commerce listings.
Dataset ID: SRC-ALPHA
Catalog: ghfht9091n3v8x2
Origin: Public web scrape of product review pages (2023-2024)
Records: 12,480
License: MIT
water-conflict-source-data
Water Conflict Training Dataset
This dataset contains labeled examples for training a multi-label water conflict classifier.
Dataset Structure
Files
positives.csv: Water conflict headlines with labels (Trigger, Casualty, Weapon)
negatives.csv: Non-conflict news headlines (pre-balanced with hard negatives)
Data Format
Both files have consistent schema:
Column
Description
Headline
News headline text
Basis
For positives: comma-separated… See the full description on the dataset page: https://huggingface.co/datasets/baobabtech/water-conflict-source-data.ccass_rapprochements_sommaires_with_sources
Training dataset for semantic textual similarity of French Court of cassation summaries
This CSV dataset is designed to train similarity models using pairs of court decision summaries. The dataset includes similar legal cases with comparable contextual frameworks and principled legal solutions derived from consistent judicial standards. The cases are categorized by the following roles:
Comparative case analyses
Judicial decision reversals
Each case entry includes a link to the… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/ccass_rapprochements_sommaires_with_sources.github_fetch_huggingface_terminal_9134_x9v2m6_source_beta
Beta Support Conversations
Anonymized customer support conversation transcripts.
Overview
Dataset ID: SRC-BETA
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Author: Nova Data Engineering
Origin: Community tech-support forum public dump (2022-2024)
Records: 18,200
Product Line: Customer Feedback Analytics
License
Apache License 2.0
Contents
Anonymized support conversation transcripts with timestamps.
sourcegithub_fetch_huggingface_terminal_9134_x9v2m6_source_gamma
Gamma Feedback Survey
Open survey responses from product feedback forms.
Overview
Dataset ID: SRC-GAMMA
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Author: Nova Market Research
Origin: Publicly distributed open survey responses (2023)
Records: 3,940
Product Line: Customer Feedback Analytics
License
Creative Commons Attribution 4.0 International (CC BY 4.0)
Contents
Open-ended survey responses and satisfaction ratings.
medical-google-chatgpt-gemini-source-overlap
Google and AI Source Overlap Across 12 Medical Niches
An open, reproducible US dataset comparing explicit ChatGPT and Gemini citations with paired Google organic Top 20 results across 12 medical niches and 432 frozen questions.
Full study: https://rotgar.com/medical/resources/google-top-20-chatgpt-gemini-source-overlap
Version DOI: https://doi.org/10.5281/zenodo.21850734
Version: 1.0
Fieldwork: August 7, 2026
Publication date: August 8, 2026
Market and language: United States… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/medical-google-chatgpt-gemini-source-overlap.coinbase_open_pricesquantem-base-model-sources
QuantEM — base model data sources
Every dataset in the corpus the QuantEM ViT-B encoder was pretrained on: public repository holdings, data contributed by external laboratories through the QuantEM outreach campaign, and in-house acquisitions.
Emitted verbatim from Supplementary Table 2 of the QuantEM manuscript — 657 rows. Please cite the
original sources listed here alongside QuantEM; rows carry a DOI or repository URL where one
exists.
Related:
ArrojoeDrigoLab/quantem — the… See the full description on the dataset page: https://huggingface.co/datasets/ArrojoeDrigoLab/quantem-base-model-sources.news-source-headlines-foxnews-nbc
News Source Headlines: FoxNews vs NBC
This dataset contains scraped news headlines from Fox News and NBC News for a binary news source classification project.
Files
expanded_headlines.csv: cleaned expanded dataset used as the main training dataset.
large_headlines.csv: larger scraped dataset used as an augmentation candidate pool.
Columns
url: original article URL
domain: article domain extracted from the URL
source: source label, either FoxNews or NBC… See the full description on the dataset page: https://huggingface.co/datasets/jessicajyzy/news-source-headlines-foxnews-nbc.github_fetch_huggingface_terminal_9134_x9v2m6_source_delta
Delta Metadata Registry
Registry metadata snapshot used internally.
Overview
Dataset ID: SRC-DELTA
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Author: Nova Product Analytics
Origin: Internal registry snapshot, non-commercial license
Records: 5,610
Product Line: Customer Feedback Analytics
License
Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
Contents
Registry metadata snapshot.
Vector_Database_With_Open-SourceAryl-Halides-Source-DataThis is the full data set for the aryl halide benchmark.
aviation-pilot-vehicle-decoherence-source-attribution-v0.1What this dataset tests
Whether a system can correctly identifythe source of pilot–vehicle loop decoherence.
Sources may be:
pilotaircraftenvironmentmixednone
Key insightCorrect attribution determinesthe correct recovery action.
Required outputs
primary_decoherence_source
source_confidence
resonance_pattern_type
escalation_likelihood
contributing_factors
attribution_rationale
Use case
Layer two of Pilot–Vehicle Loop Coherence Under Stress.Feeds adaptive intervention and crew… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-pilot-vehicle-decoherence-source-attribution-v0.1.Non-Conventional-Energy-Sources-Optimizer
Non-Conventional Energy Sources Optimizer
This dataset contains 100,000 records aimed at optimizing the efficiency of various non-conventional (renewable) energy sources such as solar and wind. The data is generated to support machine learning and deep learning models that can predict energy output and recommend configurations for maximum efficiency.
Dataset Details
Records: 100,000
Format: CSV
Use Cases: Regression models, energy output prediction, efficiency… See the full description on the dataset page: https://huggingface.co/datasets/SivaMallikarjun/Non-Conventional-Energy-Sources-Optimizer.VLMEvalKit_sourced_DocVQA_valcis5190-news-source-headlines
Fox News vs NBC News Headlines
Dataset collected for the CIS 4190 / 5190 Applied Machine Learning final
project (Spring 2026, Project B: News Source Classification).
Summary
Task: Binary text classification — predict the source (FoxNews
vs NBC) of an English news headline.
Rows: ~3,805 unique headlines (after deduplication 3,783).
Columns:
url — original article URL
headline — the headline text
source — one of FoxNews or NBC
Class balance: roughly 53% FoxNews / 47%… See the full description on the dataset page: https://huggingface.co/datasets/Ningze123/cis5190-news-source-headlines.quantem-organelle-model-sources
QuantEM — organelle model data sources
Annotated ground-truth sources behind the eight released organelle segmentation models (mitochondria, ER, nucleus, lipid droplet), with train/validation/test tile counts and annotated-crop counts per source.
Emitted verbatim from Supplementary Table 3 of the QuantEM manuscript — 45 rows. Please cite the
original sources listed here alongside QuantEM; rows carry a DOI or repository URL where one
exists.
Related:
ArrojoeDrigoLab/quantem —… See the full description on the dataset page: https://huggingface.co/datasets/ArrojoeDrigoLab/quantem-organelle-model-sources.united-states-covid-19-county-level-data-sources
United States COVID-19 County Level Data Sources - ARCHIVED
Description
The Public Health Emergency (PHE) declaration for COVID-19 expired on May 11, 2023. As a result, the Aggregate Case and Death Surveillance System will be discontinued. Although these data will continue to be publicly available, this dataset will no longer be updated.
On October 20, 2022, CDC began retrieving aggregate case and death data from jurisdictional and state partners weekly instead of daily.… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/united-states-covid-19-county-level-data-sources.dagaare_synth_source_targetThe dagaareDictTrain.tsv was generated using the Machine Translation from One Book (MTOB) technique using "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021) for LLM context.
Hartmann-4-Source-DataThis file contains some data points that can be used for the 4-dimensional transfer learning task.
Coffee-Price-Source-Websites
