datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fsrs-datasetNYC-Airbnb-Open-Dataau-racing-open-data
Australian Racing Open Data
Open, machine-readable datasets for Australian racing — the kind of data that
normally sits behind a login, a paywall, or nowhere at all.
Two of these datasets, as far as we can tell, have never existed publicly
before: track geometry (turn radii, cambers, straight lengths, first-split
distances — gathered by writing to 111 racing clubs and state bodies) and
greyhound GPS sectionals at 50-metre resolution.
Everything here is rebuilt and pushed every… See the full description on the dataset page: https://huggingface.co/datasets/brucem1967/au-racing-open-data.oceania-gov-open-data-catalog
Oceania Government Open Data — Combined Catalogue (hourly snapshot)
Combined regional catalogue of Oceania (Australia + New Zealand) public-service open
data harvested from both data.gov.au and data.govt.nz portals, including state,
territory and local-council publishers.
License declaration
License: other (see below). Records in this catalogue inherit the licence of their
source dataset. Where the source declares a standard open licence the record is tagged
with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.open-biohacking-data
Holistix Open Biohacking Data Project v1.5.0
The Holistix Open Biohacking Data Project is a public, website-first structured data system for wellness technologies, product specifications, safety information, evidence metadata, claim boundaries, and machine-readable product intelligence.
This project is separate from the private Holistix Swarm Engine. The Swarm Engine may assist with auditing and monitoring, but it is not the source of truth for the public data release.… See the full description on the dataset page: https://huggingface.co/datasets/HolistixOpen/open-biohacking-data.japan-procurement-open-data
Japan Public Procurement & Company Open Data
Machine-readable extracts of Japanese public-procurement and company open data, compiled and normalised by LoreaTec for japan-tenders.loreatec.jp and bizsearch.loreatec.jp. Everything here comes from official Japanese government sources; the value added is the cleaning, joining and the derived analysis (contract series and re-tender predictions).
Updated monthly. The authoritative, always-current copy is… See the full description on the dataset page: https://huggingface.co/datasets/loreatec/japan-procurement-open-data.korean_law_open_data_precedents
Dataset Card for Dataset Name
공지사항
인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다.
사용상 주의사항
사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다.
그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다.
사용에 참고하시기 바랍니다.
Dataset Summary
2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다.
그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.open-neo-public-test-data
Open-Neo HCC1395/HCC1395BL Public WGS Test Data
This dataset provides a public paired whole-genome test case for validating the
Open-Neo installation, input-QC, DNA evidence, purity/CNV, HLA, candidate
generation, presentation, ranking, and reporting paths.
Samples
Role
Sample
BioSample
Source material
Tumor
HCC1395 (WGS_EA_T_1)
SAMN10102573
Breast carcinoma cell line, ATCC CRL-2324
Matched normal
HCC1395BL (WGS_EA_N_1)
SAMN10102574
B-lymphoblast cell… See the full description on the dataset page: https://huggingface.co/datasets/open-neo/open-neo-public-test-data.im3_open_source_data_center_atlas_v2026.02.09
IM3 Open Source Data Center Atlas v2026.02.09 — refined database
This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and
adds a source-enriched, audited 43-column power-source table for all 1,479
source geometry records (1,474 unique IM3 IDs). The publication retains the exact
13 upstream columns plus 30 stable label, interpretation, and evidence fields.
Duplicate geometry records are intentionally retained.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.GLM-Open-Dialogue-Chinese-Dataset-v1OpenDataGen-factuality-en-v0.1This synthetic dataset was generated using the Open DataGen Python library. (https://github.com/thoddnn/open-datagen)
Methodology:
Retrieve random article content from the HuggingFace Wikipedia English dataset.
Construct a Chain of Thought (CoT) to generate a Multiple Choice Question (MCQ).
Utilize a Large Language Model (LLM) to score the results then filter it.
All these steps are prompted in the 'template.json' file located in the specified code folder.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/thoddnn/OpenDataGen-factuality-en-v0.1.atlas_opendata
ATLAS Open Data ROOT files
This dataset mirrors the skim-level layout of the ATLAS Open Data files used
by the hackathon. It contains ROOT files and the source metadata CSV only.
The top-level folders are the source skims:
1lepMET30
2J2LMET
2to4lep
4lep
GamGam/2020/MC
GamGam/2025/data
GamGam/2025/MC
Within each skim, data/ and MC/ folders are retained where they exist.
metadata.csv is the source-wide metadata table. ROOT files are stored with
Git LFS.
The diphoton files are… See the full description on the dataset page: https://huggingface.co/datasets/ho22joshua/atlas_opendata.whatsapp-open-data
WhatsApp Business Platform Open Data
Dated, sourced reference data for the WhatsApp Business Platform, published under
CC BY 4.0 by WhaTools (https://wha.tools). Each file carries its own licence,
attribution and verification date; the live source is regenerated from
https://wha.tools/api/data.
Datasets
File
What it is
Verified
Source page
rates.{json,csv}
WhatsApp Business API rate card
2026-08-03
https://wha.tools/whatsapp-api-pricing
limits.{json… See the full description on the dataset page: https://huggingface.co/datasets/WhaTools/whatsapp-open-data.futuresclock-open-data
FuturesClock Open Data
Versioned reference data for global futures contract specifications, expiry events, and trading hours. The files power the bilingual tools at FuturesClock.com and are published in both CSV and JSON formats for research, education, monitoring, and software integration.
Version 1.0.0
Dataset
Records
Last fact review
Files
Contract specifications
75
2026-08-10
contracts.csv, contracts.json
Expiry events
152
2026-08-10
expiries.csv… See the full description on the dataset page: https://huggingface.co/datasets/brucewangdatou/futuresclock-open-data.open-data-bielefeld
Open Data Bielefeld
Open Data Bielefeld ist eine unabhängige, kuratierte Sammlung öffentlich zugänglicher Datenquellen und Informationsangebote rund um Bielefeld.
Ziel ist es, relevante Quellen zu Themen wie Stadtentwicklung, Statistik, Geodaten, Mobilität, Verwaltung und Civic Tech übersichtlich auffindbar zu machen.
Die Sammlung kann künftig als Grundlage für Datenanalysen, Visualisierungen, KI-Anwendungen und lokale Agenten dienen.
Inhalte
Die Datei sources.csv… See the full description on the dataset page: https://huggingface.co/datasets/bielefeld/open-data-bielefeld.govsentry-open-data
GovSentry Open Datasets
Free, openly-licensed datasets for US government-contracting research, published by
GovSentry.
License: CC BY 4.0 — free to use, including commercially. Attribution required (see How to cite).
About / methodology: https://govsentry.ai/resources/open-datasets
Also on GitHub: https://github.com/marcusfkelley/govsentry-open-data
Use the config dropdown in the data viewer to switch between datasets.
federal-niche-winnability-2026 — flagship… See the full description on the dataset page: https://huggingface.co/datasets/Marcusfkelley/govsentry-open-data.GLM-Open-Dialogue-Chinese-Dataset-v2race-mx-open-data
Race.mx Mexican Motorsport Open Data
Open data on motorsport in Mexico, published by Race.mx, an independent English-language guide to racing in Mexico:
Winners history: every Formula 1 Mexican / Mexico City Grand Prix since 1962, Baja 1000 overall winners since 1967, Baja 500 winners since 2010, every Formula E Mexico City E-Prix, and NASCAR Mexico Series champions since 2004.
Race calendar: the current season's major races in Mexico (F1, SCORE desert racing, NASCAR Mexico… See the full description on the dataset page: https://huggingface.co/datasets/RaceMX/race-mx-open-data.GLM-Open-Dialogue-Chinese-Datasethate_speech_open_data_original_class_test_setVector_Database_With_Open-SourceOpen_Meteo_weather_dataset_from_2024_to_2025
license: mit
language: en
task_categories:
- tabular-regression
- environmental-science
- agriculture
Hungary Weather & Soil Dataset (2024–2025)
This dataset contains hourly meteorological and soil data from 8 locations across Hungary (Budapest, Debrecen, Szeged, Győr, Miskolc, Pécs, Nyíregyháza, Kecskemét) from January 1, 2024 to September 24, 2025, sourced via the Open-Meteo Historical Weather API.
🔍 Key Features
Soil temperature at 0cm, 6cm, 18cm, and… See the full description on the dataset page: https://huggingface.co/datasets/MWasil/Open_Meteo_weather_dataset_from_2024_to_2025.open-stock-reports-dataset
📊 Open Stock Reports Dataset
Quarterly Free Cash Flow (FCF) data for 3,000+ US stocks from 2019 to 2025, updated regularly.
100% open and free.
🏦 3,000+ public US companies
📅 2019–2025
🔁 Regular updates
This dataset was originally collected for a stock market statistical test for revenue to price correlation.
spain-real-estate-open-data
Spain Real Estate Open Data
Datos abiertos del mercado inmobiliario español por provincia, municipio, tipo de propiedad y operación. Source: la red federadora Eligemicasa.com — el portal que une a las inmobiliarias verificadas de toda España.
Resumen
Este dataset contiene un snapshot estructurado del mercado inmobiliario español:
52 provincias con métricas agregadas (volumen activo, precio medio/min/max).
8.131 municipios con población y volumen.
Precios medianos por… See the full description on the dataset page: https://huggingface.co/datasets/Eligemicasa/spain-real-estate-open-data.german-cities-open-data
InfraNode German Cities Open-Data Snapshot
Ein reproduzierbarer, offen lizenzierter Querschnitt von Infrastruktur- und
Umweltdaten für 84+ deutsche Städte, erzeugt aus der öffentlichen
InfraNode-API. Eine Zeile je Stadt.
Inhalt
Bereich
Felder
Quelle
Stammdaten
slug, name_de, state, ags, wikidata_qid, lat, lon, base_population, base_area_km2
Wikidata (CC0)
Wetter
weather_temperature_c, weather_humidity, weather_condition
DWD (GeoNutzV)
Luftqualität… See the full description on the dataset page: https://huggingface.co/datasets/Khaledc83/german-cities-open-data.Open-ended_Questions_dialectal_data
Dataset Summary
A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences.
For more details, please check the paper
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection
Citation Information
@inproceedings{yamani-etal-2024-kind,
title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.open_dataset_66778899000
什么都没有,看到这就退出吧!!
hinglish_open_hathi_datasethallym_AI_OpenDataset
Hallym Adult and Child Speech Dataset
This dataset contains speech recordings and transcriptions collected from adult and child speakers for AI-based speech and language research.
Dataset Overview
Total Records: 2,714
Speakers: 49 (adult: 25, child: 24)
Groups: adult, child
File Format: WAV (audio) + TXT (transcription)
Speaker Statistics
Group
Count
Gender
Age Range
Adult
25명
남/여
50~78세
Child
24명
남/여
3~8세
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_AI_OpenDataset.NYC-Airbnb-Open-Data
