datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.danbooru-tags-classified
danbooru-tags-classified
Danbooru tags split into categories, one CSV per category. Each CSV is tag,count
sorted by count descending, where count is the tag's post frequency on Danbooru.
file
tags
contents
artist.csv
83355
artist names
character.csv
57653
character names
series.csv
12405
copyright / series names
other.csv
15942
not yet assigned to a category
attire.csv
9646
clothing and worn items
object.csv
4227
objects
feature.csv
2930
body and… See the full description on the dataset page: https://huggingface.co/datasets/Jio7/danbooru-tags-classified.arxivannotations
Title
Annotation
PDF
Latex
Axion bremsstrahlung from collisions of global strings
We calculate axion radiation emitted in the collision of two straight globalstrings. The strings are supposed to be in the unexcited ground state, to beinclined with respect to each other, and to move in parallel planes. Radiationarises when the point of minimal separation between the strings moves fasterthan light. This effect exhibits a typical Cerenkov nature. Surprisingly, itallows an alternative… See the full description on the dataset page: https://huggingface.co/datasets/Dan-Kos/arxivannotations.futuresssterminal-tasksGHG-Emissions-Data
GHG Emissions Data Pipeline
Description
This repository contains a comprehensive pipeline for processing and analyzing greenhouse gas (GHG) emissions data. The pipeline integrates datasets from multiple sources, including Climate TRACE and Our World in Data, to provide insights into global emissions trends. It supports sustainability reporting, emissions tracking, and climate action planning.
Dataset Details
Sources and Methodologies
The pipeline… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/GHG-Emissions-Data.Danbooru-Top1000-Latents-NPZDanbooru-Top1000-Latents-SDXLdanbooru-tags-data-zh
danbooru-tags-data-zh
Danbooru 标签的中文翻译数据库,收录截至 2026 年 8 月图片数量大于 50 的标签,译名以社区通用叫法为准。
同一份数据同时发布在 GitHub 仓库 与 HuggingFace 数据集,推送 GitHub 后由 Actions 自动同步。
覆盖范围
收录 2026 年 8 月时图片数量大于 50 的标签
按分类收录:画师(artist)、作品/版权(copyright)、角色(character)、通用(general)、元标签(meta)
当前收录情况:
分类
标签数
画师 artist
24881
作品/版权 copyright
8413
角色 character
35382
通用 general
30664
元标签 meta
585
特点
现有的同类数据多为较早期抓取、之后未再更新,且通常只提供单一译名,不含别名,也没有对标签含义的说明。本项目在以下方面做了补充:… See the full description on the dataset page: https://huggingface.co/datasets/ame-la/danbooru-tags-data-zh.Danbooru-Dataset-csv
Danbooru Dataset CSV
面向 Danbooru 标签管理 / 打标工具的公开元数据合集。这里只放整理后的 CSV,不含任何图片。后续还会继续补充 artist、copyright 等更多表;本页只做项目总览,各文件以仓库里的 CSV 为准。
标签与 wiki 来自 Danbooru。本仓库整理表使用 MIT 协议。原图版权仍归各自作者。
当前文件
文件
内容
截止日期
行数
danbooru_dataset_general_260820.csv
general 通用标签(别名、层级、父子、分类、wiki)
2026-08-20
106,414
danbooru_character_tags.csv
character 角色标签(别名、作品、父标签、投稿数)
2026-07-20
329,747
danbooru_artist_tags.csv
artist 画师标签(译名、数据量)
—
576,842
tag-near-synonym-relations4.csv… See the full description on the dataset page: https://huggingface.co/datasets/StoryAura/Danbooru-Dataset-csv.cross_species_benchmarking**Repository: https://d-script.readthedocs.io/en/stable/data.html
**Reference: Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Systems 12, 969-982.e6 (2021).
storage-container-dimensions
Industrial Storage Container Dimensions
Planning-grade volumetric reference for the plastic containers that European and
North American warehouses actually run on: Euroboxes (Euro stacking containers),
attached-lid containers (ALCs), and VDA 4500 KLTs.
Two tables:
Config
Rows
What it is
default → containers.csv
48
One typical row per nominal size. External and internal dimensions, usable capacity, the spread of what vendors publish, lid and nesting behaviour… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/storage-container-dimensions.LLMs-Sentiment-Augmented-Bitcoin-Dataset
Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities
The work was carried out by:
Danilo Corsi
Cesare Campagnano
Description
This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media.
We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.cheboksary_transport
ChebTransport Daily GPS Dataset
This dataset contains daily GPS and metadata records of public transport vehicles in Cheboksary, Russia, for the period from 2025-04-22 to 2025-07-03.
Each file corresponds to a "transport day" (which may start and end at different times depending on the actual end of public transport service, not at midnight).
Data Source
The data was parsed from the website buscheb.ru, which aggregates public transport data for the city of Cheboksary as a… See the full description on the dataset page: https://huggingface.co/datasets/daniilakk/cheboksary_transport.multilingual-islr-mediapipe
Multilingual ISLR MediaPipe Landmarks
Dataset Description
This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning.
The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.IEEE-118_overloadTODO
Bernett_benchmarking
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Dataset Sources [optional]
**Repository: https://doi.org/10.6084/m9.figshare.21591618.v3
**Reference:
Bernett, J., Blumenthal, D. B. & List, M. Cracking the black box of deep sequence-based protein–protein interaction prediction. Briefings in Bioinformatics 25, bbae076… See the full description on the dataset page: https://huggingface.co/datasets/danliu1226/Bernett_benchmarking.danbooru-tags-classified
danbooru-tags-classified
Danbooru tags split into categories, one CSV per category. Each CSV is tag,count
sorted by count descending, where count is the tag's post frequency on Danbooru.
file
tags
contents
artist.csv
83355
artist names
character.csv
57653
character names
series.csv
12405
copyright / series names
other.csv
15942
not yet assigned to a category
attire.csv
9646
clothing and worn items
object.csv
4227
objects
feature.csv
2930
body and… See the full description on the dataset page: https://huggingface.co/datasets/LEEEFA/danbooru-tags-classified.flipfinder-usa
🎥 Project Walkthrough Video
🏠 FlipFinder USA
Identifying Undervalued Real Estate Investment Opportunities Across the United States
Author: Dan | HuggingFace: @dant555
📋 Project Overview
This project transforms a general-purpose Zillow real estate dataset into a focused investment screening tool. Using Exploratory Data Analysis (EDA), I engineered a binary target variable (is_good_flip) to identify properties that are genuinely… See the full description on the dataset page: https://huggingface.co/datasets/dant555/flipfinder-usa.Nanobody_Sequence_DatasetRepresentative sequence dataset extracted after clustering of Integrated NANOBODY® Database for Immunoinformatics (INDI) by MMseqs2 program.
For more information, please visit https://github.com/DynaX-C/EvoNB.
Mutation_effect_dataset**Repository: https://ftp.ebi.ac.uk/pub/databases/intact/current/various/mutations.tsv
**Reference: Kerrien, S. et al. The IntAct molecular interaction database in 2012. Nucleic Acids Research 40, D841–D846 (2012).
sentiment-analysis-catalan-reviews
CSXSC: Classificador de Sentiments de Xarxes Socials en Català
This repository contains the CSXSC (Classificador de Sentiments a Xarxes Socials en Català) dataset, a comprehensive corpus designed for sentiment analysis of Catalan-language content from social media.
The dataset contains 23,788 text entries, each classified as positive, negative, or neutral. It was specifically constructed to address the significant class imbalance often found in user-generated content, resulting in a… See the full description on the dataset page: https://huggingface.co/datasets/Danie1Arias/sentiment-analysis-catalan-reviews.Barcenas-HumorNegroDataset en español con 500 chistes de humor negro y una explicación.
Datos creados de manera sintética por Claude 3 Haiku y Llama 3 70B Instruct.
El proceso para crear el dataset fue el recopilar de varias fuentes chistes de humor negro en español para luego ser utilizadas en los mejores modelos como Gemini 1.5 Pro, Claude 3, etc.
Con eso genere cientos de chistes de humor negro en español para tener más datos y hacer un super recopilatorio de chistes de humor negro en español, aproximadamente… See the full description on the dataset page: https://huggingface.co/datasets/Danielbrdz/Barcenas-HumorNegro.elizaThis repository contains synthetic ELIZA chatbot conversations.
See https://github.com/princeton-nlp/ELIZA-Transformer for more details.
150k-anime-danbooru-fluxIEEE_14_datasetSTRING_V12_TrainingSet**Repository: https://stringdb-downloads.org/download/protein.physical.links.v12.0.txt.gz
**Reference: Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research 51, D638–D646 (2023).
mquad-v1
MQuAD
The Medical Question and Answering dataset(MQuAD) has been refined, including the following datasets. You can download it through the Hugging Face dataset. Use the DATASETS method as follows.
Quick Guide
from datasets import load_dataset
dataset = load_dataset("danielpark/MQuAD-v1")
Medical Q/A datasets gathered from the following websites.
eHealth Forum
iCliniq
Question Doctors
WebMD
Data was gathered at the 5th of May 2017.
The MQuAD provides embedded question… See the full description on the dataset page: https://huggingface.co/datasets/danielpark/mquad-v1.anima-earned-datascale
anima-earned-datascale — H_9968 natural-corpus operator supply vs DATA scale
The top-rung fixed-length in-band corpus for anima hypothesis H_9968: does the natural-corpus
supply of a transferable recombination operator (negation flips sentiment polarity independent of the
stem) grow with data scale, with sentence length held fixed?
This is the p9 "one unopened cell" screener corpus, measured by the certified anima-py evaluate --earned
instrument (corpus-statistics, never touches… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-earned-datascale.addition
