datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
a-large-scale-fish-dataset
[!Note]
This is a copy version from A large scale fish dataset in Kaggle. And if you download this, you should follow the same license as original(CC-BY-NC4.0) and cite it as original readme says.
Original Dataset Authors : O. Ulucan, D. Karakaya, M. Turkan
For eaiser use, the dataset has been formated to 5 columns : image_id, image, mask, class_id, class_name, which is easier for you to load in hugging face and use.
A Large-Scale Dataset for Segmentation and Classification
Authors: O.… See the full description on the dataset page: https://huggingface.co/datasets/FriedParrot/a-large-scale-fish-dataset.opc-sft-stage1-largescale_diverse_instructlarge-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.large-scale-multimodal-multilingual-summarization-datasetPlease cite this paper if you use our code or data:
@inproceedings{verma-etal-2023-large,
title = "Large Scale Multi-Lingual Multi-Modal Summarization Dataset",
author = "Verma, Yash and
Jangra, Anubhav and
Verma, Raghvendra and
Saha, Sriparna",
booktitle = "Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics",
month = may,
year = "2023",
address = "Dubrovnik, Croatia",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/Zenquiorra/large-scale-multimodal-multilingual-summarization-dataset.EmergencyTrafficDetection_Large-Scale-Audio-datasetlarge-scale-log-analytics-hdfs-datalarge-scale-gender-biases-analysisasia-owid-cumulative-number-of-large-scale-ai-systems-by-country
Cumulative Number Of Large Scale Ai Systems By Country | Asia (Our World in Data)
🌏 49 observations · 7 Asia countries · 2019–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 49 observations of Cumulative Number Of Large Scale Ai Systems By Country data across 7 Asia countries, spanning 2019–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Cumulative Number Of Large… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-cumulative-number-of-large-scale-ai-systems-by-country.large-scale-hate-speech-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1:
The original dataset that includes 100,000 tweets in English. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v1.Large-Scale_Multilingual_Disambiguation_Glosses
[!NOTE]
Dataset origin: http://lrec2016.lrec-conf.org/en/shared-lrs/
Description
A multilingual large-scale corpus of automatically disambiguated glosses drawn from different resources integrated in BabelNet (such as Wikipedia, Wiktionary, WordNet, OmegaWiki and Wikidata). Sense annotations for both concepts and named entities are provided. In total, over 40 millions definitions have been disambiguated for 264 languages.
Citation
@InProceedings{CAMACHOCOLLADOS16.629… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Large-Scale_Multilingual_Disambiguation_Glosses.rubber_duck_dataset_smoothed_10fps_256x256_LARGE_scaled_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "UR5e",
"total_episodes": 25,
"total_frames": 1200,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nsimonato25/rubber_duck_dataset_smoothed_10fps_256x256_LARGE_scaled_augmented.MycoBase-Large-Scale-Text-to-SQL
MycoBase: A Biologically Literate Text-to-SQL Dataset
MycoBase is a synthetic but biologically accurate dataset designed for stress-testing Text-to-SQL systems. It represents a research information system for the study of fungi, covering everything from taxonomy and genomics to morphology and cultivation.
Dataset Highlights
Schema Complexity: 2,016 tables with over 9,000 foreign key relationships.
Data Volume: 320,270 rows of realistic mycology data.
Realistic Names:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/MycoBase-Large-Scale-Text-to-SQL.large-scale-hate-speech-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2:
The modified dataset that includes 68,597 tweets in English. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v2.average-income-of-large-scale-food-producers-ppp-for-african-countries
Average Income of Large Scale Food Producers Ppp for African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/average-income-of-large-scale-food-producers-ppp-for-african-countries.large-scale-hate-speechThis repository contains the utilized dataset in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer". This study mainly focuses hate speech detection in Turkish and English. In addition, domain transfer success between hate domains is also examined.
There are two dataset versions.
Dataset v1: The original dataset that includes 100,000 tweets per English and Turkish, published in LREC 2022. The annotations with more than 60% agreement are included.
Dataset v2: A… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech.productivity-of-large-scale-food-producers-for-african-countries
Productivity of Large Scale Food Producers for African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/productivity-of-large-scale-food-producers-for-african-countries.EzHit-large-scale-screeningCampusDepth-A-Large-Scale-Day-Night-RGB-Dataset-for-Monocular-Depth-Estimation
Advanced Driver Assistance Systems
(ADAS) require accurate and reliable perception of the
surrounding environment to ensure vehicle safety and
reduce the risk of collisions. Depth estimation plays a
crucial role in understanding object distance and spatial
relationships in traffic scenes. Traditional depth sensing
approaches, such as stereo camera systems and LiDAR,
provide accurate depth information but suffer from high
cost, increased hardware complexity, calibration… See the full description on the dataset page: https://huggingface.co/datasets/Vaibhav14/CampusDepth-A-Large-Scale-Day-Night-RGB-Dataset-for-Monocular-Depth-Estimation.Large_Scale_Midjourney_DatasetAbout Dataset
Begin on your image generation adventure with MidjourneyDataset
This dataset contains tens of millions of Midjourney prompt and image datasets to training and fine-tune your image generation models!
This dataset offers 12 million (and counting) Midjourney prompt+image datasets for your training needs, categorized by model version, style, composition and operations, serve as vital aid for boosting machine learning prowess in text generation. We will deliver the dataset through… See the full description on the dataset page: https://huggingface.co/datasets/MidjourneyDataset/Large_Scale_Midjourney_Dataset.FECA_ECOLI_Tsuboyama_2023_2D1U_substitutions_singles_stability_PE_REGRasia-owid-income-large-scale-food-producers
Income Large Scale Food Producers | Asia (Our World in Data)
🌏 28 observations · 13 Asia countries · 2005–2021 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 28 observations of Income Large Scale Food Producers data across 13 Asia countries, spanning 2005–2021.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Income Large Scale Food Producers
Geographic coverage… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-income-large-scale-food-producers.large_scale_apksafrica-owid-income-large-scale-food-producers
Income Large Scale Food Producers | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-income-large-scale-food-producers.large_scale_txt_processing_pipelineRL20_AQUAE_Tsuboyama_2023_1GYZ_substitutions_singles_stability_PE_REGReurope-unsdg-productivity-of-large-scale-food-producers-agricultural-pd-agr-lsfpeurope-owid-cumulative-number-of-large-scale-ai-systems-by-country
Cumulative Number Of Large Scale Ai Systems By Country | Europe (Our World in Data)
🇪🇺 63 observations · 9 Europe countries · 2019–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 63 observations of Cumulative Number Of Large Scale Ai Systems By Country data across 9 Europe countries, spanning 2019–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Cumulative Number Of… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-cumulative-number-of-large-scale-ai-systems-by-country.tudobonus-largescale
