datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.modis-lake-powell-toy-dataset
MODIS Water Lake Powell Toy Dataset
Dataset Summary
Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.abidhussai512_tesla-byd-toyota-stock-prices-and-volume-20182026
Tesla, BYD, Toyota Stock Prices & Volume 2018–2026
Tesla, BYD, Toyota Stock Data 2018–2026 with Daily Prices & Trading Volume
Dataset Info
Source: Kaggle
Original Size: 0.18 MB
Kaggle Downloads: 82
Files: 1
Files
auto_company_comparison.csv
Mirrored from Kaggle
PromptCloudHQ_toy-products-on-amazon
Toy Products on Amazon
10,000 toy products on Amazon.com
Dataset Info
Source: Kaggle
Original Size: 8.18 MB
Kaggle Downloads: 13,213
Files: 1
Files
amazon_co-ecommerce_sample.csv
Mirrored from Kaggle
bridgev2-vita-toykitchen-manifests
BridgeV2 VITA ToyKitchen-like Manifests
This repository contains manifest files for a reconstructed VITA-style BridgeV2 ToyKitchen-like pick-and-place subset.
Source dataset
The source dataset is:
Gaugou/BridgeV2
This repository does not duplicate the original BridgeV2 videos. It provides episode IDs and metadata for selecting the subset from the source dataset.
Split
Train: 2,986 episodes
Test: 287 episodes
Total selected: 3,273 episodes
Selection… See the full description on the dataset page: https://huggingface.co/datasets/praedico/bridgev2-vita-toykitchen-manifests.ToyotaMotorsStockData
Toyota Motors Stock Data (1980-2024)
This is a dataset copied from Kaggle. You can see the original dataset here: https://www.kaggle.com/datasets/mhassansaboor/toyota-motors-stock-data-2980-2024
The following is the original readme of this dataset:
About Dataset
📊 Toyota Stock Dataset (1980-2024)
🌟 This dataset offers daily stock trading data for Toyota Motor Corporation (ticker: TM) spanning from 1980 to 2024, sourced from Yahoo Finance. It provides an… See the full description on the dataset page: https://huggingface.co/datasets/tablegpt/ToyotaMotorsStockData.toy_ds1toy-diabetesOSWorldFilestoytoyota-paint-attributesViSoIE-toys
SoIEA Toy Dataset
Toy dataset for testing and developing the "Social Implicit Emotion Analysis" (SoIEA) pipeline.
📄 License
This project employs a dual-license strategy to ensure open collaboration while protecting the integrity of the research data.
1. Dataset
The data files (including *.json, *.csv, and raw content) are licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).
You are free to:
Share — copy and… See the full description on the dataset page: https://huggingface.co/datasets/HiAmNear/ViSoIE-toys.modis-lake-powell-toy-dataset
MODIS Water Lake Powell Toy Dataset
Dataset Summary
Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/wateryhcho/modis-lake-powell-toy-dataset.toysentimentToyota_radiatorrobot-nav-toy-dataset
Humanoid Robot Navigation Dataset
Overview
This dataset contains state–action pairs created for humanoid robot navigation tasks.
It is designed for simple robotics learning scenarios such as:
Behavior cloning
Reinforcement learning
Policy learning
Navigation control experiments
The dataset maps robot state information (position and orientation) to discrete navigation actions.
Dataset Structure
File:
train.csv
Each row represents one timestep observation and… See the full description on the dataset page: https://huggingface.co/datasets/bengusu80/robot-nav-toy-dataset.biolama_umls_mcq_toytoy_wmt24_mqmtoyToy_Dataset
AMAMMERε Overview
The AMAMMERε Dataset delves into the subtle cultural biases present in common sense reasoning datasets, highlighting predominantly Western or U.S.-centric perspectives. By comparing cultural commonsense reasoning between Ghana (representing a low-resource language group) and the USA, this dataset shows the differences in daily activities such as shopping, meal planning, and transportation across these distinct cultural contexts.
Highlights:
525 Questions: Curated… See the full description on the dataset page: https://huggingface.co/datasets/Christabel/Toy_Dataset.toy_dataml_course_toy_dataset_housing_pricetoy-llm-eval-dataset
