datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STARK_10k
STARK: Spatial-Temporal reAsoning benchmaRK
STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Dataset Summary
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.LongVA-TPO-10k
10kTemporal Preference Optimization Dataset for LongVA
LongVA-TPO-10k, introduced by paper Temporal Preference Optimization for Long-form Video Understanding
europarl_for_language_detection_10kPlotPalette-10K
Plot Palette is a curated dataset designed for fine-tuning large language models (LLMs) on creative writing tasks. Sourced from various literary sources and generated using the Mistral 8x7B language model. The scripts used to generate the data can be found here.
Data Fields
'id': A unique identifier for each prompt-response pair.
'category': The category to which the prompt-response pair belongs (e.g., creative_writing, generation, poem, brainstorm, question_answer). --- (… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/PlotPalette-10K.spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
Crypto-Address-Annotation-10K
Codatta Crypto Address Annotations (Sample)
Overview
This dataset is a 10,000-row sample of the comprehensive Codatta Crypto Address Annotations database. The full database serves as a massive repository of over 500 million labeled address pairs across multiple blockchains.
The data provides critical metadata aimed at solving the problem of fragmented and siloed blockchain information. It includes entity names, functional categories (e.g., Exchanges, DeFi, Scam)… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Crypto-Address-Annotation-10K.VideoDPO-10k@misc{liu2024videodpoomnipreferencealignmentvideo,
title={VideoDPO: Omni-Preference Alignment for Video Diffusion Generation},
author={Runtao Liu and Haoyu Wu and Zheng Ziqiang and Chen Wei and Yingqing He and Renjie Pi and Qifeng Chen},
year={2024},
eprint={2412.14167},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2412.14167},
}
@misc{wang2024vidprommillionscalerealpromptgallery,
title={VidProM: A Million-scale Real… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/VideoDPO-10k.10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.atlas-crispr-10k-benchmark
🧬 ATLAS CRISPR 10k Benchmark
Contribution communauté LeWorldModel — Benchmark CRISPR 10k guides ARNUtilisé pour fine-tuner aguennoune17/negenWM-jepa-v2 — ATLAS NWM Sprint 3Self-supervised I-JEPA · Encodage téléologique (κ, τ, λ)
Description
10 000 guides ARN Cas9 de 20 nucléotides consolidés depuis 12 études expérimentales
de criblage CRISPR génomique à grande échelle. Ce dataset est le benchmark officiel du
Sprint 3 ATLAS NWM v2 — entraînement I-JEPA… See the full description on the dataset page: https://huggingface.co/datasets/aguennoune17/atlas-crispr-10k-benchmark.Student-Mental-Health-Counseling-10K
Student Mental Health Counseling 10K
Dataset Overview
This dataset is derived from the original chillies/student-mental-health-counseling-vn dataset, which contains student mental health counseling conversations in Vietnamese.
In this version, 10,000 randomly sampled rows from the original dataset have been translated into English using Google Translator, making the data more accessible for English-speaking researchers and developers.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/arafatanam/Student-Mental-Health-Counseling-10K.SEC-10Q-10K-Statement-tablesmagpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1Puzzles_10kSMS-dataset-sample-10klicense: mit
language:
en
📌 Free 10K SMS Preview Dataset
This is a preview subset of the full OTP + OTP INTENT + Phishing dataset.
👉 For full 73K dataset:
https://huggingface.co/datasets/gandharvbakshi/SMS-dataset-OTP-OTP_INTENT_Phishing
chess_spatial_reasoning_10kalzheimers-multi-target-10k-dataset
📚 BioDockify: Multi-Target Alzheimer's Chemical Space & Virtual Screening Dataset (10,000 Verified Compounds)
Principal Investigator: Tajuddin Shaik (tajo9128@gmail.com)Affiliation: Faculty of Pharmacy, Bharath Institute of Higher Education and Research (BIHER), Chennai, IndiaPlatform: www.biodockify.com | ai.biodockify.com
📌 Dataset Summary
This repository contains the complete 10,000 curated, literature-grounded chemical space dataset for Alzheimer's… See the full description on the dataset page: https://huggingface.co/datasets/BioDockify/alzheimers-multi-target-10k-dataset.synthetic-nsclc-10kMoltSafe-10K
MoltSafe-10K
MoltSafe-10K contains safety annotations for 10,000 posts and comments from Moltbook, an agentic social network
where autonomous agents communicate with one another. The underlying corpus is
AIcell/moltbook-data on the
Hugging Face Hub. Every node has a binary safety verdict, a severity level, a
malicious-intent taxonomy, and OWASP GenAI risk codes.
We release the annotations used in the development our study to promote research into Moltbook security. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/MoltSafe-10K.GameLabel-10kGameLabel-10k Dataset Card
This dataset contains was created in collaboration with the game developers of Armchair Commander. It contains 9800 human preferences over pairs of Flux-Schnell generated images, with over 6800 unique prompts. All labels were crowdsourced from Armchair Commander players.
Usage Example
from datasets import load_dataset
from PIL import Image
import base64
from io import BytesIO
dataset = load_dataset("Jonathan-Zhou/GameLabel-10k")
# For some reason, when using… See the full description on the dataset page: https://huggingface.co/datasets/Jonathan-Zhou/GameLabel-10k.SEC-10Q-10K-Coverpage10k_history_v510k_history_summarySummarized over 10,000 rows of historical information, mostly US History. Initally generated from historical text using gpt-4o-mini, summarized using Mistral-Nemo.
waimai_10kecommerce-fraud-detection-synthetic-10k-sampl
🛡️ Synthetic E-Commerce Fraud & AML Detection Dataset (10k Evaluation Sample)
⚠️ NOTICE: This is a truncated 10,000-row evaluation sample strictly for schema verification and local testing.
💳 [OBTAIN THE 10-MILLION ROW COMMERCIAL LICENSE HERE] > https://buy.stripe.com/8x26oIad4eH9eJf6gJ5wI01
🚀 Quick Start (Load via Hugging Face)
Data scientists can instantly load this evaluation slice into their Pandas/Python environment using the datasets library:
from… See the full description on the dataset page: https://huggingface.co/datasets/chinna887/ecommerce-fraud-detection-synthetic-10k-sampl.10k_reports_gemma_v2
Dataset Card for Financial Document Analysis Dataset
Dataset Description
This dataset comprises structured conversational entries designed to facilitate the training and evaluation of models that analyze and summarize financial documents. Each entry includes a conversation ID, a specific step in the conversation, a system-generated prompt, a user question, and the corresponding model-generated response.
Fields Overview
conv_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/yatharth97/10k_reports_gemma_v2.10k_recipes
CookBookAI EDA
This Exploratory Data Analysis (EDA) is related to the following Hugging Face Space:
CookBookAI Space
Data Overview:
The dataset consists of 10,000 synthetically generated recipes.
Full Analysis:
To view the full EDA process, you can visit the notebook directly:
EDA_AppLegacy.ipynb
1. Data Validation & Structure
We began by performing rigorous validation, checking for row duplicates, empty columns, and title repetitions.
Duplicate Analysis: The… See the full description on the dataset page: https://huggingface.co/datasets/Liori25/10k_recipes.ahsanaseer_top-rated-tmdb-movies-10k
TMDB Movies Dataset
Dataset of 10k top rated TMDB movies for text preprocessing (NLP)
Dataset Info
Source: Kaggle
Original Size: 1.43 MB
Kaggle Downloads: 8,035
Files: 1
Files
top10K-TMDB-movies.csv
Mirrored from Kaggle
Urdu-Turn-Detection-10k
Urdu Turn Detection Dataset 🗣️
A high-quality dataset of 10,000 Urdu sentences labeled for Turn Detection (End-of-Turn). This dataset is designed to help conversational AI systems determine if a user has finished speaking (Complete) or is pausing/trailing off (Incomplete).
Dataset Details
Total Samples: 10,000
Language: Urdu (ur) - Nastaliq/Arabic Script only.
Cleanliness: - 100% Urdu Script (No Roman/English).
Avg. Sentence Length: - ~7.7 words (33 characters)… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/Urdu-Turn-Detection-10k.think-10k
think-10k
A dataset with extract rows the dataset in this collection.
List of categories:
general_qa
code
science_qa
math
creative_writing
brainstorming
summarization
information_extraction
classification
Dataset structure
main/train.csv -- the full 10k training datasft/train.csv -- 2k rows for SFT warmup before RLrl/train.csv -- 8k forws for RL
