CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes6.8k downloads11mo agoHugging Face02blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes991 downloads2y agoHugging Face03ruili0 /LongVA-TPO-10k 10kTemporal Preference Optimization Dataset for LongVA LongVA-TPO-10k, introduced by paper Temporal Preference Optimization for Long-form Video Understanding textvideo-text-to-text10K<n<100K4 likes957 downloads2y agoHugging Face04simoneteglia /europarl_for_language_detection_10ktext100K<n<1M0 likes739 downloads3y agoHugging Face05Hatman /PlotPalette-10K Plot Palette is a curated dataset designed for fine-tuning large language models (LLMs) on creative writing tasks. Sourced from various literary sources and generated using the Mistral 8x7B language model. The scripts used to generate the data can be found here. Data Fields 'id': A unique identifier for each prompt-response pair. 'category': The category to which the prompt-response pair belongs (e.g., creative_writing, generation, poem, brainstorm, question_answer). --- (… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/PlotPalette-10K.text1K<n<10K1 likes266 downloads2y agoHugging Face06code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes250 downloads9mo agoHugging Face07Humanbased-AI /Crypto-Address-Annotation-10K Codatta Crypto Address Annotations (Sample) Overview This dataset is a 10,000-row sample of the comprehensive Codatta Crypto Address Annotations database. The full database serves as a massive repository of over 500 million labeled address pairs across multiple blockchains. The data provides critical metadata aimed at solving the problem of fragmented and siloed blockchain information. It includes entity names, functional categories (e.g., Exchanges, DeFi, Scam)… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Crypto-Address-Annotation-10K.texttoken-classification10K<n<100K1 likes240 downloads10mo agoHugging Face08chungimungi /VideoDPO-10k@misc{liu2024videodpoomnipreferencealignmentvideo, title={VideoDPO: Omni-Preference Alignment for Video Diffusion Generation}, author={Runtao Liu and Haoyu Wu and Zheng Ziqiang and Chen Wei and Yingqing He and Renjie Pi and Qifeng Chen}, year={2024}, eprint={2412.14167}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2412.14167}, } @misc{wang2024vidprommillionscalerealpromptgallery, title={VidProM: A Million-scale Real… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/VideoDPO-10k.textvideo-classification10K<n<100K1 likes156 downloads2y agoHugging Face09Aipresso /10k_rows_cleaned_prompts 10K Rows Cleaned Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models. 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.texttext-generation1M<n<10M0 likes126 downloads11mo agoHugging Face10aguennoune17 /atlas-crispr-10k-benchmark 🧬 ATLAS CRISPR 10k Benchmark Contribution communauté LeWorldModel — Benchmark CRISPR 10k guides ARNUtilisé pour fine-tuner aguennoune17/negenWM-jepa-v2 — ATLAS NWM Sprint 3Self-supervised I-JEPA · Encodage téléologique (κ, τ, λ) Description 10 000 guides ARN Cas9 de 20 nucléotides consolidés depuis 12 études expérimentales de criblage CRISPR génomique à grande échelle. Ce dataset est le benchmark officiel du Sprint 3 ATLAS NWM v2 — entraînement I-JEPA… See the full description on the dataset page: https://huggingface.co/datasets/aguennoune17/atlas-crispr-10k-benchmark.tabularfeature-extraction10K<n<100K0 likes113 downloads5mo agoHugging Face11arafatanam /Student-Mental-Health-Counseling-10K Student Mental Health Counseling 10K Dataset Overview This dataset is derived from the original chillies/student-mental-health-counseling-vn dataset, which contains student mental health counseling conversations in Vietnamese. In this version, 10,000 randomly sampled rows from the original dataset have been translated into English using Google Translator, making the data more accessible for English-speaking researchers and developers. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/arafatanam/Student-Mental-Health-Counseling-10K.texttext-classification10K<n<100K2 likes77 downloads1y agoHugging Face12purnasai /SEC-10Q-10K-Statement-tablestext10K<n<100K0 likes73 downloads3y agoHugging Face13CreitinGameplays /magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1text10K<n<100K1 likes61 downloads2y agoHugging Face14Guizhen /Puzzles_10ktext10K<n<100K0 likes48 downloads2y agoHugging Face15gandharvbakshi /SMS-dataset-sample-10klicense: mit language: en 📌 Free 10K SMS Preview Dataset This is a preview subset of the full OTP + OTP INTENT + Phishing dataset. 👉 For full 73K dataset: https://huggingface.co/datasets/gandharvbakshi/SMS-dataset-OTP-OTP_INTENT_Phishing text10K<n<100K0 likes48 downloads10mo agoHugging Face16oscar128372 /chess_spatial_reasoning_10ktext10K<n<100K1 likes46 downloads2y agoHugging Face17BioDockify /alzheimers-multi-target-10k-dataset 📚 BioDockify: Multi-Target Alzheimer's Chemical Space & Virtual Screening Dataset (10,000 Verified Compounds) Principal Investigator: Tajuddin Shaik (tajo9128@gmail.com)Affiliation: Faculty of Pharmacy, Bharath Institute of Higher Education and Research (BIHER), Chennai, IndiaPlatform: www.biodockify.com | ai.biodockify.com 📌 Dataset Summary This repository contains the complete 10,000 curated, literature-grounded chemical space dataset for Alzheimer's… See the full description on the dataset page: https://huggingface.co/datasets/BioDockify/alzheimers-multi-target-10k-dataset.texttabular-classification10K<n<100K0 likes44 downloads23d agoHugging Face18scalarlogicgroup /synthetic-nsclc-10ktabular10K<n<100K0 likes43 downloads2mo agoHugging Face19upb-nlp /MoltSafe-10K MoltSafe-10K MoltSafe-10K contains safety annotations for 10,000 posts and comments from Moltbook, an agentic social network where autonomous agents communicate with one another. The underlying corpus is AIcell/moltbook-data on the Hugging Face Hub. Every node has a binary safety verdict, a severity level, a malicious-intent taxonomy, and OWASP GenAI risk codes. We release the annotations used in the development our study to promote research into Moltbook security. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/MoltSafe-10K.texttext-classification10K<n<100K0 likes43 downloads18d agoHugging Face20Jonathan-Zhou /GameLabel-10kGameLabel-10k Dataset Card This dataset contains was created in collaboration with the game developers of Armchair Commander. It contains 9800 human preferences over pairs of Flux-Schnell generated images, with over 6800 unique prompts. All labels were crowdsourced from Armchair Commander players. Usage Example from datasets import load_dataset from PIL import Image import base64 from io import BytesIO dataset = load_dataset("Jonathan-Zhou/GameLabel-10k") # For some reason, when using… See the full description on the dataset page: https://huggingface.co/datasets/Jonathan-Zhou/GameLabel-10k.tabular1K<n<10K0 likes36 downloads2y agoHugging Face21purnasai /SEC-10Q-10K-Coverpagetexttoken-classification1K<n<10K0 likes32 downloads3y agoHugging Face22ambrosfitz /10k_history_v5text1K<n<10K0 likes32 downloads2y agoHugging Face23ambrosfitz /10k_history_summarySummarized over 10,000 rows of historical information, mostly US History. Initally generated from historical text using gpt-4o-mini, summarized using Mistral-Nemo. textsummarization10K<n<100K0 likes32 downloads2y agoHugging Face24dirtycomputer /waimai_10ktext10K<n<100K1 likes31 downloads4y agoHugging Face25chinna887 /ecommerce-fraud-detection-synthetic-10k-sampl 🛡️ Synthetic E-Commerce Fraud & AML Detection Dataset (10k Evaluation Sample) ⚠️ NOTICE: This is a truncated 10,000-row evaluation sample strictly for schema verification and local testing. 💳 [OBTAIN THE 10-MILLION ROW COMMERCIAL LICENSE HERE] > https://buy.stripe.com/8x26oIad4eH9eJf6gJ5wI01 🚀 Quick Start (Load via Hugging Face) Data scientists can instantly load this evaluation slice into their Pandas/Python environment using the datasets library: from… See the full description on the dataset page: https://huggingface.co/datasets/chinna887/ecommerce-fraud-detection-synthetic-10k-sampl.texttabular-classification10K<n<100K0 likes31 downloads1mo agoHugging Face26yatharth97 /10k_reports_gemma_v2 Dataset Card for Financial Document Analysis Dataset Dataset Description This dataset comprises structured conversational entries designed to facilitate the training and evaluation of models that analyze and summarize financial documents. Each entry includes a conversation ID, a specific step in the conversation, a system-generated prompt, a user question, and the corresponding model-generated response. Fields Overview conv_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/yatharth97/10k_reports_gemma_v2.text10K<n<100K0 likes30 downloads2y agoHugging Face27Liori25 /10k_recipes CookBookAI EDA This Exploratory Data Analysis (EDA) is related to the following Hugging Face Space: CookBookAI Space Data Overview: The dataset consists of 10,000 synthetically generated recipes. Full Analysis: To view the full EDA process, you can visit the notebook directly: EDA_AppLegacy.ipynb 1. Data Validation & Structure We began by performing rigorous validation, checking for row duplicates, empty columns, and title repetitions. Duplicate Analysis: The… See the full description on the dataset page: https://huggingface.co/datasets/Liori25/10k_recipes.texttext-generation10K<n<100K0 likes30 downloads8mo agoHugging Face28jason1966 /ahsanaseer_top-rated-tmdb-movies-10k TMDB Movies Dataset Dataset of 10k top rated TMDB movies for text preprocessing (NLP) Dataset Info Source: Kaggle Original Size: 1.43 MB Kaggle Downloads: 8,035 Files: 1 Files top10K-TMDB-movies.csv Mirrored from Kaggle tabular10K<n<100K0 likes30 downloads6mo agoHugging Face29PuristanLabs1 /Urdu-Turn-Detection-10k Urdu Turn Detection Dataset 🗣️ A high-quality dataset of 10,000 Urdu sentences labeled for Turn Detection (End-of-Turn). This dataset is designed to help conversational AI systems determine if a user has finished speaking (Complete) or is pausing/trailing off (Incomplete). Dataset Details Total Samples: 10,000 Language: Urdu (ur) - Nastaliq/Arabic Script only. Cleanliness: - 100% Urdu Script (No Roman/English). Avg. Sentence Length: - ~7.7 words (33 characters)… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/Urdu-Turn-Detection-10k.texttext-classification10K<n<100K0 likes29 downloads10mo agoHugging Face30chimbiwide /think-10k think-10k A dataset with extract rows the dataset in this collection. List of categories: general_qa code science_qa math creative_writing brainstorming summarization information_extraction classification Dataset structure main/train.csv -- the full 10k training datasft/train.csv -- 2k rows for SFT warmup before RLrl/train.csv -- 8k forws for RL texttext-generation10K<n<100K0 likes29 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.