CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shibing624 /alpaca-zh Dataset Card for "alpaca-zh" 本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。 Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.texttext-generation10K<n<100K145 likes7.2k downloads3y agoHugging Face02ShinMK3 /Mega-Brain-Distill Mega-Brain-Distill Curated merge of the top 10% highest-scoring examples from 584 community-uploaded LLM distillation/reasoning-trace datasets on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces, etc.), deduplicated within and across all of them — many of these source repos are the same underlying dump re-uploaded by different users. Auto-generated by run.py — do not hand-edit, it will be overwritten on the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.tabulartext-generation10K<n<100K2 likes5.4k downloads2mo agoHugging Face03Shivamg031 /say-idc-media-vaulttextn<1K0 likes5.3k downloads20d agoHugging Face04Shivamkak /STUZero-Atari-Dynamics STUZero Atari Dynamics Dataset Offline dynamics training datasets collected from trained EfficientZero V2 (EZv2) benchmark models on Atari games. Each game's data is stored in a subfolder named {game}_{steps} indicating the game and the number of training steps of the source checkpoint. While all models were trained for 120K steps, best results in some games were attained at earlier checkpoints. The model with best eval scores was used to curate data for each game.… See the full description on the dataset page: https://huggingface.co/datasets/Shivamkak/STUZero-Atari-Dynamics.textreinforcement-learningn<1K0 likes5k downloads6mo agoHugging Face05shinonomelab /cleanvid-15m_map CleanVid Map (15M) 🎥 TempoFunk Video Generation Project CleanVid-15M is a large-scale dataset of videos with multiple metadata entries such as: Textual Descriptions 📃 Recording Equipment 📹 Categories 🔠 Framerate 🎞️ Aspect Ratio 📺 CleanVid aim is to improve the quality of WebVid-10M dataset by adding more data and cleaning the dataset by dewatermarking the videos in it. This dataset includes only the map with the urls and metadata, with 3,694,510 more entries than… See the full description on the dataset page: https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map.tabulartext-to-video10M<n<100M23 likes4.8k downloads3y agoHugging Face06shijianjian /ZDPShift ZDPShift: Beyond the Zero-Disparity Plane in Stereo Every public stereo benchmark assumes positive disparity valuesd = fB/Z ≥ 0. Mordern stereoscopic display — cinema 3D, VR, HMDs — actively uses d < 0. ZDPShift bridges the gap: the same artist-authored open-movie content rendered at five ZDP shifts Δ ∈ {−16, 0, +16, +24, +32} pixels, giving you a controlled continuum from textbook-positive to substantially-crossed disparities, with analytical ground truth at every pixel.… See the full description on the dataset page: https://huggingface.co/datasets/shijianjian/ZDPShift.imagedepth-estimation1K<n<10K0 likes4.4k downloads1mo agoHugging Face07shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face08BangumiBase /shiunjikenokodomotachi Bangumi Image Base of Shiunji-ke No Kodomotachi This is the image base of bangumi Shiunji-ke no Kodomotachi, we detected 39 characters, 4423 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shiunjikenokodomotachi.image1K<n<10K0 likes2.8k downloads1y agoHugging Face09shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face10jameszhou-gl /gpt-4v-distribution-shift License This repository is licensed under the MIT License. Description This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift. These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios. Using the Dataset For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.imagen<1K0 likes2.6k downloads3y agoHugging Face11Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.4k downloads4y agoHugging Face12Shitao /bge-m3-data Dataset Summary This depository contains all the fine-tuning data for the bge-m3 model, including: Dataset Language MS MARCO English NQ English HotpotQA English TriviaQA English SQuAD English COLIEE English PubMedQA English NLI from SimCSE English DuReader Chinese mMARCO-zh Chinese T2Ranking Chinese Law-GPT Chinese cMedQAv2 Chinese NLI-zh Chinese LeCaRDv2 Chinese Mr.TyDi 11 languages MIRACL 16 languages MLDR 13 languages Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/bge-m3-data.text100K<n<1M55 likes2k downloads2y agoHugging Face13shijli /aitod-v2 AI-TOD-v2 AI-TOD-v2, the tiny-object detection benchmark in aerial images, packed once with the official v2 annotations kept whole, so it loads in one line and no data path has to be configured: from datasets import load_dataset ds = load_dataset("shijli/aitod-v2") # 11214 train / 2804 validation / 14018 test AI-TOD cuts 28036 images of 800 x 800 pixels from xView, DOTA-v1.5, VisDrone2018-Det, Airbus Ship Detection and DIOR, and annotates eight classes whose mean object size… See the full description on the dataset page: https://huggingface.co/datasets/shijli/aitod-v2.imageobject-detection10K<n<100K3 likes1.9k downloads18d agoHugging Face14vidore /shiftproject_test_beirBEIR version of vidore/shiftproject_test. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face15shibing624 /nli_zh纯文本数据,格式:(sentence1, sentence2, label)。常见中文语义匹配数据集,包含ATEC、BQ、LCQMC、PAWSX、STS-B共5个任务。texttext-classification100K<n<1M48 likes1.7k downloads4y agoHugging Face16shijiezhou /VLM4D VLM4D VLM4D is a benchmark for evaluating the spatiotemporal reasoning capabilities of Vision Language Models (VLMs). It contains real and synthetic videos paired with multiple-choice questions that require models to reason about translation, rotation, perspective, motion continuity, counting, and false-positive events. The dataset was introduced in VLM4D: Towards Spatiotemporal Awareness in Vision Language Models, accepted to ICCV 2025. Project page: https://vlm4d.github.io/… See the full description on the dataset page: https://huggingface.co/datasets/shijiezhou/VLM4D.textvideo-text-to-textn<1K4 likes1.7k downloads3mo agoHugging Face17ShirohAO /tuxun Anonymization For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information. GeoComp Dataset description Inspired by geoguessr.com, we developed a free geolocation game platform that tracks participants' competition histories. Unlike most geolocation websites, including Geoguessr, which rely solely on samples from Google Street… See the full description on the dataset page: https://huggingface.co/datasets/ShirohAO/tuxun.tabular10M<n<100M13 likes1.6k downloads9mo agoHugging Face18shi-labs /physical-ai-bench-conditional-generation Physical AI Bench - Conditional Generation Paper | Code This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below. Dataset Category Sample Nums Agibot World Robotics 200 OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.textvideo-to-videon<1K0 likes1.6k downloads10mo agoHugging Face19BangumiBase /shijousaikyounodaimaoumurabitoanitenseisuru Bangumi Image Base of Shijou Saikyou No Daimaou, Murabito A Ni Tensei Suru This is the image base of bangumi Shijou Saikyou no Daimaou, Murabito A ni Tensei suru, we detected 70 characters, 4747 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shijousaikyounodaimaoumurabitoanitenseisuru.image1K<n<10K0 likes1.6k downloads2y agoHugging Face20ShiroOnigami23 /skin-cancer-ham10000-datasetimage1K<n<10K1 likes1.4k downloads8mo agoHugging Face21CarsonnnNN /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M1 likes1.3k downloads7mo agoHugging Face22shibing624 /nli-zh-allThe SNLI corpus (version 1.0) is a merged chinese sentence similarity dataset, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE).texttext-classification10K<n<100K46 likes1.2k downloads3y agoHugging Face23BangumiBase /shiguangdailirenii Bangumi Image Base of Shiguang Dailiren Ii This is the image base of bangumi Shiguang Dailiren II, we detected 48 characters, 3677 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shiguangdailirenii.image1K<n<10K0 likes1.2k downloads2y agoHugging Face24shivDwd /W_LSTMix_test_datasettext10M<n<100M0 likes1.2k downloads1y agoHugging Face25BangumiBase /shirobako Bangumi Image Base of Shirobako This is the image base of bangumi Shirobako, we detected 52 characters, 3771 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shirobako.image1K<n<10K0 likes1.2k downloads3y agoHugging Face26shivank21 /mmconflict-editable-values-1k MMConflict Editable Values 2K This dataset contains 2,000 source images with visible atomic values for multimodal conflict research. It has 100 images in each of 20 categories. Every image comes from a photograph, scan, captured website, software screenshot, or page of a source document. The dataset does not contain generated images or project-rendered examples. Each row records the source, source URL, license, attribution, visible value, question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.imageimage-to-text1K<n<10K0 likes1.1k downloads23d agoHugging Face27BangumiBase /shinnonakamajanaitoyuushanopartywooidasaretanodehenkyoudeslowlifesurukotonishimashita2nd Bangumi Image Base of Shin No Nakama Ja Nai To Yuusha No Party Wo Oidasareta Node, Henkyou De Slow Life Suru Koto Ni Shimashita 2nd This is the image base of bangumi Shin no Nakama ja Nai to Yuusha no Party wo Oidasareta node, Henkyou de Slow Life suru Koto ni Shimashita 2nd, we detected 69 characters, 4925 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shinnonakamajanaitoyuushanopartywooidasaretanodehenkyoudeslowlifesurukotonishimashita2nd.image1K<n<10K0 likes1.1k downloads2y agoHugging Face28mteb /shiftproject_test_beirBEIR version of vidore/shiftproject_test. imagedocument-question-answering1K<n<10K0 likes969 downloads8mo agoHugging Face29shiwk24 /MathCanvas-Edit MathCanvas-Edit Dataset                   🚀 Data Usage from datasets import load_dataset dataset = load_dataset("shiwk24/MathCanvas-Edit") print(dataset) 📖 Overview MathCanvas-Edit is a large-scale dataset containing 5.2 million step-by-step editing trajectories, forming a crucial component of the [MathCanvas] framework. MathCanvas is designed to endow Unified Large Multimodal Models (LMMs) with intrinsic… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Edit.imageimage-to-image1M<n<10M3 likes957 downloads10mo agoHugging Face30shiwk24 /MathCanvas-Instruct MathCanvas-Instruct Dataset                   🚀 Data Usage from datasets import load_dataset dataset = load_dataset("shiwk24/MathCanvas-Instruct") print(dataset) 📖 Overview MathCanvas-Instruct is a high-quality, fine-tuning dataset with 219K examples of interleaved visual-textual reasoning paths. It is the core component for the second phase of the [MathCanvas] framework: Strategic Visual-Aided Reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Instruct.imageimage-text-to-text100K<n<1M6 likes940 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.