CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shi-labs /oneformer_demo0 likes119k downloads4y agoHugging Face02shihao1895 /bridge-rlds Dataset Structure These datasets are used for MemoryVLA training. This is the standard setting and can be directly used for other models as well.All data follow the RLDS format from the Bridge dataset. bridge_orig — 60k+ episodes, widowx robot robotics0 likes74k downloads11mo agoHugging Face03CoderOfCode /ship-tracking-data2 likes35k downloads2mo agoHugging Face04Dario-Shit4 /behavior-1k_2025-challenge-demosThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "R1Pro", "total_episodes": 10000, "total_frames": 119094660, "total_tasks": 50, "total_videos": 90000, "chunks_size": 10000, "fps": 30, "splits": { "train": "0:10000" }, "data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dario-Shit4/behavior-1k_2025-challenge-demos.tabularrobotics100M<n<1B0 likes7.3k downloads2mo agoHugging Face05shibing624 /alpaca-zh Dataset Card for "alpaca-zh" 本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。 Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.texttext-generation10K<n<100K145 likes7.2k downloads3y agoHugging Face06yaacovgg /shiur-clips-flac0 likes6.9k downloads2mo agoHugging Face07tweettemposhift /tweet_temporal_shift""" _TWEET_TEMPORAL_CITATION =0 likes6.9k downloads3y agoHugging Face08ShinMK3 /Mega-Brain-Distill Mega-Brain-Distill Curated merge of the top 10% highest-scoring examples from 584 community-uploaded LLM distillation/reasoning-trace datasets on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces, etc.), deduplicated within and across all of them — many of these source repos are the same underlying dump re-uploaded by different users. Auto-generated by run.py — do not hand-edit, it will be overwritten on the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.tabulartext-generation10K<n<100K2 likes5.4k downloads2mo agoHugging Face09Shivamg031 /say-idc-media-vaulttextn<1K0 likes5.3k downloads20d agoHugging Face10Shiym /ViT-FineTuneimageimage-classification10K<n<100K0 likes5.2k downloads2y agoHugging Face11Shivamkak /STUZero-Atari-Dynamics STUZero Atari Dynamics Dataset Offline dynamics training datasets collected from trained EfficientZero V2 (EZv2) benchmark models on Atari games. Each game's data is stored in a subfolder named {game}_{steps} indicating the game and the number of training steps of the source checkpoint. While all models were trained for 120K steps, best results in some games were attained at earlier checkpoints. The model with best eval scores was used to curate data for each game.… See the full description on the dataset page: https://huggingface.co/datasets/Shivamkak/STUZero-Atari-Dynamics.textreinforcement-learningn<1K0 likes5k downloads6mo agoHugging Face12shinonomelab /cleanvid-15m_map CleanVid Map (15M) 🎥 TempoFunk Video Generation Project CleanVid-15M is a large-scale dataset of videos with multiple metadata entries such as: Textual Descriptions 📃 Recording Equipment 📹 Categories 🔠 Framerate 🎞️ Aspect Ratio 📺 CleanVid aim is to improve the quality of WebVid-10M dataset by adding more data and cleaning the dataset by dewatermarking the videos in it. This dataset includes only the map with the urls and metadata, with 3,694,510 more entries than… See the full description on the dataset page: https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map.tabulartext-to-video10M<n<100M23 likes4.8k downloads3y agoHugging Face13shixuesong /openloris-scene OpenLORIS-Scene Datasets This page is an index of the OpenLORIS-Scene datasets. Terms of Use The OpenLORIS-Scene datasets are released with the CC BY-ND 4.0 license, which means you can do anything with the data, even for commercial purposes, except distributing derivative datasets (contact us at openloris@gmail.com if you would like to do so). We would appreciate it if you citeour paper when appropriate. Cite X Shi, D Li et al. “Are We Ready for Service… See the full description on the dataset page: https://huggingface.co/datasets/shixuesong/openloris-scene.1 likes4.5k downloads9mo agoHugging Face14shijianjian /ZDPShift ZDPShift: Beyond the Zero-Disparity Plane in Stereo Every public stereo benchmark assumes positive disparity valuesd = fB/Z ≥ 0. Mordern stereoscopic display — cinema 3D, VR, HMDs — actively uses d < 0. ZDPShift bridges the gap: the same artist-authored open-movie content rendered at five ZDP shifts Δ ∈ {−16, 0, +16, +24, +32} pixels, giving you a controlled continuum from textbook-positive to substantially-crossed disparities, with analytical ground truth at every pixel.… See the full description on the dataset page: https://huggingface.co/datasets/shijianjian/ZDPShift.imagedepth-estimation1K<n<10K0 likes4.4k downloads1mo agoHugging Face15Shirk6 /Challenge-phase1-dataset-rlinf0 likes4k downloads4mo agoHugging Face16shihao1895 /libero-rlds Dataset Structure These datasets are used for MemoryVLA training. This is the standard LIBERO setting and can be directly used for other models as well.All data follow the RLDS format from the LIBERO benchmark, where each task initially contains 50 trajectories and failed rollouts are filtered out.NOTE: LIBERO-90 is also included. libero_spatial_no_noops — 10 tasks libero_object_no_noops — 10 tasks libero_goal_no_noops — 10 tasks libero_10_no_noops — 10 tasks… See the full description on the dataset page: https://huggingface.co/datasets/shihao1895/libero-rlds.robotics0 likes3.8k downloads11mo agoHugging Face17Shilin-LU /vae_cache_minecraft_480p_9s0 likes3.6k downloads1y agoHugging Face18shintaro-ozaki /entity-explanationimagetext-generation100B<n<1T2 likes3.2k downloads11mo agoHugging Face19shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face20BangumiBase /shiunjikenokodomotachi Bangumi Image Base of Shiunji-ke No Kodomotachi This is the image base of bangumi Shiunji-ke no Kodomotachi, we detected 39 characters, 4423 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shiunjikenokodomotachi.image1K<n<10K0 likes2.8k downloads1y agoHugging Face21shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face22jameszhou-gl /gpt-4v-distribution-shift License This repository is licensed under the MIT License. Description This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift. These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios. Using the Dataset For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.imagen<1K0 likes2.6k downloads3y agoHugging Face23Shiki42 /robotwin_scan_object_place_dual_shoes_serialized_2000 likes2.4k downloads2mo agoHugging Face24shibing624 /medical纯文本数据,中文医疗数据集,包含预训练数据的百科数据,指令微调数据和奖励模型数据。text-generationn<1K442 likes2.4k downloads2y agoHugging Face25Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.4k downloads4y agoHugging Face26EvanOLeary /shinka-cvdp-benchmark-full0 likes2k downloads4mo agoHugging Face27Shitao /bge-m3-data Dataset Summary This depository contains all the fine-tuning data for the bge-m3 model, including: Dataset Language MS MARCO English NQ English HotpotQA English TriviaQA English SQuAD English COLIEE English PubMedQA English NLI from SimCSE English DuReader Chinese mMARCO-zh Chinese T2Ranking Chinese Law-GPT Chinese cMedQAv2 Chinese NLI-zh Chinese LeCaRDv2 Chinese Mr.TyDi 11 languages MIRACL 16 languages MLDR 13 languages Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/bge-m3-data.text100K<n<1M55 likes2k downloads2y agoHugging Face28Shionone84 /Shionone840 likes2k downloads3y agoHugging Face29Shitao /MLDR Dataset Summary MLDR is a Multilingual Long-Document Retrieval dataset built on Wikipeida, Wudao and mC4, covering 13 typologically diverse languages. Specifically, we sample lengthy articles from Wikipedia, Wudao and mC4 datasets and randomly choose paragraphs from them. Then we use GPT-3.5 to generate questions based on these paragraphs. The generated question and the sampled article constitute a new text pair to the dataset. The prompt for GPT3.5 is “You are a curious AI… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/MLDR.text-retrieval82 likes1.9k downloads3y agoHugging Face30shiyi666777 /RGB-Event-ISP-Datasetimage100B<n<1T0 likes1.9k downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.