CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hakari-bench /NanoMTEB-Scandinavian NanoMTEB-Scandinavian This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoMTEB-Scandinavian is a compact retrieval benchmark for Scandinavian-language MTEB-style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social-media retrieval. Usage from datasets import load_dataset dataset_id = "hakari-bench/NanoMTEB-Scandinavian" split… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Scandinavian.text10K<n<100K0 likes4k downloads3mo agoHugging Face02Scandium-Labs /Scandium-Dataset Dataset Card — Scandium-Dataset v1.0.0 Summary Scandium-Dataset provides a harmonized, quality-scored foundation of DFT-computed structural and thermodynamic properties across 267,230 materials from Materials Project, OQMD, and JARVIS-DFT. It supports the early screening stage of battery materials discovery — filtering by phase stability, electronic structure, and structural family — before downstream property prediction (ionic conductivity, mechanical stability… See the full description on the dataset page: https://huggingface.co/datasets/Scandium-Labs/Scandium-Dataset.tabularother100K<n<1M2 likes2.5k downloads2mo agoHugging Face03dgorbatov /vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10 Trajectory Ranking Dataset This dataset contains trajectory ranking results for autonomous navigation scenarios. Dataset Statistics Total examples: 39558 Chunks processed: 40 Upload date: 2025-09-13T00:44:30.335177 Features Image data with terrain analysis Trajectory rankings and reasoning Quality and diversity analysis Terrain and trajectory descriptions imageimage-classification10K<n<100K0 likes702 downloads1y agoHugging Face04alexandrainst /scandi-reddit Dataset Card for ScandiReddit Dataset Summary ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit. All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept. The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.texttext-generation10M<n<100M5 likes636 downloads2y agoHugging Face05mteb /scandisenttext10K<n<100K0 likes626 downloads1y agoHugging Face06suz22 /RoboTwin_scan_object_randomizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 500, "total_frames": 80479, "total_tasks": 499, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:500" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_scan_object_randomized.imagerobotics10K<n<100K0 likes566 downloads6mo agoHugging Face07GumJump /scanqa_images_16_keyframes_120_non_keyframes_min_532_long_edgeimage10K<n<100K0 likes406 downloads11mo agoHugging Face08davidilag /scandinavian_faroeseaudio100K<n<1M0 likes393 downloads2y agoHugging Face09Shiki42 /ctr-scan-object-uniform50-20260917 IdleMask review — passed (2026-09-19) Reviewed by the dataset owner: observation.arm_active_mask is correct and this revision is a formally usable CTR dataset. This section supersedes previous active/idle-mask descriptions below. For each of the 50 episodes and each physical arm, only the initial contiguous scheduling delay may have mask 0. From first duty through the final frame the mask is always 1. Scan synchronization waits, cooperative holds, the shared scan tail and… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-scan-object-uniform50-20260917.tabularrobotics10K<n<100K0 likes374 downloads6d agoHugging Face10threefruits /SCAND_traj_selectionimage10K<n<100K0 likes308 downloads1y agoHugging Face11loay /arabic-ocr-synthetic-scans-faker-300k Arabic OCR Synthetic Scans (Faker 300k) A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness. Dataset Summary Samples: ~300,000 synthetic Arabic document pages Image format: JPEG, ~800×1200 px (embedded in Parquet) Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.imageimage-to-text100K<n<1M7 likes290 downloads8mo agoHugging Face12jinaai /europeana-it-scans_beirThis is a copy of https://huggingface.co/datasets/jinaai/europeana-it-scans reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/europeana-it-scans_beir.image1K<n<10K0 likes262 downloads1y agoHugging Face13mateoguaman /vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5 vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5 Description VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5. Processing Parameters {} Dataset Configuration Train dataset: mixer: mateoguaman/coda_every1_25pct_sub5: 1.0 mateoguaman/iphone_stairs_ramps: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5.image100K<n<1M0 likes255 downloads1y agoHugging Face14ns69956 /ur5e_scanner_fp_deltaTCPThis dataset was created using LeRobot format. Dataset Structure { "codebase_version": "v2.1", "robot_type": "ur5e", "total_episodes": 60, "total_frames": 31577, "total_tasks": 1, "total_videos": 120, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:59" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ns69956/ur5e_scanner_fp_deltaTCP.tabularrobotics10K<n<100K0 likes241 downloads1y agoHugging Face15kardosdrur /scandi_eurovoctext100K<n<1M0 likes219 downloads3y agoHugging Face16jxie /scanobjectnn Dataset Card for "scanobjectnn" More Information needed 10K<n<100K0 likes198 downloads3y agoHugging Face17Lots-of-LoRAs /task131_scan_long_text_generation_action_command_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.texttext-generation1K<n<10K0 likes185 downloads2y agoHugging Face18mateoguaman /scand_every1_50pct_sub5 scand_every1_50pct_sub5 Description Processed scand dataset with filter_every_nth=1, 50% of data, and num_subsampled_points=5 Processing Parameters mateoguaman/scand: exclude_outliers_pct: 3 filter_by_curvature: true filter_every_nth: 1 horizon: 300: 0.5 500: 0.5 num_subsampled_points: 5 Dataset Configuration Train dataset: mixer: mateoguaman/scand: 0.5 split: train Validation dataset: mixer: mateoguaman/scand:… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/scand_every1_50pct_sub5.image100K<n<1M0 likes175 downloads1y agoHugging Face19Shiki42 /ctr-scan-object-uniform100-20260921 scan_object Uniform100 Open in LeRobot Dataset Visualizer 100 episodes from the same 50 source scene seeds, 25 FPS, LeRobot v3, Aloha AgileX. Parent: Shiki42/ctr-scan-object-uniform-20260916 at 766e18802a11bdde94f6dab6892c13d82e6a5bd3 (repaired gripper target labels). This is a byte-exact republication of the pinned parent for the non-mainline 100-sample comparison arm; the 50-episode arm remains ctr-scan-object-uniform50-20260917. Every source seed appears exactly twice, as… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-scan-object-uniform100-20260921.tabularrobotics10K<n<100K0 likes171 downloads4d agoHugging Face20mateoguaman /vlmn_scand_spot_sub5 vlmn_scand_spot_sub5 Description VLN Navigation dataset with 50% of scand data and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5. Processing Parameters {} Dataset Configuration Train dataset: mixer: mateoguaman/scand_every1_50pct_sub5: 1.0 mateoguaman/spot_every1_sub5: 1.0 split: train Validation dataset: mixer: mateoguaman/scand_every1_50pct_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_scand_spot_sub5.image100K<n<1M0 likes169 downloads1y agoHugging Face21thivy /scandinavian-embedding-training-datatext1M<n<10M1 likes159 downloads7mo agoHugging Face22leo66666 /scannet_countingimagen<1K0 likes148 downloads7mo agoHugging Face23IAMJB /scanned-arxiv-papers-idtext100K<n<1M1 likes146 downloads2y agoHugging Face24mateoguaman /vlmn_tartandrive100_scand50_coda25_spot100_sub5_filtered_trajectories_training_25_fixedimage100K<n<1M0 likes146 downloads1y agoHugging Face25threefruits /SCAND_path_selectionimage1K<n<10K0 likes137 downloads2y agoHugging Face26Lots-of-LoRAs /task127_scan_long_text_generation_action_command_all Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task127_scan_long_text_generation_action_command_all Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task127_scan_long_text_generation_action_command_all.texttext-generation1K<n<10K0 likes135 downloads2y agoHugging Face27ns69956 /ur5e_scanner_tcp_TRAINThis dataset was created using LeRobot format. Dataset Structure { "codebase_version": "v2.1", "robot_type": "ur5e", "total_episodes": 50, "total_frames": 24200, "total_tasks": 2, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:49" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ns69956/ur5e_scanner_tcp_TRAIN.tabularrobotics10K<n<100K0 likes133 downloads1y agoHugging Face28ScandLM /danish_culturax Danish Culturax Dataset This dataset is simply a reformatting of uonlp/CulturaX. Some minor formatting errors have been corrected. Usage from datasets import load_dataset dataset = load_dataset("ScandLM/danish_culturax") text10M<n<100M0 likes127 downloads2y agoHugging Face29davidilag /scandinavian-100haudio100K<n<1M0 likes123 downloads2y agoHugging Face30IAMJB /scanned-arxiv-paperstextn<1K0 likes122 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.