CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01trentmkelly /police-scanner-audio Police Scanner Audio Dataset A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds. Dataset Overview This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications. Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.audioautomatic-speech-recognition1K<n<10K1 likes4.2k downloads1y agoHugging Face02hakari-bench /NanoMTEB-Scandinavian NanoMTEB-Scandinavian This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoMTEB-Scandinavian is a compact retrieval benchmark for Scandinavian-language MTEB-style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social-media retrieval. Usage from datasets import load_dataset dataset_id = "hakari-bench/NanoMTEB-Scandinavian" split… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Scandinavian.text10K<n<100K0 likes4k downloads3mo agoHugging Face03Scandium-Labs /Scandium-Dataset Dataset Card — Scandium-Dataset v1.0.0 Summary Scandium-Dataset provides a harmonized, quality-scored foundation of DFT-computed structural and thermodynamic properties across 267,230 materials from Materials Project, OQMD, and JARVIS-DFT. It supports the early screening stage of battery materials discovery — filtering by phase stability, electronic structure, and structural family — before downstream property prediction (ionic conductivity, mechanical stability… See the full description on the dataset page: https://huggingface.co/datasets/Scandium-Labs/Scandium-Dataset.tabularother100K<n<1M2 likes2.5k downloads2mo agoHugging Face04cloudaocr /arabic-synthetic-scanned-booksdocument1K<n<10K1 likes2.4k downloads23d agoHugging Face05fjd /scannet-processed-testimage1 likes2k downloads3y agoHugging Face06HarrisonPENG /scannetppimage1M<n<10M1 likes1.7k downloads5mo agoHugging Face07dgorbatov /vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10 Trajectory Ranking Dataset This dataset contains trajectory ranking results for autonomous navigation scenarios. Dataset Statistics Total examples: 39558 Chunks processed: 40 Upload date: 2025-09-13T00:44:30.335177 Features Image data with terrain analysis Trajectory rankings and reasoning Quality and diversity analysis Terrain and trajectory descriptions imageimage-classification10K<n<100K0 likes702 downloads1y agoHugging Face08cywan /scannetpp_processedimage1M<n<10M0 likes677 downloads1mo agoHugging Face09alexandrainst /scandi-reddit Dataset Card for ScandiReddit Dataset Summary ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit. All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept. The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.texttext-generation10M<n<100M5 likes636 downloads2y agoHugging Face10mteb /scandisenttext10K<n<100K0 likes626 downloads1y agoHugging Face11fjd /scannet_sampleimagen<1K1 likes470 downloads3y agoHugging Face12kairunwen /scannet_temp scannet_temp This repository contains split tar.gz archives for the full local scannet dataset tree. Included scans/* scenes with unpacked usable content splits/* split files Excluded scene-level original *_2d-*.zip files results/ color_90/ pose_90.txt selected_ids_90.txt Download huggingface-cli download kairunwen/scannet_temp --repo-type dataset --local-dir ./scannet_temp_hf Extract mkdir -p extracted for f in… See the full description on the dataset page: https://huggingface.co/datasets/kairunwen/scannet_temp.text0 likes464 downloads4mo agoHugging Face13WHB139426 /Scannet Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding This repository contains the ScanNet dataset (3D scene data and 2D frame data) and refined annotations used for the paper Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding. TAB is a dynamic agentic framework designed for zero-shot 3D Visual Grounding (3D-VG). By operating directly on raw RGB-D streams, TAB reformulates 3D… See the full description on the dataset page: https://huggingface.co/datasets/WHB139426/Scannet.textzero-shot-object-detection1K<n<10K0 likes437 downloads6mo agoHugging Face14EmbodiedCity /ScanReQAtext10K<n<100K0 likes427 downloads4mo agoHugging Face15GumJump /scanqa_images_16_keyframes_120_non_keyframes_min_532_long_edgeimage10K<n<100K0 likes406 downloads11mo agoHugging Face16davidilag /scandinavian_faroeseaudio100K<n<1M0 likes393 downloads2y agoHugging Face17yifei-liu /output_3d_bounding_scannetppv2_vllm_old_descriptiontext10K<n<100K0 likes350 downloads8mo agoHugging Face18trust-and-safety /abuse-scanner-bot-datasettextn<1K0 likes340 downloads1y agoHugging Face19threefruits /SCAND_traj_selectionimage10K<n<100K0 likes308 downloads1y agoHugging Face20ZzZZCHS /processed_scannetimage100K<n<1M0 likes294 downloads2y agoHugging Face21timpal0l /scandisent Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/timpal0l/scandisent.texttext-classification10K<n<100K1 likes292 downloads3y agoHugging Face22loay /arabic-ocr-synthetic-scans-faker-300k Arabic OCR Synthetic Scans (Faker 300k) A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness. Dataset Summary Samples: ~300,000 synthetic Arabic document pages Image format: JPEG, ~800×1200 px (embedded in Parquet) Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.imageimage-to-text100K<n<1M7 likes290 downloads8mo agoHugging Face23jinaai /europeana-it-scans_beirThis is a copy of https://huggingface.co/datasets/jinaai/europeana-it-scans reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/europeana-it-scans_beir.image1K<n<10K0 likes262 downloads1y agoHugging Face24mateoguaman /vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5 vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5 Description VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5. Processing Parameters {} Dataset Configuration Train dataset: mixer: mateoguaman/coda_every1_25pct_sub5: 1.0 mateoguaman/iphone_stairs_ramps: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5.image100K<n<1M0 likes255 downloads1y agoHugging Face25tetracta /llm-xray-lesion-scans VG1 status · 22 September 2026 (Europe/Istanbul) The 7B class opened on 21 September 2026. Any registered account may scan the repositories on the eligible list of the scope page — exact Apache-2.0 revisions, listed there with their status — within the free allowance of 20 browser scans and five distinct source models per calendar month. 7B-class repositories run as single-model quantization simulations. A comparison of two 7B-class checkpoints is currently accepted by the… See the full description on the dataset page: https://huggingface.co/datasets/tetracta/llm-xray-lesion-scans.tabularothern<1K1 likes241 downloads3d agoHugging Face26kardosdrur /scandi_eurovoctext100K<n<1M0 likes219 downloads3y agoHugging Face27zr-zhou /Scannetpp3dn<1K0 likes215 downloads3mo agoHugging Face28alexandrainst /scandi-qaScandiQA is a dataset of questions and answers in the Danish, Norwegian, and Swedish languages. All samples come from the Natural Questions (NQ) dataset, which is a large question answering dataset from Google searches. The Scandinavian questions and answers come from the MKQA dataset, where 10,000 NQ samples were manually translated into, among others, Danish, Norwegian, and Swedish. However, this did not include a translated context, hindering the training of extractive question answering models. We merged the NQ dataset with the MKQA dataset, and extracted contexts as either "long answers" from the NQ dataset, being the paragraph in which the answer was found, or otherwise we extract the context by locating the paragraphs which have the largest cosine similarity to the question, and which contains the desired answer. Further, many answers in the MKQA dataset were "language normalised": for instance, all date answers were converted to the format "YYYY-MM-DD", meaning that in most cases these answers are not appearing in any paragraphs. We solve this by extending the MKQA answers with plausible "answer candidates", being slight perturbations or translations of the answer. With the contexts extracted, we translated these to Danish, Swedish and Norwegian using the DeepL translation service for Danish and Swedish, and the Google Translation service for Norwegian. After translation we ensured that the Scandinavian answers do indeed occur in the translated contexts. As we are filtering the MKQA samples at both the "merging stage" and the "translation stage", we are not able to fully convert the 10,000 samples to the Scandinavian languages, and instead get roughly 8,000 samples per language. These have further been split into a training, validation and test split, with the former two containing roughly 750 samples. The splits have been created in such a way that the proportion of samples without an answer is roughly the same in each split.textquestion-answering10K<n<100K8 likes188 downloads4y agoHugging Face29Lots-of-LoRAs /task131_scan_long_text_generation_action_command_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.texttext-generation1K<n<10K0 likes185 downloads2y agoHugging Face30YiquanLi /ScanNet_for_ScanQA_SQA3Dtext1K<n<10K1 likes183 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.