datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
police-scanner-audio
Police Scanner Audio Dataset
A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds.
Dataset Overview
This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications.
Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.NanoMTEB-Scandinavian
NanoMTEB-Scandinavian
This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoMTEB-Scandinavian is a compact retrieval benchmark for Scandinavian-language MTEB-style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social-media retrieval.
Usage
from datasets import load_dataset
dataset_id = "hakari-bench/NanoMTEB-Scandinavian"
split… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Scandinavian.Scandium-Dataset
Dataset Card — Scandium-Dataset v1.0.0
Summary
Scandium-Dataset provides a harmonized, quality-scored foundation of DFT-computed structural and thermodynamic properties across 267,230 materials from Materials Project, OQMD, and JARVIS-DFT. It supports the early screening stage of battery materials discovery — filtering by phase stability, electronic structure, and structural family — before downstream property prediction (ionic conductivity, mechanical stability… See the full description on the dataset page: https://huggingface.co/datasets/Scandium-Labs/Scandium-Dataset.arabic-synthetic-scanned-booksscannet-processed-testscannetppvlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
scannetpp_processedscandi-reddit
Dataset Card for ScandiReddit
Dataset Summary
ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit.
All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept.
The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.scandisentscannet_samplescannet_temp
scannet_temp
This repository contains split tar.gz archives for the full local scannet dataset tree.
Included
scans/* scenes with unpacked usable content
splits/* split files
Excluded
scene-level original *_2d-*.zip files
results/
color_90/
pose_90.txt
selected_ids_90.txt
Download
huggingface-cli download kairunwen/scannet_temp --repo-type dataset --local-dir ./scannet_temp_hf
Extract
mkdir -p extracted
for f in… See the full description on the dataset page: https://huggingface.co/datasets/kairunwen/scannet_temp.Scannet
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
This repository contains the ScanNet dataset (3D scene data and 2D frame data) and refined annotations used for the paper Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding.
TAB is a dynamic agentic framework designed for zero-shot 3D Visual Grounding (3D-VG). By operating directly on raw RGB-D streams, TAB reformulates 3D… See the full description on the dataset page: https://huggingface.co/datasets/WHB139426/Scannet.ScanReQAscanqa_images_16_keyframes_120_non_keyframes_min_532_long_edgescandinavian_faroeseoutput_3d_bounding_scannetppv2_vllm_old_descriptionabuse-scanner-bot-datasetSCAND_traj_selectionprocessed_scannetscandisent
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/timpal0l/scandisent.arabic-ocr-synthetic-scans-faker-300k
Arabic OCR Synthetic Scans (Faker 300k)
A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness.
Dataset Summary
Samples: ~300,000 synthetic Arabic document pages
Image format: JPEG, ~800×1200 px (embedded in Parquet)
Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.europeana-it-scans_beirThis is a copy of https://huggingface.co/datasets/jinaai/europeana-it-scans reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/europeana-it-scans_beir.vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
Description
VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0
mateoguaman/iphone_stairs_ramps: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5.llm-xray-lesion-scans
VG1 status · 22 September 2026 (Europe/Istanbul)
The 7B class opened on 21 September 2026. Any registered account may scan the repositories on the eligible list of the scope page — exact Apache-2.0 revisions, listed there with their status — within the free allowance of 20 browser scans and five distinct source models per calendar month. 7B-class repositories run as single-model quantization simulations. A comparison of two 7B-class checkpoints is currently accepted by the… See the full description on the dataset page: https://huggingface.co/datasets/tetracta/llm-xray-lesion-scans.scandi_eurovocScannetppscandi-qaScandiQA is a dataset of questions and answers in the Danish, Norwegian, and Swedish
languages. All samples come from the Natural Questions (NQ) dataset, which is a large
question answering dataset from Google searches. The Scandinavian questions and answers
come from the MKQA dataset, where 10,000 NQ samples were manually translated into,
among others, Danish, Norwegian, and Swedish. However, this did not include a
translated context, hindering the training of extractive question answering models.
We merged the NQ dataset with the MKQA dataset, and extracted contexts as either "long
answers" from the NQ dataset, being the paragraph in which the answer was found, or
otherwise we extract the context by locating the paragraphs which have the largest
cosine similarity to the question, and which contains the desired answer.
Further, many answers in the MKQA dataset were "language normalised": for instance, all
date answers were converted to the format "YYYY-MM-DD", meaning that in most cases
these answers are not appearing in any paragraphs. We solve this by extending the MKQA
answers with plausible "answer candidates", being slight perturbations or translations
of the answer.
With the contexts extracted, we translated these to Danish, Swedish and Norwegian using
the DeepL translation service for Danish and Swedish, and the Google Translation
service for Norwegian. After translation we ensured that the Scandinavian answers do
indeed occur in the translated contexts.
As we are filtering the MKQA samples at both the "merging stage" and the "translation
stage", we are not able to fully convert the 10,000 samples to the Scandinavian
languages, and instead get roughly 8,000 samples per language. These have further been
split into a training, validation and test split, with the former two containing
roughly 750 samples. The splits have been created in such a way that the proportion of
samples without an answer is roughly the same in each split.task131_scan_long_text_generation_action_command_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.ScanNet_for_ScanQA_SQA3D
