datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NanoMTEB-Scandinavian
NanoMTEB-Scandinavian
This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoMTEB-Scandinavian is a compact retrieval benchmark for Scandinavian-language MTEB-style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social-media retrieval.
Usage
from datasets import load_dataset
dataset_id = "hakari-bench/NanoMTEB-Scandinavian"
split… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Scandinavian.Scandium-Dataset
Dataset Card — Scandium-Dataset v1.0.0
Summary
Scandium-Dataset provides a harmonized, quality-scored foundation of DFT-computed structural and thermodynamic properties across 267,230 materials from Materials Project, OQMD, and JARVIS-DFT. It supports the early screening stage of battery materials discovery — filtering by phase stability, electronic structure, and structural family — before downstream property prediction (ionic conductivity, mechanical stability… See the full description on the dataset page: https://huggingface.co/datasets/Scandium-Labs/Scandium-Dataset.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
scandi-reddit
Dataset Card for ScandiReddit
Dataset Summary
ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit.
All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept.
The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.scandisentRoboTwin_scan_object_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 80479,
"total_tasks": 499,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_scan_object_randomized.scanqa_images_16_keyframes_120_non_keyframes_min_532_long_edgescandinavian_faroesectr-scan-object-uniform50-20260917
IdleMask review — passed (2026-09-19)
Reviewed by the dataset owner: observation.arm_active_mask is correct and this revision is a formally usable CTR dataset. This section supersedes previous active/idle-mask descriptions below.
For each of the 50 episodes and each physical arm, only the initial contiguous scheduling delay may have mask 0. From first duty through the final frame the mask is always 1. Scan synchronization waits, cooperative holds, the shared scan tail and… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-scan-object-uniform50-20260917.SCAND_traj_selectionarabic-ocr-synthetic-scans-faker-300k
Arabic OCR Synthetic Scans (Faker 300k)
A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness.
Dataset Summary
Samples: ~300,000 synthetic Arabic document pages
Image format: JPEG, ~800×1200 px (embedded in Parquet)
Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.europeana-it-scans_beirThis is a copy of https://huggingface.co/datasets/jinaai/europeana-it-scans reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/europeana-it-scans_beir.vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
Description
VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0
mateoguaman/iphone_stairs_ramps: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5.ur5e_scanner_fp_deltaTCPThis dataset was created using LeRobot format.
Dataset Structure
{
"codebase_version": "v2.1",
"robot_type": "ur5e",
"total_episodes": 60,
"total_frames": 31577,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:59"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ns69956/ur5e_scanner_fp_deltaTCP.scandi_eurovocscanobjectnn
Dataset Card for "scanobjectnn"
More Information needed
task131_scan_long_text_generation_action_command_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.scand_every1_50pct_sub5
scand_every1_50pct_sub5
Description
Processed scand dataset with filter_every_nth=1, 50% of data, and num_subsampled_points=5
Processing Parameters
mateoguaman/scand:
exclude_outliers_pct: 3
filter_by_curvature: true
filter_every_nth: 1
horizon:
300: 0.5
500: 0.5
num_subsampled_points: 5
Dataset Configuration
Train dataset:
mixer: mateoguaman/scand: 0.5
split: train
Validation dataset:
mixer: mateoguaman/scand:… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/scand_every1_50pct_sub5.ctr-scan-object-uniform100-20260921
scan_object Uniform100
Open in LeRobot Dataset Visualizer
100 episodes from the same 50 source scene seeds, 25 FPS, LeRobot v3, Aloha
AgileX. Parent: Shiki42/ctr-scan-object-uniform-20260916 at 766e18802a11bdde94f6dab6892c13d82e6a5bd3
(repaired gripper target labels). This is a byte-exact republication of the
pinned parent for the non-mainline 100-sample comparison arm; the 50-episode
arm remains ctr-scan-object-uniform50-20260917.
Every source seed appears exactly twice, as… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-scan-object-uniform100-20260921.vlmn_scand_spot_sub5
vlmn_scand_spot_sub5
Description
VLN Navigation dataset with 50% of scand data and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/scand_every1_50pct_sub5: 1.0
mateoguaman/spot_every1_sub5: 1.0
split: train
Validation dataset:
mixer: mateoguaman/scand_every1_50pct_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_scand_spot_sub5.scandinavian-embedding-training-datascannet_countingscanned-arxiv-papers-idvlmn_tartandrive100_scand50_coda25_spot100_sub5_filtered_trajectories_training_25_fixedSCAND_path_selectiontask127_scan_long_text_generation_action_command_all
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task127_scan_long_text_generation_action_command_all
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task127_scan_long_text_generation_action_command_all.ur5e_scanner_tcp_TRAINThis dataset was created using LeRobot format.
Dataset Structure
{
"codebase_version": "v2.1",
"robot_type": "ur5e",
"total_episodes": 50,
"total_frames": 24200,
"total_tasks": 2,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:49"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ns69956/ur5e_scanner_tcp_TRAIN.danish_culturax
Danish Culturax Dataset
This dataset is simply a reformatting of uonlp/CulturaX. Some minor formatting errors have been corrected.
Usage
from datasets import load_dataset
dataset = load_dataset("ScandLM/danish_culturax")
scandinavian-100hscanned-arxiv-papers
