datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.betty-dota2
Betty Dota 2 — Decision Context Dataset
Overview
9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives.
Built to train Transformer and RL models that understand the game state at each moment in time.
Dataset Structure
matches.parquet — 9,385 rows
One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.bybit-linear-perps-dotusdtcsharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
dot-airline-ontime
DOT Airline On-Time Performance 2018–2024
45,968,068 reported flight records across all 84 months, with all 109 BTS source fields, exact delay values and separate cancellation/diversion outcomes.
This repository contains the 1,000-row public sample, covering all 84 months.
The full package is a one-time $99 snapshot, with monthly CSV and Parquet files.
It has 111 Parquet columns: 109 BTS fields plus source_month and source_row.
The deterministic sample uses 12 evenly spaced rows… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/dot-airline-ontime.dotnet-runtime
.NET Runtime Fine-Tuning Data and Index
This directory contains data for fine-tuning models and building RAGs for the dotnet/runtime repository.
Overview
data/: Contains all datasets and indexes.
raw/sample/: Sample PRs and diffs collected from GitHub.
raw_data.tar: Archive of collected PRs and diffs from GitHub.
samples/: Json files with processed samples suitable for dataset generation.
processed/: Parquet files for fine-tuning (e.g., train.parquet, test.parquet).… See the full description on the dataset page: https://huggingface.co/datasets/kotlarmilos/dotnet-runtime.dotting-test
Dotting Test
Dotting is a glyph-level benchmark for Turkish text in AI-generated images. It tests whether image
models preserve the dotless ı and other Turkish diacritics at the pixel level.
This package is generated from the Dotting project outputs for fge-auto/dotting-test.
Creator: Fırat Gelbal.
Released under Creative Commons Attribution 4.0 International (CC BY 4.0).
Attribution should credit Fırat Gelbal and Dotting Test.
Related Links
Project site:… See the full description on the dataset page: https://huggingface.co/datasets/fge-auto/dotting-test.msmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.dot-Generic-100k
dot-Generic-100k
A frozen evaluation suite from DoTime.
Episodes: 100000
Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata.
Load with:
from dotime.benchmarks import load_benchmark
suite = load_benchmark("dot-Generic-100k") # pulls this repo at tag v1.0.0
Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable
archive of record.
d_otherschema-dot-org
Geolocated text from the Web Data Commons schema.org GeoCoordinates subset
12,427,530 geolocated text records, extracted from the class-specific
GeoCoordinates subset of the Web Data Commons schema.org data set series
(release 2024-12). Each record pairs one coordinate pair published on a web page
with the text published next to it on that same page.
Each source stream is deduplicated by host-local runs: a coordinate-and-name
pair is kept once per contiguous host run. A host… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/schema-dot-org.dota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.dot-Identifiability-v1
dot-Identifiability-v1
A frozen evaluation suite from DoTime.
Episodes: 10800
Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata.
Load with:
from dotime.benchmarks import load_benchmark
suite = load_benchmark("dot-Identifiability-v1") # pulls this repo at tag v1.1.0
Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable
archive of record.
red-dotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/red-dot.fiction_dot_live
Fiction.live Public Stories
This dataset contains public story metadata and story text from Fiction.live, exported as zstd-compressed Parquet files. The collection covers active, finished, and hiatus stories across teen, mature, unrated, and NSFW content ratings.
The dataset includes adult and user-generated content. Downstream users should filter by content_rating, tags, and story metadata as appropriate for their use case.
Files
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/fiction_dot_live.lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 562160,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000.dot-Continuous-v1
dot-Continuous-v1
A frozen evaluation suite from DoTime.
Episodes: 9999
Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata.
Load with:
from dotime.benchmarks import load_benchmark
suite = load_benchmark("dot-Continuous-v1") # pulls this repo at tag v1.0.0
Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable
archive of record.
dot-RegimeSwitch-v1
dot-RegimeSwitch-v1
A frozen evaluation suite from DoTime.
Episodes: 9999
Schema: parquet shards + manifest.json (md5-checksummed), Croissant metadata.
Load with:
from dotime.benchmarks import load_benchmark
suite = load_benchmark("dot-RegimeSwitch-v1") # pulls this repo at tag v1.0.0
Generated reproducibly by scripts/build_release.py. Zenodo DOI is the citable
archive of record.
lekiwi_lego_dotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 59,
"total_frames": 31492,
"total_tasks": 1,
"total_videos": 59,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:59"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/lekiwi_lego_dot.lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 100,
"total_frames": 55262,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100.lego_mimicgen_5_6block_dot_subgoal_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 938696,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_mimicgen_5_6block_dot_subgoal_lerobot_v3_1000.dot_collect_empty_bottle_black_white_wrist_200k_bs8_lba0_testing1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 1320,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/dot_collect_empty_bottle_black_white_wrist_200k_bs8_lba0_testing1.dot_collect_empty_bottle_black_white_wrist_200k_bs8_lba0_testing2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 1303,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/dot_collect_empty_bottle_black_white_wrist_200k_bs8_lba0_testing2.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.lastfm-64-dot
Dataset Overview
dataset: lastfm-64-dot
Metadata
Creation Time: 2025-01-06 11:09:48+0000
Update Time: 2025-01-07 11:48:10+0000
Source: https://github.com/erikbern/ann-benchmarks
Task: N/A
Train Samples: N/A
Test Samples: N/A
License: DISCLAIMER AND LICENSE NOTICE:
This dataset is intended for benchmarking and research purposes only.
The source data used in this dataset retains its original license and copyright. Users must comply with the respective licenses of the… See the full description on the dataset page: https://huggingface.co/datasets/open-vdb/lastfm-64-dot.dot_collect_empty_bottle_black_white_wrist_200k_bs8_lba0_testing0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 1304,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/dot_collect_empty_bottle_black_white_wrist_200k_bs8_lba0_testing0.laser_routing_dot_points_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_laser_routing",
"total_episodes": 100,
"total_frames": 16928,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/laser_routing_dot_points_lerobot_v3_100.steer_place_yellow_dot_red_arrow_example_ep201
Placement Yellow-Dot Red-Arrow Sanity Dataset
This is a LeRobot-format one-episode sanity export derived from local HDF5 placement data.
It uses source recording episode_00201.hdf5, one of the five episodes newer
than the existing 197-episode placement export.
Dataset size:
episodes: 1
frames: 259
videos: 3
export fps: 100
frame stride from 100 Hz source: 1
source episode index: 201
Source goal label:
full-resolution target: (564.0, 223.0) px
224x224 overlay target: (98.7… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/steer_place_yellow_dot_red_arrow_example_ep201.lego_mimicgen_5_6block_dot_subgoal_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 100,
"total_frames": 94036,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_mimicgen_5_6block_dot_subgoal_lerobot_v3_100.
