datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bartholomew-dataset-v3
BART Dataset v3
The final version of the BART pretraining corpus, focused on removing anything that betrays a
post-1930 origin. This is our best vintage dataset yet.
Documents
146,031 (97.52% of v2)
Characters
102,798,688,961 (96.73% of v2)
Tokens
~23B (estimated)
Shards
473 (one per v2 shard, same basename)
Source
BART Dataset v2
Cutoff
1930
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v3.bartholomew-dataset-v1
BART Dataset v1
The first version of the BART pretraining corpus: pre-1930 English books drawn from
Institutional Books 1.0
and filtered hard on OCR quality, language, date, and tokenizability.
Documents
160,263
Characters
118,745,375,871
Tokens
~27B (estimated)
Shards
473 (472 train + 1 val)
Source
Institutional Books 1.0 (242B tokens, ~983K documents)
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v1.variouscryptodata
variouscryptodata
Crypto market datasets collected as a by-product of our own research and
published so they are not lost. One sub-folder per dataset; each appended
nightly where collection is still running.
folder
what
coverage
cadence
polymarket_updown_orderbook/
Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference
2026-05-24 → present
appended nightly (previous UTC day)
hyperliquid_trades/
Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.bartholomew-dataset-v2
BART Dataset v2
The second version of the BART pretraining corpus, focused on stripping low-quality text —
boilerplate and OCR corruption — out of
v1.
Documents
149,745 (93.44% of v1)
Characters
106,274,384,672 (89.50% of v1)
Tokens
~24B (estimated)
Shards
473 (one per v1 shard, same basename)
Source
BART Dataset v1
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:
Institutional Books 1.0 →
v1 →
v2 →
v3… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v2.EventHubDatasetspeech_commands
Dataset Card for "speech_commands"
More Information needed
vcc2026-broad-development-20260908
VCC 2026 broad development artifacts
The B7 RunPod science-output recovery is complete: 256 files, 5,106,104,214 bytes.
Contents
Files
Checkpoints
64
Fit receipts
64
Prediction arrays
64
Prediction receipts
64
Use archive-index.json for current locations, SHA-256 values, sizes, and immutable per-file revisions. STATUS.md records completed work and remaining scientific comparisons. The recovery receipt proves all 109 files missing from the previous archive… See the full description on the dataset page: https://huggingface.co/datasets/barthazian/vcc2026-broad-development-20260908.bart-midtrain
BART Midtrain
The midtraining corpus for BART —
pre-1930 mathematics, science, technology, and medicine — plus the full pipeline that built it and
every training mixture it was blended into.
Built by Unbounded Labs.
Corpus documents
11,409
Corpus characters
2,543,809,124
Corpus tokens
~604M
Removed by cleaning
24% of documents (15,075 → 11,409)
Subject focus
math, science, technology, medicine
Cutoff
1930
Schema
single string column text… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/bart-midtrain.rayst3rdice2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 3660,
"total_tasks":1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/dice2.tape_to_bin2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 4896,
"total_tasks":1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/tape_to_bin2.tape_to_binThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 739,
"total_tasks":1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/tape_to_bin.dice4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 6198,
"total_tasks":1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/dice4.bart-tokenizer-tryout
build_dataset.py
Dataset Summary
A nlp sentiment dataset with pointcloud text modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: curriculum
Augmentation: randaugment
Splits & Sampling
Split strategy: leave one out
Sampling: curriculum
Quality & Labeling
Quality filtering: adaptive
Labeling: weak supervision
Files
build_dataset.py — main artifact of this repository… See the full description on the dataset page: https://huggingface.co/datasets/lukasm-uel0417/bart-tokenizer-tryout.bartenderkaminoglass
Bangumi Image Base of Bartender: Kami No Glass
This is the image base of bangumi Bartender: Kami no Glass, we detected 26 characters, 3350 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/bartenderkaminoglass.SynGallery-1024
SynGallery-1024: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition
The 1024×1024 high-resolution edition of
patryk-bartkowiak/SynGallery.
A synthetic dataset for instance-level artwork recognition: 4,898 real
paintings (MET Open Access) hung in a procedurally randomized 3D art-gallery
scene, each rendered from 5 camera viewpoints at 1024×1024 — 24,490
synthetic RGB images paired with their source photos and museum metadata
(title, artist, date, medium… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-1024.bartowski-imatrix-v5-semantic
Bartowski iMatrix Calibration v5 (Semantic Chunking)
A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure.
Dataset Summary
Metric
Value
Total samples
2,075
Chunking method
V5-optimized semantic boundary detection
Chunk size
200+ characters (no upper limit, preserves document integrity)
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.bart-robot-arm
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/josvdwest/bart-robot-arm.Leo__bart-large__1645784880GEM__bart_base_schema_guided_dialog__1645547915SynGallery-abl4-tex-light-glass-frame
SynGallery-abl4-tex-light-glass-frame: + frame variety
Rung 4 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting, glass and frame molding variant + color/roughness/metallic, while freezing camera pose (the only frozen factor). Same schema, source images and index↔painting… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl4-tex-light-glass-frame.summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1SynGallery-abl0
SynGallery-abl0: fixed-environment baseline
Rung 0 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies nothing — the gallery is one fixed configuration for all 24,490 images, while freezing wall/floor/roof textures + floor material, lighting, glass, frame variant/color, camera pose. Same schema… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl0.hard_pick_and_place_45This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 25,
"total_frames": 5836,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bartek-niedzielski/hard_pick_and_place_45.details_bartowski__internlm2-chat-20b-llama
Dataset Card for Evaluation run of bartowski/internlm2-chat-20b-llama
Dataset automatically created during the evaluation run of model bartowski/internlm2-chat-20b-llama on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bartowski__internlm2-chat-20b-llama.SynGallery
SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition
A synthetic dataset for instance-level artwork recognition: 4,898 real
paintings (MET Open Access) hung in a procedurally randomized 3D art-gallery
scene, each rendered from 5 camera viewpoints at 512×512 — 24,490
synthetic RGB images paired with their source photos and museum metadata
(title, artist, date, medium, …). The environment is randomized per scene — wall/floor/roof textures… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery.object-placing-on-Sutton-Barto_20260916_201400This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ssabats/object-placing-on-Sutton-Barto_20260916_201400.eval_dice2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 3840,
"total_tasks":1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/eval_dice2.SynGallery-abl1-tex
SynGallery-abl1-tex: + texture variety
Rung 1 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies wall/floor/roof textures + floor material (on top of the baseline), while freezing lighting, glass, frame variant/color, camera pose. Same schema, source images and index↔painting mapping as every… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl1-tex.hard_pick_and_place_70This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 25,
"total_frames": 5836,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bartek-niedzielski/hard_pick_and_place_70.
