datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretokenized-dolma
The Pretokenized Dolma Dataset
A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library.
Overview
Key Features:
Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280
Sequence length: 2049 tokens (2048 + 1 for next-token prediction)
Sharded into 10,000 Parquet files (~78MB each)
420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.pico-banana-smolvlm-format-with-rejected-answer
pico-banana-smolvlm-format-with-rejected-answer
Balanced image-level tampering detection dataset in SmolVLM-style format
with chosen/rejected answer pairs, derived from the pico-banana MCQ
pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training.
Dataset overview
Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a
rejected_answer field: the answer from the counterpart sample (same
edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.pico-8-games
PICO-8 Games Dataset
The first multimodal dataset of PICO-8 games. 10,967 cartridges scraped from the Lexaloffle BBS, each decomposed into Lua source code, pixel-art spritesheets, tile maps, sound effects, music patterns, and metadata.
Label screenshots from the top 48 games by star count
What's Inside
Every PICO-8 cartridge is a self-contained game packed into a single file. This dataset cracks each one open into its component parts:
The… See the full description on the dataset page: https://huggingface.co/datasets/Fraser/pico-8-games.picorpusrapberry_pi_pico_all-dataset
🤗 rapberry_pi_pico_all
This dataset was automatically generated and verified using the Universal PDF & Rendergit Code Dataset Generator Pipeline (Phase 1-4).
It contains high-quality synthetic code pairs, technical SFT Q&A, DPO (Direct Preference Optimization) preference pairs, and multi-turn technical chat sequences in both English and Turkish.
📊 Dataset Summary & Splits
tr_dpo_dataset.jsonl: 756 örnek (samples)
tr_sft_dataset.jsonl: 8,865 örnek (samples)… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/rapberry_pi_pico_all-dataset.franka-revo2-pico4-hand-demo-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka_research3_dexhand",
"total_episodes": 7,
"total_frames": 1293,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/franka-revo2-pico4-hand-demo-test.franka-pico4-red-blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"tcp.x",
"tcp.y",
"tcp.z",
"tcp.r1",
"tcp.r2",
"tcp.r3",
"tcp.r4",
"tcp.r5",
"tcp.r6",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-pico4-red-block.picoctf
PicoCTF Challenges
This dataset was originally only available on GitHub under agpl-3.0 license. I ported it to Hugging Face after making small corrections.
franka-pico4-yellow-blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"tcp.x",
"tcp.y",
"tcp.z",
"tcp.r1",
"tcp.r2",
"tcp.r3",
"tcp.r4",
"tcp.r5",
"tcp.r6",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-pico4-yellow-block.franka-pico4-blue-blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"tcp.x",
"tcp.y",
"tcp.z",
"tcp.r1",
"tcp.r2",
"tcp.r3",
"tcp.r4",
"tcp.r5",
"tcp.r6",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-pico4-blue-block.b601-pico4-demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "seeed_b601_rt_follower",
"total_episodes": 3,
"total_frames": 1909,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-pico4-demo.tron2-pico4-demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_tcp.x",
"left_tcp.y",
"left_tcp.z",
"left_tcp.r1",
"left_tcp.r2",
"left_tcp.r3",
"left_tcp.r4",
"left_tcp.r5"… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/tron2-pico4-demo.pretokenized-paloma-tinsy
The Tinsy Pretokenized Paloma Dataset
A small version of the pretokenized-paloma benchmark dataset.
This dataset is a sub-sampled version of pretokenized-paloma, and can be used in place of pretokenized-paloma.
We release the exact scripts we use to create this dataset in our pico-lm/pico-dataset GitHub repo.
pretokenized-dolma-tinsy
The Tinsy Pretokenized Dolma Dataset
A tiny little baby-version of the pretokenized-dolma dataset.
Meant to be used in a jupyter notebook to test things out, or quickly look at the structure of the data.
tron2-pico4-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_tcp.x",
"left_tcp.y",
"left_tcp.z",
"left_tcp.r1",
"left_tcp.r2",
"left_tcp.r3",
"left_tcp.r4",
"left_tcp.r5"… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/tron2-pico4-test.b601-pico4-demo1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "seeed_b601_rt_follower",
"total_episodes": 3,
"total_frames": 2167,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-pico4-demo1.b601-bi-pico4-demo-0611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_seeed_b601_rt_follower",
"total_episodes": 4,
"total_frames": 1418,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-bi-pico4-demo-0611.cleaned_ebmnlp_pico
Dataset Card for "cleaned_ebmnlp_pico"
More Information needed
b601-bi-pico4-demo-7camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_seeed_b601_rt_follower",
"total_episodes": 1,
"total_frames": 967,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-bi-pico4-demo-7cam.minipile_density-proportioned_picoProduced with the MiniCorpus pipeline, a reproduction and investigation of the MiniPile data-distillation method (Kaddour, 2023).
This dataset is derived from and based on the contents of The Pile Deduplicated.
ebmnlp_pico
Dataset Card for "ebmnlp_pico"
More Information needed
PICO-breast-cancer
PICO breast cancer dataset
This dataset has been extracted from PICO-Corpus. The corpus consists of 1,011 abstracts of breast cancer randomized controlled
trials extracted from PubMed. The PICO breast cancer dataset contains a total of 26 entities, compared to the usual 4 found in PICO corpora.
Specifically, the following image extracted by the dataset's authors shows the hierarchy of the entities.
The preprocessed dataset, ready to serve as inputs for MLMs such as BERT-like… See the full description on the dataset page: https://huggingface.co/datasets/cuevascarlos/PICO-breast-cancer.unlicense-pico8
Pico-8 Unlicense Games Collection
Dataset Summary
A curated collection of all the games with code under Unlicense tagged PICO-8 from itch.io (29 games as of 2026-04-14). This dataset provides direct access to both the raw cartridge data and Lua source code for each game (extracted from the web player).
Note: The assets are sometimes licensed under non-commercial terms - not everything's free for commercial use. Check the asset_license column for game-specific details.… See the full description on the dataset page: https://huggingface.co/datasets/thomasgauthier/unlicense-pico8.pico_laptop_reachThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_leader",
"total_episodes": 8,
"total_frames": 3054,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/NONHUMAN-RESEARCH/pico_laptop_reach.pretokenized-paloma
The Pretokenized Paloma Benchmark Dataset
This dataset is a compact, pre-tokenized evaluation dataset designed to complement the pretokenized-dolma training set. Built from the Paloma corpus (Allen Institute), this benchmark was designed to not contain any data overlap with Dolma and is ideal for evaluating models trained on it.
Overview
Features:
Pre-tokenized with the same tokenizer as pretokenized-dolma: allenai/OLMo-7B-0724-hf
Sequence length: 2048 tokens
Ideal… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-paloma.seeder_pico_thinking_function_callingpico_covid19pico-human-corpus_nerfair_processed
Benchmark dataset PICO
This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow )
Original dataset: https://github.com/sociocom/PICO-Corpus/tree/main/pico_corpus_brat_annotated_files
pico_ebmnlp
Dataset Card for "pico_ebmnlp"
More Information needed
pico_thinking_function_calling
