mimi
Datasets
All datasets matching “mimi”fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.mimicgen_datasets
Dataset Card for MimicGen Datasets
Dataset Summary
This repository contains the official release of datasets for the CoRL 2023 paper "MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations".
The datasets contain over 48,000 task demonstrations across 12 tasks, grouped into the following categories:
source: 120 human demonstrations across 12 tasks used to automatically generate the other datasets
core: 26,000 task demonstrations… See the full description on the dataset page: https://huggingface.co/datasets/amandlek/mimicgen_datasets.mimimimic-shards-v2
mimic-shards-v2
Pre-built training shards for MIMIC —
behavior-cloned Super Smash Bros. Melee bots. Derived from the raw ranked
replays in erickfm/melee-ranked-replays; this repo is the cached tensor
form so training pods never rebuild from .slp.
Layout
One folder per character (23 total: 22 roster characters + fox/). Each
folder is directly usable as a MIMIC --data-dir:
<char>/
train_shard_*.pt # per-game v2 tensor shards
val_shard_*.pt # 10%… See the full description on the dataset page: https://huggingface.co/datasets/erickfm/mimic-shards-v2.MIMICIT
Bo Li*,♠,1
Yuanhan Zhang*,♠,1
Liangyu Chen*,1
Jinghao Wang*,1
Fanyi Pu*,1
Jingkang Yang1
Chunyuan Li2
Ziwei Liu✉,1
1S-Lab, Nanyang Technological University
2Microsoft Research, Redmond
♠Co-Project Lead
* Equal Contribution
✉ Corresponding Author
Note 1: To reduce memory consumption during image loading and improve loading speed, we are converting the JSON format of images to the Parquet format. For… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/MIMICIT.mimicgen-aligned-gt-depth
MimicGen Aligned GT Depth Sidecars
Per-frame ground-truth simulator depth for the MimicGen demonstration set,
aligned to the original RGB demo frames. Generated by replaying each demo's
HDF5 simulator states with robosuite + MuJoCo EGL rendering.
Producer: scripts/export_mimicgen_aligned_gt_depth.py in the 3DA_unified
training repo. Original MimicGen data: https://mimicgen.github.io/ (CC BY 4.0).
Layout
26 tasks x ~1000 demos = ~26000 NPZ files, all at the repo root.… See the full description on the dataset page: https://huggingface.co/datasets/SeonghuJeon/mimicgen-aligned-gt-depth.
