datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iclr-wm-backup-public
ICLR Watermark Benchmark — backup overflow (public part)
Companion to the private repo Aak975/iclr-wm-backup, which reached its
storage quota. Together the two repos form ONE backup — every file exists in
exactly one of them, with the same layout:
archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json
restore one archive: cat part-* | zstd -d | tar -x
MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M}
This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.ft49
ft49 — baselining backup
Backup of the ft49 baselining working directory (Monash M3), created 2026-09-03.
facellm/ — the facellm/ tree (76,427 files, 315 MB raw) as a single zstd-compressed tar
stream: part-0000 + MANIFEST.json ({"parts", "sha256" (whole stream), "total_bytes"}).
phase1_output/appearance_action_hair_00400_unfrozen.zip — 3.3 GB, stored as-is.
phase2_output/htcc_phase2_attributes_output.zip — 24.7 GB, stored as-is.
Restore
python3 -m pip install… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/ft49.udposUniversal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers.angle_peg_stereo_merged_07_29_stateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/angle_peg_stereo_merged_07_29_state.mysdxl-dataset
Image-Prompt Dataset
An image-prompt dataset scraped and assembled with
MySDXL for training
latent diffusion models.
Dataset structure
Each row contains one image with its corresponding text prompt.
Column
Type
Description
image
Image
RGB image (lossless PNG, original resolution)
prompt
string
Text prompt describing the image
negative_prompt
string
Negative prompt (empty string if none)
Stored as Parquet shards (data/train-*.parquet).
Load… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/mysdxl-dataset.vaigai-dataset
Vaigai Dataset
aakkaasshh/vaigai-dataset is an Indic-focused multilingual text corpus aggregated for tokenizer training, built with the goal of reaching tokenizer quality on par with dedicated Indic tokenizers (e.g. Sarvam AI's).
One Parquet file per language is stored at the root of this repo (e.g. bn.parquet), with no nested folders.
Every time new data is fetched for a language, it is merged with that language's existing file and deduplicated on the text column, so re-running… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/vaigai-dataset.sindhi-corpus-505m
Sindhi Corpus 505M
The largest open-source, deduplicated Sindhi language pretraining corpus.
~505 million tokens across 742K documents, covering news, literature, legal, religious, encyclopedic, and web-crawled Sindhi text. Built for training Sindhi language models, tokenizers, and NLP tools.
Dataset Summary
Stat
Value
Documents
~742,379
Tokens (estimated)
~505 million
Language
Sindhi (sd) — Arabic script
Format
Parquet (single text column)… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/sindhi-corpus-505m.eval_so101-pick-cube-v2-fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 89373,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aakashv100/eval_so101-pick-cube-v2-fixed.eval_so101-pick-cube-v2-randomThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 89414,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aakashv100/eval_so101-pick-cube-v2-random.cdr_bigbio_processedautotrain-data-auto-nlp-poc
AutoTrain Dataset for project: auto-nlp-poc
Dataset Description
This dataset has been automatically processed by AutoTrain for project auto-nlp-poc.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"context": "Simultaneously with these conflicts, bison, a keystone species and the primary protein source that Native people had survived on for… See the full description on the dataset page: https://huggingface.co/datasets/aak7912/autotrain-data-auto-nlp-poc.colpali_train_set
Dataset Description
This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%).
Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages.
Dataset
#examples (query-page pairs)
Language
DocVQA… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/colpali_train_set.so101-pick-place-positions
SO-101 Pick-and-Place (Multi-Position)
Teleoperated pick-and-place demonstrations on the SO-ARM 101 (follower + leader), recorded with LeRobot v3.0.
Task: pick a cube from a numbered sheet position and place it in a fixed box.
Summary
Field
Value
Robot
SO-101 follower (so_follower)
Control
Human teleoperation (SO-101 leader)
Episodes
50
Frames
60,186
FPS
30
Cameras
gripper_cam, top_cam (640×480, H.264)
Action / state
6-DoF joint… See the full description on the dataset page: https://huggingface.co/datasets/aakashv100/so101-pick-place-positions.sindhi-tokenized-505msilicone_rope_pull_up_stereoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/silicone_rope_pull_up_stereo.angle_peg_stereo_06_10_base_to_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/angle_peg_stereo_06_10_base_to_cam.tissue_retraction_07_08_stateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/tissue_retraction_07_08_state.Sindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.sindhi-corpus-langid-cleanjeeangle_blackpeg_stereo_07_07_base_to_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/angle_blackpeg_stereo_07_07_base_to_cam.tissue_retraction_stereo_07_03_base_to_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/tissue_retraction_stereo_07_03_base_to_cam.angle_peg_stereo_06_08_base_to_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/angle_peg_stereo_06_08_base_to_cam.angle_peg_stereo_05_27_base_to_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/angle_peg_stereo_05_27_base_to_cam.silicone_rope_pull_up_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/silicone_rope_pull_up_merged.tissue_retraction_stereo_07_08_base_to_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/tissue_retraction_stereo_07_08_base_to_cam.silicone_rope_pull_up_07_07_stereoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/silicone_rope_pull_up_07_07_stereo.tissue_retraction_06_30_stateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/tissue_retraction_06_30_state.tissue_retraction_07_03_stateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/tissue_retraction_07_03_state.eval_so101-pick-cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aakashv100/eval_so101-pick-cube.
