datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TokenShrink-OCR
TokenShrink-OCR Dataset
Introduction
This is a large-scale dataset containing 120,000 images, designed for Optical Character Recognition (OCR) tasks.
All images are derived from the ImageNet database, providing a challenging collection of text against complex backgrounds, varied lighting conditions, and diverse fonts.
Dataset Structure
All image files are stored in a sharded structure.
All data (train, validation, test) has been split into small… See the full description on the dataset page: https://huggingface.co/datasets/LukB4UJump/TokenShrink-OCR.classify_tokens_dataset_demo_reThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 30,
"total_frames": 23081,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dogeum/classify_tokens_dataset_demo_re.classify_tokens_dataset_demoaugmented_prompts_images_77_tokens
Dataset Card for "augmented_prompts_images_77_tokens"
More Information needed
MoulSot-Tokens-v1
MoulSot-Tokens-v1
Discrete speech tokens for Moroccan Darija. 101 hours of transcribed speech,
encoded to a single-codebook neural audio codec and paired with text, ready to
train a text-to-speech model that predicts tokens directly.
Size
Source audio (16 kHz parquet)
~12 GB
This dataset (tokens, text)
153 MB
Same speech, ~78× smaller. At the codec level that is 256 kbps of PCM
reduced to 0.8 kbps — a 320× reduction in bits — since in this case, one second of… See the full description on the dataset page: https://huggingface.co/datasets/Tilas/MoulSot-Tokens-v1.guide-tokens-v1-8k1kaugmented_images_40_tokens
Dataset Card for "augmented_images_40_tokens"
More Information needed
rubber_duck_extended_tokensscicap-caption-no-more-than-100-tokens-no-subfigscicap-caption-no-more-than-100-tokens-yes-subfigdesign-tokens-datasetstack_tokens_datasetrubber_duck_tokensposition-conditioning-4K-with-class-tokens
