CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LukB4UJump /TokenShrink-OCR TokenShrink-OCR Dataset Introduction This is a large-scale dataset containing 120,000 images, designed for Optical Character Recognition (OCR) tasks. All images are derived from the ImageNet database, providing a challenging collection of text against complex backgrounds, varied lighting conditions, and diverse fonts. Dataset Structure All image files are stored in a sharded structure. All data (train, validation, test) has been split into small… See the full description on the dataset page: https://huggingface.co/datasets/LukB4UJump/TokenShrink-OCR.image100K<n<1M0 likes125 downloads11mo agoHugging Face02dogeum /classify_tokens_dataset_demo_reThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 30, "total_frames": 23081, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:30" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dogeum/classify_tokens_dataset_demo_re.imagerobotics10K<n<100K0 likes76 downloads18d agoHugging Face03dogeum /classify_tokens_dataset_demoimage10K<n<100K0 likes47 downloads19d agoHugging Face04rbeauchamp /augmented_prompts_images_77_tokens Dataset Card for "augmented_prompts_images_77_tokens" More Information needed image1K<n<10K0 likes39 downloads3y agoHugging Face05Tilas /MoulSot-Tokens-v1gated MoulSot-Tokens-v1 Discrete speech tokens for Moroccan Darija. 101 hours of transcribed speech, encoded to a single-codebook neural audio codec and paired with text, ready to train a text-to-speech model that predicts tokens directly. Size Source audio (16 kHz parquet) ~12 GB This dataset (tokens, text) 153 MB Same speech, ~78× smaller. At the codec level that is 256 kbps of PCM reduced to 0.8 kbps — a 320× reduction in bits — since in this case, one second of… See the full description on the dataset page: https://huggingface.co/datasets/Tilas/MoulSot-Tokens-v1.imagetext-to-speech10K<n<100K1 likes39 downloads1mo agoHugging Face06arshadshk /guide-tokens-v1-8k1kimage1K<n<10K0 likes36 downloads6mo agoHugging Face07rbeauchamp /augmented_images_40_tokens Dataset Card for "augmented_images_40_tokens" More Information needed imagen<1K1 likes30 downloads3y agoHugging Face08AdrianoC /rubber_duck_extended_tokensimage1K<n<10K0 likes27 downloads2y agoHugging Face09bhalladitya /scicap-caption-no-more-than-100-tokens-no-subfigimagen<1K0 likes22 downloads2y agoHugging Face10bhalladitya /scicap-caption-no-more-than-100-tokens-yes-subfigimagen<1K0 likes22 downloads2y agoHugging Face11agiorchestrator /design-tokens-datasetimagen<1K0 likes12 downloads7mo agoHugging Face12dogeum /stack_tokens_datasetimage1K<n<10K0 likes9 downloads4mo agoHugging Face13AdrianoC /rubber_duck_tokensimagen<1K0 likes5 downloads2y agoHugging Face14kushaaagr /position-conditioning-4K-with-class-tokensimage1K<n<10K0 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.