datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tinyllama_pretrained_datatiny_llavavideoTinyLLaVA-Video
This dataset combines data from multiple sources for pre-training and fine-tuning.
Pretrain Data: Four subsets of LLaVA-Video-178K (0_30_s_academic_v0_1, 30_60_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_youtube_v0_1), supplemented with filtered Video-LLaVA data (https://huggingface.co/datasets/LanguageBind/Video-LLaVA) and data from Valley (https://github.com/RupertLuo/Valley). The video data can be downloaded from the linked datasets, and cleaned annotations are provided… See the full description on the dataset page: https://huggingface.co/datasets/pbwpbw/tiny_llavavideo.TinyStoriesChineseTinyStories数据集的中文翻译版。只翻译了story字段(翻译后字段为story_zh):
{
"story": "\n\nLily and Ben are friends. They like to play in the park. One day, they see a big tree with a swing. Lily wants to try the swing. She runs to the tree and climbs on the swing.\n\"Push me, Ben!\" she says. Ben pushes her gently. Lily feels happy. She swings higher and higher. She laughs and shouts.\nBen watches Lily. He thinks she is cute. He wants to swing too. He waits for Lily to stop. But Lily does not stop. She swings… See the full description on the dataset page: https://huggingface.co/datasets/adam89/TinyStoriesChinese.comix_v0_tiny_pages
Comic Books Tiny Dataset v0 - Pages (Testing)
Small test dataset of comic book pages for rapid development and testing.
⚠️ This is a TINY dataset for testing only. For production, use comix_v0_pages.
What's Included
Each page has:
{page_id}.jpg - Page image
{page_id}.json - Metadata (detections, captions, page class)
{page_id}.seg.npz - Segmentation masks (SAMv2)
Quick Start
from datasets import load_dataset
import numpy as np
# Load tiny pages dataset
pages… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix_v0_tiny_pages.comix-v0_1-books-tiny
Comic Books Dataset v0.1 - Books
Full dataset of book-level metadata from Digital Comic Museum.
This is the PRODUCTION dataset. For testing, use comix_v0_tiny_books.
What's Included
Each book has:
{book_id}.json - Book metadata with page references
Purpose
This dataset provides book-level metadata to group pages from comix-v0_1-pages.
Workflow:
Download comix-v0_1-pages (with images)
Download comix-v0_1-books (metadata only)
Use WebDataset pipeline to group… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix-v0_1-books-tiny.SFHQ-Tiny-512-Part1TinyStoriesChineseTinyStories数据集的中文翻译版。只翻译了story字段(翻译后字段为story_zh):
{
"story": "\n\nLily and Ben are friends. They like to play in the park. One day, they see a big tree with a swing. Lily wants to try the swing. She runs to the tree and climbs on the swing.\n\"Push me, Ben!\" she says. Ben pushes her gently. Lily feels happy. She swings higher and higher. She laughs and shouts.\nBen watches Lily. He thinks she is cute. He wants to swing too. He waits for Lily to stop. But Lily does not stop. She swings… See the full description on the dataset page: https://huggingface.co/datasets/78yang/TinyStoriesChinese.laion-audio-tinytinystories_1300kcomix_v0_tiny_books
Comic Books Tiny Dataset v0 - Books (Testing)
Small test dataset of book-level metadata for rapid development and testing.
⚠️ This is a TINY dataset for testing only. For production, use comix_v0_books.
What's Included
Each book has:
{book_id}.json - Book metadata with page references
Purpose
This dataset provides book-level metadata to group pages from comix_v0_tiny_pages.
Workflow:
Download comix_v0_tiny_pages (with images)
Download comix_v0_tiny_books… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix_v0_tiny_books.protenix_tiny_rna_emb_15508audiosnippets-tiny
Tiny audio snippets
