datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiktok-hooks-finetune
Tiktok Caption and Hook Dataset
Grabbed the initial dataset from https://x.com/iamgdsa/status/1884294758484611336
Ran quick language classification atop it (probably is bad, but it gets the job done) , and created 3 new conversation columns:
conversations - based on given input variables, generate a full set of caption + hook
conversations_caption - based on given input variables including hook, generate a caption
conversations_hook - based on given input variables including… See the full description on the dataset page: https://huggingface.co/datasets/benxh/tiktok-hooks-finetune.nova-dataset-finetune-2
Nova: Voice-to-Text Companion Dataset
Dataset Description
This dataset contains fine-tuning data for Nova, a real-time voice-to-text companion that can transcribe, translate, and assist with various voice-controlled workflows.
Dataset Summary
Total Conversations: 845 conversation objects
Total Messages: 4,605 individual messages
Languages: English, Vietnamese, and multilingual support
Format: Structured conversations with tool calls and responses
Use Case:… See the full description on the dataset page: https://huggingface.co/datasets/Buiilding/nova-dataset-finetune-2.
