datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stem-videosDataset Description:
This dataset is a large-scale collection of STEM educational video data, containing 100,000 hours of video content, designed to support the development and training of advanced video understanding, multimodal AI, vision-language models (VLMs), educational AI systems, video captioning, video reasoning, and large-scale machine learning applications.
Additionally, this dataset can be used in pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/stem-videos.rollout_stem_grasp_corrections
