50m
Datasets
All datasets matching “50m”SenseNova-Vision-Corpus-50M
Vision as Unified Multimodal Generation
English | 简体中文
This repository contains the dataset for the paper Vision as Unified Multimodal Generation.
SenseNova Vision Corpus 50M
Overview
SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M.reflection-50m
SPP Reflection 50M
The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero — the production half-corpus run, and the dataset the
released models were actually trained on.
🔬 Small sample (same format): dlab-spp/reflection-sample-2k
📉 Earlier 10M run: dlab-spp/reflection-10m
🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications
Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.SIFT-50M
Dataset Card for SIFT-50M
SIFT-50M (Speech Instruction Fine-Tuning) is a 50-million-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). It is built from publicly available speech corpora containing a total of 14K hours of speech and leverages LLMs and off-the-shelf expert models. The dataset spans five languages, covering diverse aspects of speech understanding and controllable speech generation instructions. SIFT-50M… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/SIFT-50M.vlite36-50M-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite36-50M-dataset.wildchat-50m-extended-resultsvlite7-mini-50m-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite7-mini-50m-dataset.
