datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-omni-open-source-balanced-1200
Qwen3-Omni 30A3 open-source balanced subset
This dataset contains 1,200 samples selected from public-source-labelled portions of the Qwen3-Omni 30A3 posttrain recipe. The 8 topics are balanced at 150 samples each. Every item includes topic, public_dataset, public_dataset_confidence, source_id, and provenance_json fields. public_dataset is the canonical per-item public-dataset label.
Loading
The data/train-*.jsonl shards are ordinary Hugging Face JSONL data files… See the full description on the dataset page: https://huggingface.co/datasets/Transl/qwen3-omni-open-source-balanced-1200.open-source-ai-models-dataset
OpenModelMap — The Largest Open-Source AI Models Dataset (Chinese + English)
2,484 models · 35 fields · 9 sources · Updated daily
This dataset provides the most comprehensive structured metadata for open-source AI models, with a focus on Chinese model coverage. Every model includes benchmark scores, hardware requirements, GPU compatibility, license information, and deployment methods.
What's Inside
Field
Description
id
HuggingFace model ID
name… See the full description on the dataset page: https://huggingface.co/datasets/duola15/open-source-ai-models-dataset.
