datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLVisionQA-QBenchDataset for Paper: Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.
Images: images.tar
dev-labels: llvisionqa_dev.json
test-labels: llvisionqa_test.json
See Github for Usage: https://github.com/vqassessment/q-bench.
Feel free to cite us.
@article{wu2023qbench,
title={Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision},
author={Wu, Haoning and Zhang, Zicheng and Zhang, Erli and Chen, Chaofeng and Liao, Liang and Wang, Annan and… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LLVisionQA-QBench.teochew_wild
Teochew-Wild:首个正字标注的野外潮州话数据集
本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正);
Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。
文件说明 (File Structure Explanation)
├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式
├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.lab2-ngramsvalue-for-instruction-tuning
Dataset Overview
This dataset is derived from the existing datasets ETHICS, SOCIAL-CHEM-101, and UNIMORAL, with additional annotations for both normative ethics and moral foundation labels for each scenario. The dataset is wrapped with instruction-tuning template, and can be directly used for instruction tuning. For more information, see the github repo
batch_test_fixed
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
6
Examples
52
Shard size
10
Updated
2026-07-13 10:07 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_fixed")
ds = load_dataset("TeoStarshine/batch_test_fixed", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb / math)… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_fixed.qwen35-continuation-bench
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
1
Examples
100
Shard size
500
Updated
2026-07-09 18:29 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen35-continuation-bench")
ds = load_dataset("TeoStarshine/qwen35-continuation-bench", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen35-continuation-bench.qwen_continuation_dataset2
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:41 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset2")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset2", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset2.nobatched_test_fixed
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
5
Examples
50
Shard size
10
Updated
2026-07-13 10:26 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/nobatched_test_fixed")
ds = load_dataset("TeoStarshine/nobatched_test_fixed", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/nobatched_test_fixed.qwen_continuation_dataset3
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:48 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset3")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset3", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset3.qwen_continuation_dataset4
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:53 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset4")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset4", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset4.batch_test_fixed8
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
5
Examples
50
Shard size
10
Updated
2026-07-13 10:40 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_fixed8")
ds = load_dataset("TeoStarshine/batch_test_fixed8", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb /… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_fixed8.batch_test_entropy
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
5
Examples
50
Shard size
10
Updated
2026-07-13 11:01 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_entropy")
ds = load_dataset("TeoStarshine/batch_test_entropy", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb /… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_entropy.dataset2dataset3alpaca_mini_slice
