datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLVisionQA-QBenchDataset for Paper: Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.
Images: images.tar
dev-labels: llvisionqa_dev.json
test-labels: llvisionqa_test.json
See Github for Usage: https://github.com/vqassessment/q-bench.
Feel free to cite us.
@article{wu2023qbench,
title={Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision},
author={Wu, Haoning and Zhang, Zicheng and Zhang, Erli and Chen, Chaofeng and Liao, Liang and Wang, Annan and… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LLVisionQA-QBench.eval_record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 11,
"total_frames": 19650,
"total_tasks": 1,
"total_videos": 11,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:11"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/teodortio/eval_record-test.theconstruct2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "g1",
"total_episodes": 20,
"total_frames": 7858,
"total_tasks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/teotomic/theconstruct2.diffusiondb_ner
Description
Extended dataset infered by the name entity recognition model en_ner_prompting. This model has been trained on hand-annotated prompts from poloclub/diffusiondb.
This dataset is hence infered by this model and can comprise mistakes, especially on certain categories (cf. model card).
The entities comprise 7 main categories and 11 subcategories for a total of 16 categories, extracted from a topic analysis made with BERTopic.
The topic analysis can be explored the… See the full description on the dataset page: https://huggingface.co/datasets/teo-sanchez/diffusiondb_ner.theconstruct_appleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "g1",
"total_episodes": 21,
"total_frames": 4241,
"total_tasks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/teotomic/theconstruct_apple.batch_test_fixed
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
6
Examples
52
Shard size
10
Updated
2026-07-13 10:07 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_fixed")
ds = load_dataset("TeoStarshine/batch_test_fixed", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb / math)… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_fixed.qwen35-continuation-bench
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
1
Examples
100
Shard size
500
Updated
2026-07-09 18:29 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen35-continuation-bench")
ds = load_dataset("TeoStarshine/qwen35-continuation-bench", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen35-continuation-bench.github-issues-datasetqwen_continuation_dataset2
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:41 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset2")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset2", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset2.nobatched_test_fixed
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
5
Examples
50
Shard size
10
Updated
2026-07-13 10:26 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/nobatched_test_fixed")
ds = load_dataset("TeoStarshine/nobatched_test_fixed", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/nobatched_test_fixed.qwen_continuation_dataset3
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:48 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset3")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset3", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset3.qwen_continuation_dataset4
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:53 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset4")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset4", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset4.batch_test_fixed8
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
5
Examples
50
Shard size
10
Updated
2026-07-13 10:40 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_fixed8")
ds = load_dataset("TeoStarshine/batch_test_fixed8", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb /… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_fixed8.batch_test_entropy
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
5
Examples
50
Shard size
10
Updated
2026-07-13 11:01 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_entropy")
ds = load_dataset("TeoStarshine/batch_test_entropy", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb /… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_entropy.
