datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
teogonia
Bangumi Image Base of Teogonia
This is the image base of bangumi Teogonia, we detected 59 characters, 4942 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/teogonia.LSVQ-videosThis is an unofficial copy of the videos in the LSVQ dataset (Ying et al, CVPR, 2021), the largest dataset available for Non-reference Video Quality Assessment (NR-VQA); this is to facilitate research studies on this dataset given that we have received several reports that the original links of the dataset is not available anymore.
See FAST-VQA (Wu et al, ECCV, 2022) or DOVER (Wu et al, ICCV, 2023) repo on its converted labels (i.e. quality scores for videos).
The file links to the labels in… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LSVQ-videos.HumanML3DhtdsSMPLX-amass
Dataset Card for "SMPLX-amass"
More Information needed
humanml3d-amassLLVisionQA-QBenchDataset for Paper: Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.
Images: images.tar
dev-labels: llvisionqa_dev.json
test-labels: llvisionqa_test.json
See Github for Usage: https://github.com/vqassessment/q-bench.
Feel free to cite us.
@article{wu2023qbench,
title={Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision},
author={Wu, Haoning and Zhang, Zicheng and Zhang, Erli and Chen, Chaofeng and Liao, Liang and Wang, Annan and… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LLVisionQA-QBench.teochew_wild
Teochew-Wild:首个正字标注的野外潮州话数据集
本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正);
Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。
文件说明 (File Structure Explanation)
├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式
├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.seamless_nat_audio-e4078c4aselma-corpus1monolingual-quechua-iic
Dataset Card for Monolingual-Quechua-IIC
Dataset Summary
We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers.
Supported Tasks and Leaderboards
More Information Needed
Languages
Southern… See the full description on the dataset page: https://huggingface.co/datasets/Teo1-7/monolingual-quechua-iic.lab2-ngramsextreme-floods-kg
📖 Dataset Summary
Floods are among the most frequent and devastating disasters worldwide, yet data describing them is often scattered across unstructured reports, geospatial sources, and satellite imagery.This dataset unifies those heterogeneous data sources into a structured, ontology-aligned Knowledge Graph (KG) format.
Each event is represented with:
Metadata: disaster type, location, date, country.
Textual Descriptions: humanitarian situation reports from ReliefWeb.… See the full description on the dataset page: https://huggingface.co/datasets/teoaivalis/extreme-floods-kg.socio-ecological-regions
Socio-Environmental Regions Dataset
This dataset contains environmental and socioeconomic data organized by geographical regions and countries. The data is extracted from a multi-axis KG and covers 149 processed geographical regions and countries, mapping their physical terrain characteristics, local species, and historical flood profiles.
Environmental and Exposure Axes
Each region contains data structured across three primary categories:
Biological Axis:… See the full description on the dataset page: https://huggingface.co/datasets/teoaivalis/socio-ecological-regions.diffusiondb_ner
Description
Extended dataset infered by the name entity recognition model en_ner_prompting. This model has been trained on hand-annotated prompts from poloclub/diffusiondb.
This dataset is hence infered by this model and can comprise mistakes, especially on certain categories (cf. model card).
The entities comprise 7 main categories and 11 subcategories for a total of 16 categories, extracted from a topic analysis made with BERTopic.
The topic analysis can be explored the… See the full description on the dataset page: https://huggingface.co/datasets/teo-sanchez/diffusiondb_ner.hq-kineticsvalue-for-instruction-tuning
Dataset Overview
This dataset is derived from the existing datasets ETHICS, SOCIAL-CHEM-101, and UNIMORAL, with additional annotations for both normative ethics and moral foundation labels for each scenario. The dataset is wrapped with instruction-tuning template, and can be directly used for instruction tuning. For more information, see the github repo
us-warn-act-layoff-notices
US WARN Act Layoff & Closure Notices — normalized snapshot (2026-07-29)
5,660 US WARN Act layoff and closure notices, normalized to one schema
across 5 states: TX (2,358), CA (1,632), IL (1,369), NC (236), NY (65).
Covering 528,129 affected workers and 4,635 distinct employers,
with effective dates from 2010-10-30 to 2028-02-01.
This is a frozen snapshot
Taken 2026-07-29. It does not update. WARN filings arrive continuously,
so this file begins going stale the day… See the full description on the dataset page: https://huggingface.co/datasets/teosv1/us-warn-act-layoff-notices.LLMs-Turkish-TEOG-Leaderboard
TEOG Scores Leaderboard
Welcome to the TEOG Scores Leaderboard! This repository contains the results of evaluating various large language models (LLMs) on the TEOG (Temel Eğitimden Ortaöğretime Geçiş) exam dataset. The TEOG exam is a standardized test in Turkey used for high school admissions, and this dataset provides a benchmark for assessing the performance of LLMs in Turkish educational tasks. Please remember that full score for TEOG is 500 points.
More Models Are… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/LLMs-Turkish-TEOG-Leaderboard.hd-vila-metadoc-handwriting-extraction-complex-1batch_test_fixed
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
6
Examples
52
Shard size
10
Updated
2026-07-13 10:07 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/batch_test_fixed")
ds = load_dataset("TeoStarshine/batch_test_fixed", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source dataset (fineweb / math)… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/batch_test_fixed.simpson-conversations-sharegptsocrates-instruction-datasetFilter-CoF-RL-full-balanceqwen35-continuation-bench
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
1
Examples
100
Shard size
500
Updated
2026-07-09 18:29 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen35-continuation-bench")
ds = load_dataset("TeoStarshine/qwen35-continuation-bench", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen35-continuation-bench.VietnamsesLegalDatasetself-driving
Self-driving dataset
Dataset containing model training
github-issues-datasetqwen_continuation_dataset2
Qwen Continuation Dataset
Generated with qwen_continuation_dataset.
Statistics
Shards
2
Examples
20
Shard size
10
Updated
2026-07-12 10:41 UTC
Usage
from datasets import load_dataset
ds = load_dataset("TeoStarshine/qwen_continuation_dataset2")
ds = load_dataset("TeoStarshine/qwen_continuation_dataset2", streaming=True)
Fields
Field
Description
source_id
source document ID
source_name
source… See the full description on the dataset page: https://huggingface.co/datasets/TeoStarshine/qwen_continuation_dataset2.
