datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
teogonia
Bangumi Image Base of Teogonia
This is the image base of bangumi Teogonia, we detected 59 characters, 4942 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/teogonia.LSVQ-videosThis is an unofficial copy of the videos in the LSVQ dataset (Ying et al, CVPR, 2021), the largest dataset available for Non-reference Video Quality Assessment (NR-VQA); this is to facilitate research studies on this dataset given that we have received several reports that the original links of the dataset is not available anymore.
See FAST-VQA (Wu et al, ECCV, 2022) or DOVER (Wu et al, ICCV, 2023) repo on its converted labels (i.e. quality scores for videos).
The file links to the labels in… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LSVQ-videos.TEOChatlasTEOChatlas is the first instruction-following dataset for temporal EO data. It contains 554,071 examples spanning dozens of temporal instruction-following tasks.HumanML3DhtdsSMPLX-amass
Dataset Card for "SMPLX-amass"
More Information needed
humanml3d-amassLLVisionQA-QBenchDataset for Paper: Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.
Images: images.tar
dev-labels: llvisionqa_dev.json
test-labels: llvisionqa_test.json
See Github for Usage: https://github.com/vqassessment/q-bench.
Feel free to cite us.
@article{wu2023qbench,
title={Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision},
author={Wu, Haoning and Zhang, Zicheng and Zhang, Erli and Chen, Chaofeng and Liao, Liang and Wang, Annan and… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LLVisionQA-QBench.DIVIDE-MaxWellteochew_wild
Teochew-Wild:首个正字标注的野外潮州话数据集
本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正);
Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。
文件说明 (File Structure Explanation)
├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式
├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.public_iqa_vqa_databasesteologu-capitalmarkets-fragment
TEOLOGU Capital Markets — fragment WP
Pagină completă (hero foto + order ticket interactiv + ticker + 22 secțiuni + notice + CTA), design Revolut v3, namespace tw-*, fixuri mobile incluse.
Varianta 1 — LOADER (recomandat, 1 snippet mic)
Creează pagina, adaugă un bloc Custom HTML cu conținutul din loader.html și salvează.
Varianta 2 — COD COMPLET fără loader (3 lipiri)
Lipește în ordinea 1→2→3 în același bloc Custom HTML, una imediat după alta:… See the full description on the dataset page: https://huggingface.co/datasets/alexandruo/teologu-capitalmarkets-fragment.seamless_nat_audio-e4078c4ateologu-capitalmarkets-v2-fragment
TEOLOGU Capital Markets — Varianta 2 "Editorial Light"
Același conținut integral din prompt, altă grafică față de varianta 1:
canvas deschis cald, titluri serif (Playfair Display), hairlines numerotate,
accent auriu + cobalt, poze în duoton, mockup FX Converter interactiv.
Fișiere
cm2-a.html — CSS complet + deschidere container
cm2-b.html — hero + converter FX + ticker + secțiunile 01–14
cm2-c.html — secțiunile 15–29, ecosistem, CTA, notice, footer, script… See the full description on the dataset page: https://huggingface.co/datasets/alexandruo/teologu-capitalmarkets-v2-fragment.teologu-pay-fragmentselma-corpus1teologu-knowledge-fragmentteologu-wealth-fragment
TEOLOGU Wealth — fragment WP pentru teologu.com/wealth-2
Fragmentul complet (namespaced tw-*, cu fix defensiv Lifecycle) este împărțit în 5 fișiere care se lipeșc în ordine, unul după altul, în ACEEAȘI zonă de conținut (un singur bloc/widget Custom HTML):
part-1-of-5-head.html — comentariu, fonturi, CSS (variabile, hero, phone, marquee, secțiuni, panel-uri, grid-uri, tile-uri, chips, flow, split)
part-2-of-5-css2-hero.html — restul CSS (cards, ladder, horizons, treemap… See the full description on the dataset page: https://huggingface.co/datasets/alexandruo/teologu-wealth-fragment.teologu-marketplace-fragmentteologu-ai-fragmentteologu-security-fragmentteologu-research-fragmentLLDescribe-QBenchDataset for Paper: Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.
See Github: https://github.com/vqassessment/q-bench.
Feel free to cite us.
@article{wu2023qbench,
title={Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision},
author={Wu, Haoning and Zhang, Zicheng and Zhang, Erli and Chen, Chaofeng and Liao, Liang and Wang, Annan and Li, Chunyi and Sun, Wenxiu and Yan, Qiong and Zhai, Guangtao and Lin, Weisi},
year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LLDescribe-QBench.ShareGPT-WebDatamonolingual-quechua-iic
Dataset Card for Monolingual-Quechua-IIC
Dataset Summary
We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers.
Supported Tasks and Leaderboards
More Information Needed
Languages
Southern… See the full description on the dataset page: https://huggingface.co/datasets/Teo1-7/monolingual-quechua-iic.teologu-infrastructure-fragmentteologu-tech-v2-fragmentLongVideoBench_Miniteologu-youth-fragmentteologu-data-intelligence-fragment
