ido
Datasets
All datasets matching “ido”idolish7
Bangumi Image Base of Idolish7
This is the image base of bangumi IDOLiSH7, we detected 27 characters, 3443 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/idolish7.PopVQA
PopVQA: Popular Entity Visual Question Answering
PopVQA is a dataset designed to study the performance gap in vision-language models (VLMs) when answering factual questions about entities presented in images versus text.
Paper: https://huggingface.co/papers/2412.14133
Code: https://github.com/ido-co/vlm-modality-gap
🔍 Motivation
PopVQA was curated to explore the disparity in model performance when answering factual questions about an entity described in text… See the full description on the dataset page: https://huggingface.co/datasets/idoco/PopVQA.s1s2_water
S1S2-Water (mirror of v1.0.1)
This repository is an unofficial mirror of the S1S2-Water dataset, created and published by its original authors. I am only re-hosting a copy of version v1.0.1 on Hugging Face to provide a faster and more convenient download alternative to Zenodo, whose transfer speeds can be slow for a dataset of this size (~170 GB).
All credit, ownership, and rights belong to the original authors. Please refer to the original source for the authoritative version… See the full description on the dataset page: https://huggingface.co/datasets/idomogalla/s1s2_water.SVHighlights
SVHighlights: Towards Extremely Long Sport Video Highlight Detection
Donggyu Lee*, Youngbin Ki*, Jeonghun Kang, Taehwan Kim — UNIST
KDD 2026 · Datasets & Benchmarks Track (*equal contribution)
SVHighlights is the first highlight-detection benchmark for extremely long
sports videos — 320 full-length broadcasts averaging 2.00 hours
across 8 sports (40 videos each: american football, baseball, basketball,
ice hockey, racing, rugby, soccer, volleyball), totaling 640.18… See the full description on the dataset page: https://huggingface.co/datasets/idong1004/SVHighlights.idol-songs-jpEnglish version follows the Japanese text.
IdolSongsJp: アイドルグループ楽曲スタイルにもとづく音楽コーパス
日本のアイドルグループを模した 15 のオリジナル楽曲からなるコーパスです。
ティザームービー
コーパス構成
楽曲と歌唱者
15 楽曲のうち、8 楽曲が女性グループ、7 曲が男性グループによる楽曲です。
女性歌唱者は 10 名、男性歌唱者は 8 名であり、歌唱メンバーは楽曲によって異なります。
作曲者・編曲者・作詞者はすべてアイドルグループに楽曲提供経験を持つプロフェッショナルです。各歌唱者はセミプロフェッショナルもしくはプロフェッショナルのボーカリストです(実際のアイドルではありません)。
楽曲の一覧と制作者を下記に示します。楽曲長は切り捨てです。BPM は代表値を示しています。
楽曲 ID
楽曲名
作詞
作曲
編曲
楽曲長
BPM
歌唱者数
f01-intro_juice
いんとろじゅーす
ハイジナカムラ… See the full description on the dataset page: https://huggingface.co/datasets/imprt/idol-songs-jp.govuk-policy-qa-pairsThis is a dataset of synthetically generated question and answer pairs on UK government policy papers.
It comes in 2 parts:
Plain text UK government policy papers, scraped from the Gov.uk Policy papers and consultations page. These are in results.json
A series of question and answer pairs on chunk of the above documents, generated using llama_index.finetuning.generate_qa_embedding_pairs and OpenAI GPT3.5 Turbo.
