fto
Datasets
All datasets matching “fto”llava-video-178k-siglip-tokens-ftov-new
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a
redistribution of the source videos. Source:
lmms-lab/LLaVA-Video-178K -- its card
restricts use to academic research and education, and its annotations come
from GPT-4-class models (see the OpenAI usage policy).
Complete: 85000 clips.
Subset
Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.mediawiki-code2code-search
MediaWiki Code2Code Search — Dataset
Pre-computed retrieval artifacts for MediaWiki Code2Code Search, a neural system for the
semantic discovery of open-source software entities (functions, types, templates) across the
MediaWiki / Wikimedia ecosystem.
This Hugging Face dataset is a complementary mirror of the pre-computed artifacts archived on
Zenodo (10.5281/zenodo.20586256). The Hugging Face copy
makes the corpus browsable in the dataset Viewer and easy to pull with the… See the full description on the dataset page: https://huggingface.co/datasets/ftosoni/mediawiki-code2code-search.golden-fto-layer-a
Layer A — Office Action Triples for FTO Evaluation
A public dataset of (invention → cited prior art → outcome) triples
spanning three patent offices:
US slice — the USPTO Office Action Research Dataset (OARD)
Open Data Portal (ODP) API
EP slice — the EPO Open Patent Services (OPS) Register service
(WIPO ST.14 search-report citations)
JP slice — the JPO 整理標準化 (seiri-hyōjunka) ISO bulk data
for 拒絶理由通知 (rejection notices), filing years 2002–2019, published
labels only per JPO… See the full description on the dataset page: https://huggingface.co/datasets/v13s/golden-fto-layer-a.something-something-v2-siglip-tokens-ftov
Something-Something V2 SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not the source videos.
Source: Something-Something V2 (Goyal et al., "The 'something something' video database
for learning and evaluating visual common sense", 2017; Mahdisoltani et al., "On the
effectiveness of task granularity for transfer learning", 2018), distributed by Qualcomm
under its Data License Agreement - Research Use. Read that… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/something-something-v2-siglip-tokens-ftov.clevrer-siglip-tokens-ftov
CLEVRER SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a redistribution
of the source videos. Source: CLEVRER --
"CLEVRER: CoLlision Events for Video REpresentation and Reasoning" (Yi et al.,
ICLR 2020) -- official release from MIT CSAIL under CC0.
Complete: 11000 videos encoded (10000 train, 1000 validation), 0 failed (see manifest.json).
Subset
train: video_00000 ... video_09999 (10000 of the… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/clevrer-siglip-tokens-ftov.huggingface-models-processed
