text-annotations
youtube_annotations_text
Youtube Annotations Text
YouTube 注释(YouTube Annotations)是 YouTube 在 2008 年推出的一项功能,
允许视频创作者在视频上添加文本、链接和互动元素, 以增强观众的观看体验.
YouTube 已在 2019 年删除了此功能.
您可以在这里找到由 omarroth 创建的存档 YouTube Annotations,
本数据集从13亿条存档中提取出了文本.
如果您需要 x_id 与 videoId 的映射, 请使用 utilities/video_text_mapping_indexed.sqlite3 数据库.
lidc-idri-text-annotations
🩺 LIDC-IDRI Text-Annotated
We release a text-annotated version of the LIDC-IDRI dataset, where each annotation is carefully curated from structured metadata provided by radiologists (e.g., malignancy, size, shape, margin, texture, spiculation, etc.).
This enables new research directions in:
Multi-modal learning (image + text)
Text-guided medical image segmentation
Includes
Radiologist Nodule annotations (radiologist contours, malignancy scores)
Natural language… See the full description on the dataset page: https://huggingface.co/datasets/siddharthdhara17/lidc-idri-text-annotations.multilingual-image-annotations-text
Multilingual Image Annotations (Text Only)
Text-only companion to Reubencf/multilingual-image-annotations. Same rows, same google/gemma-4-31B-it annotations, but the image and boxed_image columns are removed so the dataset is small and loadable without binary image bytes.
Stats
Rows: 464
Detection-applicable: 273 (58%)
Languages: en, es, fr, hi, zh, ar, pt
Schema
Column
Type
Notes
image_id
string
UUID/stem of original file
description_en
string… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations-text.
