datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Personix
Open-Personix
Dataset Summary
Open-Personix is a structured JSON dataset maintained under Poralus.
The dataset is primarily text and metadata: each record contains a relative image path,
a natural-language caption, and descriptive annotation fields for a person-centered sample.
The dataset is designed for workflows such as:
caption generation and caption analysis
text-based filtering over person annotations
metadata-aware retrieval and evaluation
multimodal experiments… See the full description on the dataset page: https://huggingface.co/datasets/Below-Image/Open-Personix.bell-labs-technical-archive
Bell Labs Documents and Stuff
This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing.
What is in the release
Split
Documents
train
1220
validation
29
test
42
The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.herodotusbelle-multiround
Dataset Card for belle-multiround
本資料集為以 BELLE 系列指令資料為基礎,整理/轉寫之多輪對話(multi-round) 繁體中文版本,可作為繁中對話模型在多輪互動上的補強資料。
Dataset Details
Dataset Description
BELLE 是早期具規模的中文指令資料集系列。其原始版本以簡體中文為主、且多為單輪指令對話。本資料集做了兩件事:
多輪化:以 BELLE 的單輪指令為起點,請 LLM 生成自然延伸的後續輪次(追問、澄清、延伸要求),形成 multi-round 結構。
繁中化:將內容轉寫為繁體中文,並調整在地用語。
可用於補強模型在多輪對話中的脈絡保持能力。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: cc-by-nc-sa-4.0
Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/belle-multiround.
