datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-for-excel-sessions
Coding agent session traces for thomasmustier/pi-for-excel-sessions
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-for-excel-sessions.Taiwan-Text-Excellence-2B
High quality corpus for Taiwanese culture and Traditional Chinese
Taiwan Text Excellence (TTE)
Contains high quality news and articles in Traditional Chinese.
The data processing pipeline is optimized for LLM performance.
Is de-duplicated and cleaned using both rule-based and learning-based filters.
E.g., urls/emails/html tags/abnormal characters are cleaned, and numbers (full-width or half-width) are normalized.
Contains ~2 billion tokens, measured using BPE tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/liswei/Taiwan-Text-Excellence-2B.Taiwan-Text-Excellence-sentences
台灣文摘句資料集
概述
Taiwan-Text-Excellence 句子資料集是從較大的 liswei/Taiwan-Text-Excellence-2B 資料集中抽取的 200 萬個獨特中文句子的綜合集。這些句子是隨機選取的,並使用 chinese-sentence-processor 工具進行分割。此資料集非常適合各種自然語言處理任務,包括語言建模、文本生成和其他研究用途。
資料集統計
總句數: 2,000,000
訓練集: 1,600,000 個句子
測試集: 400,000 個句子
資料格式
資料集中的每一行都包含一個欄位:
**text**:包含中文句子的字串。
範例
{"text": "而新郎和女方家人的脂燭在當晚亦會合二為一,再送到母屋帳前點燃一個燈籠,保持三日不滅。"}
{"text": "這個時期簽署的現代劇至今仍是台灣戲劇的中堅力量,而這十年為之奮鬥也奠定了其後數十年的基礎。"}
{"text": "曾柏瑜今天也車票,明天還請在規劃畫相關票活動中。"}… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Taiwan-Text-Excellence-sentences.
