datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Zhoulifeng-Streaming-Dataset
Zhoulifeng-Streaming-Dataset (峰哥直播语料 SFT 数据清洗流水线)
本项目旨在将“峰哥亡命天涯”的原始直播语音转录文本(ASR Transcripts)清洗、提纯、去重并重写,最终构建出高质量的大模型指令微调(SFT)和思维链(CoT)数据集。
本流水线包含 11 个核心处理步骤,通过规则过滤、统计学过滤、语义降噪以及大模型(LLM)深度重写等手段,将原始的五万多条粗糙问答,提纯为高逻辑密度、高“峰味”浓度的高质量训练集。
🗂️ 核心处理流程 (Data Pipeline)
数据清洗呈现严格的“漏斗状”递减,以下是完整的清洗流程序号与功能说明:
Step 01: 语音文本标点恢复 (Punctuation Generation)
脚本: 01_punctuation_generate.py
输入/输出: 生成 01_punc_transcripts/ 目录
说明: 针对原始 ASR 语音识别出的无断句纯文本,利用模型恢复正确的标点符号,为后续的语义切分和 QA… See the full description on the dataset page: https://huggingface.co/datasets/hzb29/Zhoulifeng-Streaming-Dataset.streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.Streaming_Sampleafrica-synth-streaming-consumption-all
African Streaming Consumption | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: json - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-streaming-consumption-all.model0_boundary_predict_streaming
## Dataset Card: Terminal Log Boundary Prediction (Streaming)
📋 Dataset Purpose & Model 0 Overview
This dataset is designed to train "Model 0" for the Winter 2026 iteration of the AutoDocs project.
You can access the official repository here:AutoDocs (Winter 2026) Repository
Objective
The primary objective of the model is the binary classification of sequential data.
It is engineered to process continuous, timestamped terminal logs formatted in XML to… See the full description on the dataset page: https://huggingface.co/datasets/Jaiccc/model0_boundary_predict_streaming.VQA-stage2-streamingyodataset-streaming-elbrus
