datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.wavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.DataClawEval
DataClawEval
An executable benchmark for end-to-end data-engineering agents in industrial environments.
DataClawEval measures an autonomous agent's ability to inspect data, implement and debug pipelines,
and materialize correct artifacts in realistic data-engineering workflows. It contains 100
production-grounded tasks across five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino,
and FlinkSQL. Each task runs in an isolated Docker sandbox and is evaluated by a… See the full description on the dataset page: https://huggingface.co/datasets/dicemy/DataClawEval.DICE-BENCH
🎲 DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
🔗 Links for Reference
Repository: https://github.com/snuhcc/DICE-Bench
Paper: https://arxiv.org/abs/2506.22853
Project page: https://snuhcc.github.io/DICE-Bench/
Point of Contact: kyochul@snu.ac.kr
📖 Paper Description
DICE-BENCH is a benchmark that tests how well large language models can call external functions in realistic… See the full description on the dataset page: https://huggingface.co/datasets/OfficerChul/DICE-BENCH.wildchat50m-rewild-sft-385700
wildchat50m-rewild-sft-385700
A supervised fine-tuning (SFT) dataset formed by the union of three sources, each
reformatted to a single canonical conversational schema (WildChat's format is the
ground-truth). Single train split, 385,700 rows.
This is a capped variant of
nyu-dice-lab/wildchat50m-rewild-sft-1118773:
identical union and format handling, except the WildChat source is randomly
subsampled to 250,000 rows (the other two sources are kept in full).
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wildchat50m-rewild-sft-385700.wildchat50m-rewild-sft-1118773
wildchat50m-rewild-sft-1118773
A supervised fine-tuning (SFT) dataset formed by the union of three sources, each
reformatted to a single canonical conversational schema (WildChat's format is the
ground-truth). Single train split, 1,118,773 rows.
Canonical schema
Column
Type
Description
conversation_hash
string
Per-row identifier
conversation
list[{role: string, content: string}]
The chat turns
model
string
Provenance / generating-model label… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wildchat50m-rewild-sft-1118773.
