datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GitTaskBench
Dataset Card for GitTaskBench
The dataset was presented in the paper GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging.
Dataset Details
Dataset Description
GitTaskBench is a benchmark dataset designed to evaluate the capabilities of code-based intelligent agents in solving real-world tasks by leveraging GitHub repositories.It contains 54 representative tasks across 7 domains, carefully curated to reflect… See the full description on the dataset page: https://huggingface.co/datasets/Nicole-Yi/GitTaskBench.github_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-speech-transcripts
Speech Transcripts
Dataset Summary
Time-aligned transcripts of English speech audio.
Dataset Structure
Data fields: audio_path, text, start_time, end_time.
Licensing Information
This dataset is released under the CC BY-SA 4.0 license (cc-by-sa-4.0).
