datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthesized-coding-assistant-dataset
Synthesized Coding Assistant Dataset
Overview
Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice.
Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.Assistantz8
OpenAssistant Conversations Dataset (OASST1)
Dataset Summary
In an effort to democratize research on large-scale alignment, we release OpenAssistant
Conversations (OASST1), a human-generated, human-annotated assistant-style conversation
corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292
quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus
is a product of a worldwide crowd-sourcing effort… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/Assistantz8.
