datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Daring-Anteater
Dataset Card
Daring-Anteater is a comprehensive dataset for instruction tuning, covering a wide range of tasks and scenarios. The majority of the dataset is synthetically generated using NVIDIA proprietary models and Mixtral-8x7B-Instruct-v0.1, while the remaining samples are sourced from FinQA, wikitablequestions, and commercially-friendly subsets of Open-Platypus.
This dataset is used in HelpSteer2 paper, resulting in a solid SFT model for further preference tuning. We… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Daring-Anteater.empathy-conversations
AntEngage Empathy Conversation Dataset
Organization: AntEngage Technology Private Limited
Version: 1.0.0
License: CC BY-SA 4.0
Language: English
DOI: 10.7910/DVN/LWUFLG
Source & full datasheet: antengage.com/datasets
Dataset Summary
4,008 empathy-focused multi-turn conversations, generated by AntEngage's own
pipeline and quality-verified by an automated cross-checker. Each conversation
simulates an emotional support dialogue in which a speaker expresses distress… See the full description on the dataset page: https://huggingface.co/datasets/AntEngage/empathy-conversations.ante-bench
Ante Bench
A fully synthetic benchmark for contribution-acceptance protocols: how should
a project decide whether to accept a pull request when it cannot afford to read
every one of them?
42 pull requests across 3 small Python projects, in four
ground-truth classes:
class
n
what it is
legitimate
20
a genuine, correct contribution (bug fix, feature, docs, typing, performance, refactor)
plausible_wrong
8
looks right, breaks behaviour
evidence_gaming
8
evidence… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/ante-bench.
