datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sgd_dst
Schema-Guided Dialogue dataset - Dialogue State Tracking
This dataset contains the Schema-Guided Dialogue Dataset, formatted according to the prompt formats from the following two dialogue state tracking papers:
Description-Driven Dialogue State Tracking (D3ST) (Zhao et al., 2022)
Show, Don't Tell (SDT) (Gupta et al., 2022)
Data processing code: https://github.com/google-research/task-oriented-dialogueOriginal dataset:… See the full description on the dataset page: https://huggingface.co/datasets/shermansiu/sgd_dst.New_York_Times_TopicsSeCoDa
SeCoDa
Repository for the Sense Complexity Dataset (SeCoDa)
Paper
For more information on the SeCoDa, see the paper.
Publications using this dataset must include a reference to the following publication:
SeCoDa: Sense Complexity Dataset. David Strohmaier, Sian Gooding, Shiva Taslimipoor, Ekaterina Kochmar. Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), pages 5964–5969, Marseille, 11–16 May 2020
The dataset is based on the earlier… See the full description on the dataset page: https://huggingface.co/datasets/dstrohmaier/SeCoDa.dstc11-intentThis dataset is a subset of the data contained in https://github.com/amazon-science/dstc11-track2-intent-induction/tree/main
The data is filtered by speaker_role as customer and number of intents is 1.
This is an open-source dataset that can be used for demos and explorations.
license: apache-2.0
This project is licensed under the Apache-2.0 License.
Please cite the following papers if using the tasks, code, or data from this track in your work:
@misc{gung2023natcs,
title={NatCS:… See the full description on the dataset page: https://huggingface.co/datasets/splevine/dstc11-intent.3D-DST-captions
3D-DST-captions
As part of our data release in 3D-DST, we present MiniGPT4-generated captions for all 1000 classes in ImageNet-1k.
See wufeim/DST3D for synthetic data generation with 3D annotations using the captions here.
These captions can be used to produce other synthetic datasets for fair comparisons between different data generation procedures.
dstc2_dialogues3D-DST-models
3D-DST-models
As part of our data release in 3D-DST, we present aligned CAD models for all 1000 classes in ImageNet-1k.
See wufeim/DST3D for synthetic data generation with 3D annotations using the CAD models here.
Besides the .csv file as visualized in the dataset viewer above, we also provide a python script (models_3d_dst.py) to help integrate with other Python modules.
Fields
For each CAD model, there are seven fields:
synset: synset associated with each… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/3D-DST-models.dst-question-plus-que-vs-statDstPortsDstIPsDS_testds-testdst-question
Label info:
0: "fragment",
1: "statement",
2: "question",
3: "command",
4: "rhetorical question",
5: "rhetorical command",
6: "intonation-dependent utterance"
Training process:
{'loss': 1.8008, 'grad_norm': 7.2770233154296875, 'learning_rate': 1e-05, 'epoch': 0.03}
{'loss': 0.894, 'grad_norm': 27.84651756286621, 'learning_rate': 2e-05, 'epoch': 0.06}
{'loss': 0.6504, 'grad_norm': 30.617990493774414, 'learning_rate': 3e-05, 'epoch': 0.09}
{'loss': 0.5939, 'grad_norm':… See the full description on the dataset page: https://huggingface.co/datasets/jooni22/dst-question.ds_test
