festvox/cmu_hinglish_dog
Dataset Card for CMU Document Grounded Conversations Dataset Summary This is a collection of text conversations in Hinglish (code mixing between Hindi-English) and their corresponding English versions. Can be used for Translating between the two. The dataset has been provided by Prof. Alan Black's group from CMU. Supported Tasks and Leaderboards abstractive-mt Languages Dataset Structure Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/festvox/cmu_hinglish_dog.
Convert dataset to Parquet (#4)
Delete legacy JSON metadata (#3)
Support streaming (#2)
Reorder split names (#1)
add dataset_info in dataset metadata
remove dummmy data
Align/fix license metadata info (#4613)
Remove config names as yaml keys (#4367)
Update datasets task tags to align tags with models (#4067)
Update files from the datasets library (from 1.16.0)
