CoolFace
Datasetpublic

BobbleAI/Bobble-Hinglish-Sports-Dataset_BHSD

Dataset Card for Bobble Hinglish Sports Dataset (BHSD) Dataset Description The Bobble Hinglish Sports Dataset is a meticulously curated collection of 7,029 code-mixed sentences spanning various sports categories. It includes human annotations across seven distinct sports categories, along with out-of-scope values. This testing dataset is specifically designed to enhance NLP models' ability to understand Hinglish sports content, making it valuable for intent… See the full description on the dataset page: https://huggingface.co/datasets/BobbleAI/Bobble-Hinglish-Sports-Dataset_BHSD.

sourceHugging Facecc-by-nc-nd-4.0updated 1y agoView on Hugging Face
0likes57downloads
Dataset Card

Dataset Card for Bobble Hinglish Sports Dataset (BHSD)

Dataset Description

The Bobble Hinglish Sports Dataset is a meticulously curated collection of 7,029 code-mixed sentences spanning various sports categories. It includes human annotations across seven distinct sports categories, along with out-of-scope values. This testing dataset is specifically designed to enhance NLP models' ability to understand Hinglish sports content, making it valuable for intent detection tasks. To ensure high quality, the dataset has undergone thorough reviews by native speakers, capturing a broad spectrum of Hinglish usage in conversational media.

  • —Language(s) (NLP): English, Hindi
  • —License: BobbleHinglishSports_Dataset © 2024 by BobbleAI is licensed under CC BY-NC-ND 4.0

<!--## Dataset Sources [optional]

Provide the basic links for the dataset.

  • —Repository: [More Information Needed]
  • —Paper : [More Information Needed] -->

Uses

<!-- Address questions around how the dataset is intended to be used. --> Testing dataset for intent detection tasks used in chatbots, virtual assistants, and customer service automation.

How to Use

python

import pandas as pd

df = pd.read_csv('test.csv')

print(df.head())

Dataset Structure

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. --> This dataset comprises over 7,000 manually curated sentences, specifically designed to capture the unique blend of Hindi-English code-mixing. The dataset structure consists of high quality Hinglish code-mixed sentences along with their sports intent label.

Sports CategorySentences
Cricket2642
Football465
Kabaddi154
Carrom154
Ludo139
Chess126
outofscope3349
Total7029

Dataset Creation

Curation Rationale

This Hindi-English Human Annotated Dataset for Sports Categories aims to support NLP model development and benchmarking. Data is manually generated focusing on sports discussions based on the abundance of sports content in conversational media. It includes a mix of pure Hindi, pure English, and Hinglish texts, covering cricket, football, ludo, etc. Native speakers have annotated the data for sport categories. Quality is ensured through multiple review rounds and validation checks. Ethical considerations include privacy protection and bias mitigation. This test dataset facilitates inferencing and enhances Hindi-English NLP task capabilities in sports context.

<!-- Motivation for the creation of this dataset. -->

Source Data

Manually curated by experts in Hindi and English language.

<!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). -->

Performance Analysis of SOTA LLMs over proposed Hinglish Sports Dataset

<!--| COMMON PROMPT | <!--|--------------------------------------------------------------------------------| | Parameters | Model | Zero Shot | Single Shot | Two Shot | |:--------------:|:------------------:|:-----------:|:-------------:|:----------:| | 780M | Flan-T5-large | 61.56 65.17 | 63.69 55.89 | 65.00 56.90| | 6B | Gemini-1.5-Pro | 60.44 79.49 | 63.68 80.50 | 62.64 81.14| | 7B | Airavata | 62.53 43.88 | 58.17 2.05 | 59.68 2.05 | | 8B | Llama 3.1 | 75.52 | 73.93 | 69.29 | | 11B | Flan-T5-XXL | 70.52 | 72.37 | 72.37 | | 12B | Mistral-Nemo | 64.26 80.35 | 64.33 80.30 | 65.29 78.06| | 405B | Llama 3.1 | 66.88 82.16 | 70.27 82.50 | 71.72 74.09| | NA* | GPT-4o | 70.62 75.32 | 70.19 61.69 | 66.99 72.37|-->

<table style="width: 100%; border-collapse: collapse; text-align: center;"> <tr> <!-- First row with merged headers --> <th rowspan="2" style="border: 1px solid #000; padding: 8px;">Model</th> <th rowspan="2" style="border: 1px solid #000; padding: 8px; width:100px;">Parameters</th> <th colspan="3" style="border: 1px solid #000; padding: 8px;">Common Prompts</th> <th style="border: none;"></th> <!-- Empty column for spacing --> <th colspan="3" style="border: 1px solid #000; padding: 8px;">Best Prompts</th> </tr> <tr> <!-- Second row with individual headers under each section --> <th style="border: 1px solid #000; padding: 18px;">Zero Shot</th> <th style="border: 1px solid #000; padding: 8px;">Single Shot</th> <th style="border: 1px solid #000; padding: 8px;">Two Shot</th> <th style="border: none;width:10px;"></th> <!-- Empty column for spacing --> <th style="border: 1px solid #000; padding: 8px;">Zero Shot</th> <th style="border: 1px solid #000; padding: 8px;">Single Shot</th> <th style="border: 1px solid #000; padding: 8px;">Two Shot</th> </tr> <!-- Example data rows -->

<tr> <td style="border: 1px solid #000; padding: 8px;">Flan T5-XXL</td> <td style="border: 1px solid #000; padding: 8px;">11B</td> <td style="border: 1px solid #000; padding: 8px;">61.56</td> <td style="border: 1px solid #000; padding: 8px;">63.69</td> <td style="border: 1px solid #000; padding: 8px;">65.00</td> <td style="border: none;width:10px;"></td> <!-- Empty column for spacing --> <td style="border: 1px solid #000; padding: 8px;">63.48</td> <td style="border: 1px solid #000; padding: 8px;">65.87</td> <td style="border: 1px solid #000; padding: 8px;">66.17</td> </tr>

<tr> <td style="border: 1px solid #000; padding: 8px;">Llama 3.1</td> <td style="border: 1px solid #000; padding: 8px;">405B</td> <td style="border: 1px solid #000; padding: 8px;">66.88</td> <td style="border: 1px solid #000; padding: 8px;">70.27</td> <td style="border: 1px solid #000; padding: 8px;">71.72</td> <td style="border: none;width:10px;"></td> <!-- Empty column for spacing --> <td style="border: 1px solid #000; padding: 8px;">67.53</td> <td style="border: 1px solid #000; padding: 8px;">70.27</td> <td style="border: 1px solid #000; padding: 8px;">71.72</td> </tr> <tr> <td style="border: 1px solid #000; padding: 8px;">Mistral-Nemo</td> <td style="border: 1px solid #000; padding: 8px;">12B</td> <td style="border: 1px solid #000; padding: 8px;">64.26</td> <td style="border: 1px solid #000; padding: 8px;">64.33</td> <td style="border: 1px solid #000; padding: 8px;">65.29</td> <td style="border: none;width:10px;"></td> <!-- Empty column for spacing --> <td style="border: 1px solid #000; padding: 8px;">66.55</td> <td style="border: 1px solid #000; padding: 8px;">64.33</td> <td style="border: 1px solid #000; padding: 8px;">65.29</td> </tr> <tr> <td style="border: 1px solid #000; padding: 8px;">Airavata</td> <td style="border: 1px solid #000; padding: 8px;">7B</td> <td style="border: 1px solid #000; padding: 8px;">62.53</td> <td style="border: 1px solid #000; padding: 8px;">58.17</td> <td style="border: 1px solid #000; padding: 8px;">59.68</td> <td style="border: none;width:10px;"></td> <!-- Empty column for spacing --> <td style="border: 1px solid #000; padding: 8px;">62.53</td> <td style="border: 1px solid #000; padding: 8px;">59.98</td> <td style="border: 1px solid #000; padding: 8px;">63.07</td> </tr> <tr> <td style="border: 1px solid #000; padding: 8px;">Gemini 1.5-pro</td> <td style="border: 1px solid #000; padding: 8px;">NA</td> <td style="border: 1px solid #000; padding: 8px;">60.44</td> <td style="border: 1px solid #000; padding: 8px;">63.68</td> <td style="border: 1px solid #000; padding: 8px;">62.64</td> <td style="border: none;width:10px;"></td> <!-- Empty column for spacing --> <td style="border: 1px solid #000; padding: 8px;">71.23</td> <td style="border: 1px solid #000; padding: 8px;">68.84</td> <td style="border: 1px solid #000; padding: 8px;">67.93</td> </tr> <!--<tr> <td style="border: 1px solid #000; padding: 8px;">QXLab-eus1</td> <td style="border: 1px solid #000; padding: 8px;">NA</td> <td style="border: 1px solid #000; padding: 8px;">71.24</td> <td style="border: 1px solid #000; padding: 8px;">77.58</td> <td style="border: 1px solid #000; padding: 8px;">77.72</td> <td style="border: none;width:10px;"></td> <td style="border: 1px solid #000; padding: 8px;">71.24</td> <td style="border: 1px solid #000; padding: 8px;">77.58</td> <td style="border: 1px solid #000; padding: 8px;">77.72</td> </tr> --> <tr> <td style="border: 1px solid #000; padding: 8px;">GPT-4o</td> <td style="border: 1px solid #000; padding: 8px;">NA</td> <td style="border: 1px solid #000; padding: 8px;">70.62</td> <td style="border: 1px solid #000; padding: 8px;">70.19</td> <td style="border: 1px solid #000; padding: 8px;">66.99</td> <td style="border: none;width:10px;"></td> <!-- Empty column for spacing --> <td style="border: 1px solid #000; padding: 8px;">74.08</td> <td style="border: 1px solid #000; padding: 8px;">73.07</td> <td style="border: 1px solid #000; padding: 8px;">73.99</td> </tr> </table> Parameters not available <!-- Provide a quick summary of the dataset. -->

<!-- #### Data Collection and Processing -->

<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->

<!-- #### Who are the source data producers?

<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->

<!-- [More Information Needed]

<!-- ### Annotations [optional]

<!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. -->

<!-- #### Annotation process

<!-- This section describes the annotation process such as annotation tools used in the process, the amount of data annotated, annotation guidelines provided to the annotators, interannotator statistics, annotation validation, etc. -->

<!-- [More Information Needed] -->

Who are the annotators?

Human annotators are experts in native languages.

<!-- This section describes the people or systems who created the annotations. -->

<!-- [More Information Needed]

Personal and Sensitive Information

<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->

<!-- [More Information Needed]

Bias, Risks, and Limitations

<!-- This section is meant to convey both technical and sociotechnical limitations. -->

<!-- [More Information Needed]

Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

<!-- Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.

Citation [optional]

<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->

<!-- ## Glossary [optional]

<!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. -->

<!-- ## More Information [optional] -->

<!-- [More Information Needed] -->

Dataset Card Authors

Artificial Intelligence Department, Bobble AI, Gurugram, India

Dataset Card Contact

For enquiries or issues, you can contact: BobbleAI