datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english-hindi-colloquial-datasetA curated dataset of colloquial English phrases and their corresponding Hindi translations. This dataset focuses on informal language, including slang, idioms, and everyday expressions, making it ideal for training models that handle casual conversations.
Dataset Details:
Size:e.g., 500+ phrase pairs]
Source: Collected from publicly available conversational datasets, social media, and crowdsourced contributions.
Language Pair: English → Hindi
Annotations: Each phrase pair is manually verified… See the full description on the dataset page: https://huggingface.co/datasets/bajpaideeksha/english-hindi-colloquial-dataset.english-hindi-colloquial-datasetThis dataset consists of a thoughtfully assembled collection of everyday English phrases and their Hindi translations. It highlights informal language, including slang, idiomatic expressions, and common phrases, making it ideal for training models that process casual conversations.
Dataset Details:
Size:e.g., 100+ phrase pairs
Source: Gathered from publicly available datasets and crowd-contributed inputs.
Language Pair: English to Hindi
Use Cases:
Developing and optimizing translation… See the full description on the dataset page: https://huggingface.co/datasets/KumariPrerna2905/english-hindi-colloquial-dataset.SAWiT-Tamil-Colloquial-Datasethindi_colloquial_datasethindi-colloquial-dataset
Hindi Colloquial Dataset
This dataset contains pairs of English Text and Hindi Colloquial Text, designed for training machine learning models for translation .
The dataset was created as part of a hackathon organized by Swati.
Dataset Details
Size: 90 pairs of English and colloquial Hindi sentences
Languages: English, Hindi
Task: Translation, Text Generation
Content: Contains colloquial translations for everyday conversational texts in Hindi.
Example… See the full description on the dataset page: https://huggingface.co/datasets/SirirshaD/hindi-colloquial-dataset.hindi_colloquial_datasetEnglish_hinglish_colloquial-datasetTamil_Colloquial_Datasetcolloquial-telugu-datasetTamil-Colloquial-Standard-Parlance-Corpustelugu_colloquialhindi-colloquial-datasetTelugu_colloquial_language_dataset.csvDataset README
Overview
This dataset is used for machine learning and data analysis. It contains valuable insights on various attributes relevant to the study.
File Information
• Filename: your_dataset.csv
• Format: CSV (Comma-Separated Values)
• Size: [Mention size if known]
• Number of Records: [Mention the number of rows]
• Number of Features: [Mention the number of columns]
Column Descriptions
Column Name Description
ID Unique identifier for each record
Name Name of the entity
Age Age in… See the full description on the dataset page: https://huggingface.co/datasets/sindhuakkaraju/Telugu_colloquial_language_dataset.csv.telugu_colloquial_datasettamil-english-colloquial-translations
Dataset: English-Tamil (en-ta) Parallel Corpus
This dataset contains parallel sentences in English and Tamil (en-ta) that have been curated from multiple sources.
It is designed for tasks such as machine translation, language modeling, and other natural language processing (NLP) applications involving English and Tamil.
Dataset Composition
The dataset is composed of three main parts, which have been concatenated into a single file with two columns: ta (Tamil) and… See the full description on the dataset page: https://huggingface.co/datasets/nandhinivaradharajan14/tamil-english-colloquial-translations.Telugu_colloquial_language_datasettelugu-colloquial-datasetEnglish_to_Tamil_Colloquialenglish-hindi-colloquial-datasetTamil-Colloquial-Datasetcolloquial-malayalamtelugu_colloquial_datasetTelugu_Colloquial_Language_Dataset1Telugu_Colloquial_Dataset1Bhashini_colloquial_Dataset1marathi-colloquial-datasettamil-colloquial-english-translation-dataset
Dataset: English-Tamil Sentences
Overview
This dataset consists of English sentences and their corresponding Tamil translations. The sentences are categorized into short, medium, and long forms to provide a diverse range of expressions, making it useful for language learning, machine translation, and NLP applications.
Dataset Structure
Each entry in the dataset contains:
An English sentence
Its Tamil translation
The dataset includes a variety of sentences… See the full description on the dataset page: https://huggingface.co/datasets/Sruthi-sai-2004/tamil-colloquial-english-translation-dataset.english-tamil-colloquial-dataset
