CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yugrathee28 /Hinglish-dataset 🇮🇳 Hinglish Dataset — 1.4 Million Samples Industrial-Grade Code-Mixed NLP Dataset | By ScaleIndia AI · Founder: Yug Rathee This repository contains a 5,000-row teaser sample from the full 1.46 Million+ Hinglish comment dataset built by Scaling YUG (Founder: Yug Rathee(yugrathee28@gmail.com)). Provided strictly for research and evaluation purposes only. Commercial use, redistribution, or production-model training requires explicit written consent from… See the full description on the dataset page: https://huggingface.co/datasets/Yugrathee28/Hinglish-dataset.tabular1K<n<10K2 likes59 downloads5mo agoHugging Face02Ghanashyaam /CallAgentAI-Hinglish-Customer-Service CallAgent AI: Hinglish Business Conversations Dataset This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs. Why this dataset exists Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.tabulartext-generationn<1K0 likes37 downloads26d agoHugging Face03sunitha98 /hindi-driving-test-questionstabularn<1K0 likes19 downloads2y agoHugging Face04open-llm-leaderboard /1024m__PHI-4-Hindi-detailsgated Dataset Card for Evaluation run of 1024m/PHI-4-Hindi Dataset automatically created during the evaluation run of model 1024m/PHI-4-Hindi The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/1024m__PHI-4-Hindi-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face05open-llm-leaderboard /1-800-LLMs__Qwen-2.5-14B-Hindi-detailsgated Dataset Card for Evaluation run of 1-800-LLMs/Qwen-2.5-14B-Hindi Dataset automatically created during the evaluation run of model 1-800-LLMs/Qwen-2.5-14B-Hindi The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/1-800-LLMs__Qwen-2.5-14B-Hindi-details.tabular10K<n<100K1 likes14 downloads2y agoHugging Face06open-llm-leaderboard /1-800-LLMs__Qwen-2.5-14B-Hindi-Custom-Instruct-detailsgated Dataset Card for Evaluation run of 1-800-LLMs/Qwen-2.5-14B-Hindi-Custom-Instruct Dataset automatically created during the evaluation run of model 1-800-LLMs/Qwen-2.5-14B-Hindi-Custom-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/1-800-LLMs__Qwen-2.5-14B-Hindi-Custom-Instruct-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face07muggle07 /ao3-hinny-ficstabular10K<n<100K0 likes9 downloads3y agoHugging Face08ryandsilva /erc-hinglishtabular10K<n<100K0 likes8 downloads3y agoHugging Face09open-llm-leaderboard /jebish7__qwen2.5-0.5B-IHA-Hin-detailsgated Dataset Card for Evaluation run of jebish7/qwen2.5-0.5B-IHA-Hin Dataset automatically created during the evaluation run of model jebish7/qwen2.5-0.5B-IHA-Hin The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jebish7__qwen2.5-0.5B-IHA-Hin-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face10somu9 /hindi-hq-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/hindi-hq Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 439,507 Total audio hours 811.7h Codebooks 16 Avg frames/sample 83.1 Avg duration 6.6s Format JSONL file (manifest.jsonl) where each line is: { "text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/hindi-hq-tokens.tabulartext-to-speech100K<n<1M1 likes7 downloads3mo agoHugging Face11open-llm-leaderboard /jebish7__Nemotron-4-Mini-Hindi-4B-Base-detailsgated Dataset Card for Evaluation run of jebish7/Nemotron-4-Mini-Hindi-4B-Base Dataset automatically created during the evaluation run of model jebish7/Nemotron-4-Mini-Hindi-4B-Base The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jebish7__Nemotron-4-Mini-Hindi-4B-Base-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face12open-llm-leaderboard /jebish7__Nemotron-4-Mini-Hindi-4B-Instruct-detailsgated Dataset Card for Evaluation run of jebish7/Nemotron-4-Mini-Hindi-4B-Instruct Dataset automatically created during the evaluation run of model jebish7/Nemotron-4-Mini-Hindi-4B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jebish7__Nemotron-4-Mini-Hindi-4B-Instruct-details.tabular10K<n<100K0 likes4 downloads2y agoHugging Face13Agents-X /sft_data_minio3_wo_image_hintgated tabular1K<n<10K0 likes4 downloads1y agoHugging Face14rishiyadav011 /Hinglish-dataset 🇮🇳 Hinglish Dataset — 1.4 Million Samples Industrial-Grade Code-Mixed NLP Dataset | By ScaleIndia AI · Founder: Yug Rathee This repository contains a 5,000-row teaser sample from the full 1.46 Million+ Hinglish comment dataset built by Scaling YUG (Founder: Yug Rathee(yugrathee28@gmail.com)). Provided strictly for research and evaluation purposes only. Commercial use, redistribution, or production-model training requires explicit written consent from… See the full description on the dataset page: https://huggingface.co/datasets/rishiyadav011/Hinglish-dataset.tabular1K<n<10K0 likes4 downloads3mo agoHugging Face15sunitha98 /ctet-hindi-questionstabularn<1K0 likes3 downloads2y agoHugging Face16uvaidya /hindi-cc100-10btabularn<1K0 likes3 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.