CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sonalsannigrahi /cv22_azeros FLEURS (Lhotse cuts) Each language is a separate config. Load a single language's cuts as a HF Dataset of raw manifest records with, e.g.: from datasets import load_dataset ds = load_dataset("your-org/REPO_NAME", "bg_bg", split="train") If audio shards (recording.NNNNN.tar) are present alongside the cuts, the LANG/SPLIT/ folder is a valid Lhotse Shar directory. Download it (e.g. via snapshot_download) and load with Lhotse directly: from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/sonalsannigrahi/cv22_azeros.tabular1M<n<10M0 likes322 downloads3mo agoHugging Face02sonalkum /AudioSkills-Llama3 AudioSkills-XL Dataset To promote the development of open source models, we have released AudioSkills using the exact same method generated with Llama 3.1-8B Instruct instead of GPT4o in the original. Project page | Paper | Code Dataset Description AudioSkills-XL is a large-scale audio question-answering (AQA) dataset designed to develop (large) audio-language models on expert-level reasoning and problem-solving tasks over short audio clips (≤30 seconds). It… See the full description on the dataset page: https://huggingface.co/datasets/sonalkum/AudioSkills-Llama3.textaudio-text-to-text100K<n<1M0 likes51 downloads1y agoHugging Face03SonalPrabhune /RealWorldQuestioning RealWorldQuestioning Benchmark RealWorldQuestioning is a benchmark dataset of 400+ real-world user questions collected from public discussion forums (e.g., Reddit, Quora), designed to support evaluation of gender bias and information disparity in Large Language Models (LLMs). The dataset spans four business-relevant domains: Education, Jobs, Investment, and Health. Each question is annotated with: User persona (Male or Female framing) Source forum Domain category Four anonymized… See the full description on the dataset page: https://huggingface.co/datasets/SonalPrabhune/RealWorldQuestioning.tabulartext-generationn<1K1 likes49 downloads1y agoHugging Face04sonal-ssj /MedExpert MedExpert MedExpert-Benchmark features clinician-created questions and detailed annotations designed to assess the accuracy, completeness, and reliability of LLM-generated medical responses. It comprises 540 question–response pairs across two distinct specialties: Young Adult Mental Health (MH) Prenatal Care (PC) Each sample is annotated by clinical subject-matter experts for factual accuracy, completeness (omissions), and model certainty. This dataset is designed to support… See the full description on the dataset page: https://huggingface.co/datasets/sonal-ssj/MedExpert.tabular1K<n<10K4 likes44 downloads10mo agoHugging Face05AADIMIND /sona-corpus THE SONA CORPUS — Noisy-to-Clean Hindi–English Parallel Dataset A clean, bilingual dataset card you can read at a glance and use immediately. Curated by: Aditya (AADIMIND) Languages: Hindi, English Total examples: 581312 (INPUT: 256 TOKEN• TARGET: 256 TOKEN) Tasks: Text cleaning, GEC, OCR post-processing, Seq2Seq fine-tuning License: MIT Source: Hindi Wikipedia (HiWiki) processed into noisy–clean pairs Repo: https://huggingface.co/datasets/AADIMIND/sona-corpus… See the full description on the dataset page: https://huggingface.co/datasets/AADIMIND/sona-corpus.texttext-generation100K<n<1M0 likes35 downloads1y agoHugging Face06sonasimon /LoFTI LoFTI: Localization and Factuality Transfer to Indian Locales LoFTI is a benchmark dataset that can be used to evaluate an LLM’s localization and factual text transfer capabilities. LoFTI consists of factual statements about entities in source and target locations; the source locations are spread across the globe and the target locations are all within India with varying degrees of hyperlocality (country, states, cities). The entities span a wide variety of categories.… See the full description on the dataset page: https://huggingface.co/datasets/sonasimon/LoFTI.texttext-generation1K<n<10K0 likes34 downloads2y agoHugging Face0711-47 /cakewalk_sonar_x2_god_producer Cakewalk Sonar X2 — God-Level Producer Dataset (8k) Train an LLM to become a god-level music producer in Cakewalk Sonar X2 Producer Edition. This dataset contains 4,060 high-quality examples covering every aspect of professional music production in Sonar X2: Virtual instrument programming (Dimension Pro, Rapture, Z3TA+, etc.) Plugin mastery (ProChannel, Sonitus suite, Console Emulator, etc.) Sound design & synthesis Advanced editing (AudioSnap, VocalSync, Step Sequencer) Mixing… See the full description on the dataset page: https://huggingface.co/datasets/11-47/cakewalk_sonar_x2_god_producer.text1K<n<10K0 likes34 downloads5mo agoHugging Face08thiagoambiel /sonar-municipal-pl-actions Sonar Municipal — PL Actions Corpus The largest publicly released Portuguese legal text rewriting dataset: 241,111 pairs mapping the original ementa (summary) of a Brazilian municipal Projeto de Lei (PL) to its action-form textualization (ação). Action form is a direct, imperative rewrite that strips juridical boilerplate, normalizes typography, and surfaces the underlying intervention rather than the legal instrument. Companion to: ICMC-USP undergraduate thesis "Mineração… See the full description on the dataset page: https://huggingface.co/datasets/thiagoambiel/sonar-municipal-pl-actions.tabularsummarization100K<n<1M0 likes33 downloads4mo agoHugging Face0911-47 /cakewalk_sonar_8_god_producer_dataset Cakewalk Sonar 8 Producer Edition - God Level Producer Dataset The ultimate training dataset for becoming a god-level producer in Cakewalk Sonar 8 Producer Edition. This dataset trains LLMs to master every aspect of Sonar 8 Producer Edition at the highest professional level: All virtual instruments (Dimension Pro, Rapture, Pentagon I, etc.) ProChannel (Console Emulator, EQ, Compressor, Tube Saturation) Complete mixing workflows Commercial mastering chains (LP-64, Sonitus:fx)… See the full description on the dataset page: https://huggingface.co/datasets/11-47/cakewalk_sonar_8_god_producer_dataset.textn<1K0 likes19 downloads5mo agoHugging Face10sonalsannigrahi /mls_entabular10K<n<100K0 likes16 downloads3mo agoHugging Face11sonawat-sambudh-tft /ml-intern-llama-sft-data ML-Intern Llama SFT data Prepared local session logs for TRL SFTTrainer. Rows use conversational prompt/completion format. textn<1K0 likes10 downloads5mo agoHugging Face12soendup21 /sonamdatatabular1K<n<10K0 likes8 downloads2y agoHugging Face13sonalsannigrahi /mls_pttabular10K<n<100K0 likes8 downloads3mo agoHugging Face14sonalsannigrahi /yodas_sampletabular10K<n<100K0 likes7 downloads3mo agoHugging Face15sonasb /VIDHI-Indian_Law_Consumer_Enrichedtext10K<n<100K0 likes5 downloads1y agoHugging Face16sonagihu /jtdatasettextn<1K0 likes2 downloads2y agoHugging Face17sonamtenzey /instruction_dataset-edu-aitextn<1K0 likes2 downloads1y agoHugging Face18sonawilson /marvel-characters-datasettextn<1K0 likes2 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.