CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aak975 /iclr-wm-backup-public ICLR Watermark Benchmark — backup overflow (public part) Companion to the private repo Aak975/iclr-wm-backup, which reached its storage quota. Together the two repos form ONE backup — every file exists in exactly one of them, with the same layout: archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json restore one archive: cat part-* | zstd -d | tar -x MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M} This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.tabularn<1K0 likes179k downloads21d agoHugging Face02aakkaasshh /mysdxl-dataset Image-Prompt Dataset An image-prompt dataset scraped and assembled with MySDXL for training latent diffusion models. Dataset structure Each row contains one image with its corresponding text prompt. Column Type Description image Image RGB image (lossless PNG, original resolution) prompt string Text prompt describing the image negative_prompt string Negative prompt (empty string if none) Stored as Parquet shards (data/train-*.parquet). Load… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/mysdxl-dataset.imagetext-to-image1K<n<10K0 likes231 downloads3mo agoHugging Face03aakkaasshh /vaigai-dataset Vaigai Dataset aakkaasshh/vaigai-dataset is an Indic-focused multilingual text corpus aggregated for tokenizer training, built with the goal of reaching tokenizer quality on par with dedicated Indic tokenizers (e.g. Sarvam AI's). One Parquet file per language is stored at the root of this repo (e.g. bn.parquet), with no nested folders. Every time new data is fetched for a language, it is merged with that language's existing file and deduplicated on the text column, so re-running… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/vaigai-dataset.tabulartext-generation1M<n<10M0 likes190 downloads2mo agoHugging Face04aakashMeghwar01 /sindhi-corpus-505m Sindhi Corpus 505M The largest open-source, deduplicated Sindhi language pretraining corpus. ~505 million tokens across 742K documents, covering news, literature, legal, religious, encyclopedic, and web-crawled Sindhi text. Built for training Sindhi language models, tokenizers, and NLP tools. Dataset Summary Stat Value Documents ~742,379 Tokens (estimated) ~505 million Language Sindhi (sd) — Arabic script Format Parquet (single text column)… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/sindhi-corpus-505m.texttext-generation100K<n<1M2 likes170 downloads3mo agoHugging Face05aakanksha19 /cdr_bigbio_processedtext1K<n<10K1 likes121 downloads3y agoHugging Face06aakash-projects /colpali_train_set Dataset Description This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%). Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages. Dataset #examples (query-page pairs) Language DocVQA… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/colpali_train_set.imagedocument-question-answering100K<n<1M0 likes99 downloads3mo agoHugging Face07aakashMeghwar01 /sindhi-corpus-langid-cleantext100K<n<1M0 likes61 downloads3mo agoHugging Face08aakash0017 /it-support-llmtext1K<n<10K3 likes59 downloads3y agoHugging Face09aakashMeghwar01 /Sindhi-Intelligence-Core-SFT 🧠 Sindhi Intelligence Core SFT This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning. 📊 Dataset Summary This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT). 📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.texttext-generation100K<n<1M1 likes58 downloads7mo agoHugging Face10Aakali /jeeaudion<1K0 likes54 downloads2y agoHugging Face11aakarsh-nair /semeval-2025-task-8-test-casestext1K<n<10K0 likes37 downloads2y agoHugging Face12Aakash22134 /hi-en-noisy-vad-benchmark Hindi-English Noisy VAD Benchmark Version 0.1.0 is a deterministic, evaluation-only benchmark with 78 mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB). Intended use Use this dataset to compare voice-activity detectors under matched Hindi/English noise conditions and to tune thresholds. It is too small and insufficiently diverse for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.audioaudio-classificationn<1K0 likes37 downloads1mo agoHugging Face13aakarsh-nair /semeval-2025-task-8-test-cases-competitiontextn<1K0 likes36 downloads2y agoHugging Face14aakanksha180 /Indian_Railway_maintance license: cc-by-4.0 task_categories: - tabular-classification - tabular-regression - time-series language: - en tags: - railway - predictive-maintenance - failure-detection - transportation - iot - machine-learning - synthetic-data - analytics pretty_name: Indian Railway Failure Detection & Maintenance (100K) size_categories: - 100K<n<1M 🚆 Indian Railway Failure Detection & Maintenance (100K) Overview This dataset contains 100,000 synthetic yet… See the full description on the dataset page: https://huggingface.co/datasets/aakanksha180/Indian_Railway_maintance.tabular100K<n<1M0 likes36 downloads21d agoHugging Face15aakanksha19 /pico_bigbio_processedtext1K<n<10K0 likes35 downloads3y agoHugging Face16aakashmallik /research-paper-agent-reasoning-traces-unverifiedtextn<1K0 likes31 downloads2mo agoHugging Face17aakash-projects /biomedical_lectures_eng_v2 Vidore Benchmark 2 - MIT Dataset This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions). Dataset Summary Each query is in english. This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/biomedical_lectures_eng_v2.imagedocument-question-answering1K<n<10K0 likes23 downloads3mo agoHugging Face18aakashMeghwar01 /sindhi-corpus-cleantext100K<n<1M0 likes22 downloads3mo agoHugging Face19aakash-agarwal /ReDepressgated ReDepress Dataset The ReDepress Dataset originates from the paper"ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media". This dataset is designed to support research on detecting depression relapse using social media text.It provides rich temporal user data and cognitive bias annotations, allowing systems to monitor conversational dynamics and make timely inferences. Important Notice By requesting, accessing, downloading, or using this… See the full description on the dataset page: https://huggingface.co/datasets/aakash-agarwal/ReDepress.tabular1K<n<10K0 likes21 downloads7mo agoHugging Face20aakash0017 /It-support-synthetic-datatext10K<n<100K2 likes18 downloads3y agoHugging Face21Aakash941 /THAR-DatasetThe dataset consists 11,549 YouTube comments in Hindi-English code-mixed language for targeted hate speech detection against religion. Binary and multi-class tagging of YouTube comments is used. The classification of YouTube comments addresses two subtasks: Subtask-1 (Binary classification): comments are labeled as antireligion or non-antireligion. Subtask-2 (Multi-class classification): comments are labeled on the major targeted religions such as Islam, Hinduism, and Christianity, with a… See the full description on the dataset page: https://huggingface.co/datasets/Aakash941/THAR-Dataset.texttext-classification10K<n<100K1 likes18 downloads2y agoHugging Face22aakarsh-nair /semeval-2025-task-8-prompts-competitiontextn<1K0 likes18 downloads2y agoHugging Face23aakashMeghwar01 /aurat-march-processed-newstext1K<n<10K0 likes18 downloads7mo agoHugging Face24aakashMeghwar01 /english-aurat-march-newstextn<1K0 likes16 downloads5mo agoHugging Face25aakashMeghwar01 /Sindhi-Intelligence-Core-SFT-v2text10K<n<100K0 likes15 downloads7mo agoHugging Face26Aakali /commn-voice-11-translatedaudio1K<n<10K0 likes14 downloads2y agoHugging Face27aakash-projects /arxivqa_test_subsampled Dataset Description This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. The questions were generated synthetically using GPT-4 Vision. Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs. Furthermore we renamed the different columns for our purpose. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/arxivqa_test_subsampled.imagedocument-question-answeringn<1K0 likes13 downloads3mo agoHugging Face28aakashverma0000 /Testingtabularn<1K0 likes13 downloads2mo agoHugging Face29AAkhoram /Persian-Wikipedia-Corpus Overview This dataset is derived from the Persian Wikipedia Corpus project, which contains parsed articles from the Persian Wikipedia. The original data has been converted into a more accessible format and made available through the HuggingFace datasets library. Usage from datasets import load_dataset dataset = load_dataset("codersan/Persian-Wikipedia-Corpus") Persian-Wikipedia-Corpus A complete copy of Persian Wikimedia pages, The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AAkhoram/Persian-Wikipedia-Corpus.tabulartext-generation1M<n<10M0 likes13 downloads2mo agoHugging Face30aakarsh-nair /semeval-2025-task-8-prompts-traintextn<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.