CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01khanh2023 /prop_logic_puzzletext100K<n<1M0 likes5.6k downloads5mo agoHugging Face02khaled123 /dataioiemetext100K<n<1M0 likes1.4k downloads2y agoHugging Face03khalidalt /HuffPostA dataset of approximately 200K news headlines from the year 2012 to 2018 collected from HuffPost.text100K<n<1M2 likes1.4k downloads3y agoHugging Face04khaled123 /dataset-metext100K<n<1M0 likes1.3k downloads2y agoHugging Face05khamidov17 /ds-84149598ff52f2caaudio10K<n<100K0 likes1.2k downloads2mo agoHugging Face06khamidov17 /ds-5540a63e610e2a90audio1K<n<10K1 likes1k downloads2mo agoHugging Face07khaled123 /dataset-nametext1M<n<10M0 likes978 downloads2y agoHugging Face08khaled123 /datametext100K<n<1M0 likes871 downloads2y agoHugging Face09khanhvinh9 /imagenet-cimage1M<n<10M0 likes819 downloads9mo agoHugging Face10khangmacon /cybermetric-10000 Dataset Card for "cybermetric-10000" More Information needed text10K<n<100K1 likes587 downloads2y agoHugging Face11khaimaitien /multi-hop-qa-function-calling-format-V1.0This dataset is converted from khaimaitien/qa-expert-multi-hop-qa-V1.0 to OpenAI function calling format. Each data point is a list of messages with role=user, assistant or function: message that role=user, content is the question message that role=assistant, content is not None, function_call is None: --> assistant responds with text only message that role=assistant and function_call is not None --> assistant asks to execute a function call function_call is of the form: {"name": "retrieve"… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/multi-hop-qa-function-calling-format-V1.0.text10K<n<100K10 likes450 downloads3y agoHugging Face12KhaledReda /pairs_with_scores_v27text100M<n<1B0 likes385 downloads8mo agoHugging Face13khairi /kothar-dataset-it Kothar fine-tuning datasets This repo documents the instruction-tuning datasets built by this repository for a protein language model. All datasets are seeded from the Neo4j protein knowledge graph (docs/neo4j_schema.md) or from computed sequence features, built by the scripts under scripts/python/, and published as HuggingFace Hub configs at khairi/kothar-dataset-it. What's here Document Covers shared-conventions.md Sequence encoding, the computed… See the full description on the dataset page: https://huggingface.co/datasets/khairi/kothar-dataset-it.text1M<n<10M1 likes363 downloads26d agoHugging Face14Khanh14ph /asr-youtube-datasettext100K<n<1M0 likes331 downloads1y agoHugging Face15khaled123 /datasenametext100K<n<1M0 likes285 downloads2y agoHugging Face16adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes277 downloads13d agoHugging Face17ahmedheakl /arocrbench_khattKITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding This dataset is designed to evaluate the performance of Arabic OCR and document understanding systems. It includes a variety of document types and tasks. Please see paper & code for more information: GitHub Repository Project Page arXiv Paper imageimage-to-textn<1K2 likes255 downloads1y agoHugging Face18khang119966 /InternVL_Chat_V12_SFT_Dataimage1M<n<10M0 likes254 downloads2y agoHugging Face19khangmacon /cyberQA Dataset Card for "cyberQA" More Information needed text10K<n<100K1 likes251 downloads2y agoHugging Face20khadijah00 /ppe-benchmark-eval PPE Benchmark Eval Set (v1) A held-out, human-verified benchmark for evaluating vision-language models on personal protective equipment (PPE) detection — specifically hardhat and safety-vest presence — framed as a VQA-style classification task. What this is 96 images, balanced 24/24/24/24 across the four hardhat × vest combinations (yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of the karabuk-university PPE dataset on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.imagevisual-question-answeringn<1K0 likes237 downloads1mo agoHugging Face21wannaphong /KhanomTanLLM-pretrained-dataset KhanomTanLLM pretrained dataset This daataset collect all raw text for pretraining LLM. Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Tokens 53,376,211,711 Tokens English: 31,629,984,243 Tokens Thai: 12,785,565,497 Tokens Code: 8,913,084,300 Toekns Parallel data: 190,310,686 Tokens Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer All subset Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.texttext-generation10M<n<100M1 likes231 downloads2y agoHugging Face22KhalfounMehdi /arabic-latin-invoices-synthetic Synthetic Arabic/Latin Invoices — label-first multimodal dataset A large, perfectly-labeled synthetic invoice dataset for training multimodal invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin layouts, and all three numeral glyph systems. Every sample is generated label-first: a structured invoice record is sampled, then rendered to pixels via a headless browser, and bounding boxes are read back from the same DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.imageimage-to-text1K<n<10K1 likes226 downloads3mo agoHugging Face23khalidalt /openai_mmlu_arabic Dataset Card text10K<n<100K0 likes225 downloads2y agoHugging Face24oddadmix /arabic-audio-collection-mohamed-khairy Mohamed Khairy Arabic Speech Dataset Dataset Summary The Mohamed Khairy Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 430 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mohamed-khairy.audiotext-to-speech10K<n<100K7 likes222 downloads3mo agoHugging Face25Khanhpham1992 /es-futures-1mtabular1M<n<10M0 likes214 downloads5mo agoHugging Face26adeshkin /khakas-russian-dict Khakas-Russian Dictionary (Dataset) Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu 📌 Important Notice & Citation When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/. Please note: Some optical character recognition (OCR) errors may still be present in the data. 🛠 Contribution & Authorship I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.tabulartranslation10K<n<100K0 likes213 downloads5mo agoHugging Face27khashazad /amc-ruler-qwen35-32k AMC RULER 32k This dataset contains frozen inputs for the RULER benchmark. Generation metadata Benchmark: RULER Sequence length: 32,768 tokens Tokenizer: Qwen/Qwen3.5-9B Tokenizer revision: c202236 lm-eval version: 0.4.12 Task configurations: 13 Samples per configuration: 500 Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42. Task configurations niah_single_1 niah_single_2 niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-32k.tabular1K<n<10K0 likes206 downloads1mo agoHugging Face28khashazad /amc-ruler-qwen35-16k AMC RULER 16k This dataset contains frozen inputs for the RULER benchmark. Generation metadata Benchmark: RULER Sequence length: 16,384 tokens Tokenizer: Qwen/Qwen3.5-9B Tokenizer revision: c202236 lm-eval version: 0.4.12 Task configurations: 13 Samples per configuration: 500 Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42. Task configurations niah_single_1 niah_single_2 niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-16k.tabular1K<n<10K0 likes205 downloads1mo agoHugging Face29khaclinh /pp4avPP4AV is the first public dataset with faces and license plates annotated with driving scenarios. P4AV provides 3,447 annotated driving images for both faces and license plates. For normal camera data, dataset sampled images from the existing videos in which cameras were mounted in moving vehicles, running around the European cities. The images in PP4AV were sampled from 6 European cities at various times of day, including nighttime. This dataset use the fisheye images from the WoodScape dataset to select 244 images from the front, rear, left, and right cameras for fisheye camera data. PP4AV dataset can be used as a benchmark suite (evaluating dataset) for data anonymization models in autonomous driving.imageobject-detection1K<n<10K5 likes203 downloads4y agoHugging Face30jtatman /cosmopedia-openstax-khanacademy-150k-sharegpt Dataset Card for "cosmopedia-openstax-khanacademy-150k-sharegpt" More Information needed text100K<n<1M2 likes201 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.