CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01its5Q /biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio. audiotext-to-speech100K<n<1M23 likes1.3k downloads1y agoHugging Face02Itsuki-music /BACHI_Chord_Recognition BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music Paper | Project Page | Code | POP909-CL Dataset This repository contains trained model weights and classical datasets for the paper: Mingyang Yao, Ke Chen, Shlomo Dubnov and Taylor Berg-Kirkpatrick "BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music."ICASSP 2026, 2025 Abstract Automatic chord… See the full description on the dataset page: https://huggingface.co/datasets/Itsuki-music/BACHI_Chord_Recognition.textothern<1K7 likes983 downloads2mo agoHugging Face03xX-its-amit-Xx /pxr-structure-pose-pool PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands) Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2) structure-prediction challenge, released openly with per-pose labels so the community can reuse the compute already spent — and, we hope, crack the problem this data makes visible. What's here poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand (resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.tabular1K<n<10K0 likes957 downloads2mo agoHugging Face04ItsNotRohit /Food121 Dataset Details Dataset Description This dataset is the combination of the Food101, Indian Food Classification and The-massive-Indian-Food-Dataset datasets. This Dataset aims to be a viable dataset for Image Classification of Foods with an added Indian context. This dataset has 121 classes with each class having 800 images in the train split and 200 images in the test split. Maximum resolution of images is 512*512. The Food121-224 dataset has all images downscaled to a… See the full description on the dataset page: https://huggingface.co/datasets/ItsNotRohit/Food121.imageimage-classification100K<n<1M1 likes266 downloads3y agoHugging Face05its5Q /wikireading Dataset Card for Wikireading This is a dataset of book chapters scraped from a Russian website called Wikireading. Dataset Details Dataset Description Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining. The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.texttext-generation1M<n<10M9 likes236 downloads2y agoHugging Face06itsakhilyou /FinSearchCompThis repository contains the FinSearchComp dataset, a benchmark for evaluating financial search and reasoning capabilities of LLM-based agents, as presented in the paper FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning. Project Page: https://randomtutu.github.io/FinSearchComp/ FinSearchComp is the first fully open-source agent benchmark designed for realistic, open-domain financial search and reasoning. It comprises three tasks that closely… See the full description on the dataset page: https://huggingface.co/datasets/itsakhilyou/FinSearchComp.textquestion-answeringn<1K0 likes214 downloads5mo agoHugging Face07itsluketwist /NotAllCodeIsEqual NotAllCodeIsEqual This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning. It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities. We provide 2 types of dataset, that cover complementary settings: CodeNet (solution-driven complexity): The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.tabulartext-generation100K<n<1M0 likes213 downloads8mo agoHugging Face08itsG /smishing-synthetictextn<1K0 likes152 downloads1y agoHugging Face09benjaminmacklin /IT_Support_V2 Mack: IT Support & Admin Dataset 📋 Dataset Description This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent. The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging. Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.texttext-generation100K<n<1M2 likes127 downloads10mo agoHugging Face10itsrishub /synthetic-logs Synthetic Logs (Wild) Just messy, realistic-looking logs paired with their parsed version. Each row has a raw log line and what you'd want a parser to pull out of it. 100,000 rows total, split across 3 files in data/: data/logs-0001.parquet data/logs-0002.parquet data/logs-0003.parquet 2 columns: raw_log (the messy string) and parsed_json (the answer, as JSON string) Covers 130+ services — nginx, postgres, k8s, lambda, python tracebacks, etc. — in 16 formats like syslog, JSON… See the full description on the dataset page: https://huggingface.co/datasets/itsrishub/synthetic-logs.texttext-generation100K<n<1M0 likes106 downloads7d agoHugging Face11Lots-of-LoRAs /task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists.texttext-generation1K<n<10K0 likes104 downloads2y agoHugging Face12itskamo-com /sepediaudio10K<n<100K2 likes102 downloads2y agoHugging Face13ItsMaxNorm /MedAgentSim-datasets MedAgentSim Datasets GitHub: https://github.com/MAXNORM8650/MedAgentSimWebsite: https://medagentsim.netlify.app This repository contains various datasets used in the MedAgentSim project for simulating medical agent interactions. Datasets Included Dataset Rows Description medqa_v1.parquet 107 General medical question-answering OSCE examinations medqa_extended_v1.parquet 214 Extended medical QA with comprehensive coverage mimiciv_v1.parquet 288 Patient… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/MedAgentSim-datasets.textquestion-answeringn<1K1 likes102 downloads6mo agoHugging Face14alezzandro /itsm_ticketstext1K<n<10K0 likes97 downloads2y agoHugging Face15ItsNotRohit /Food121-224 Dataset Details Dataset Description This dataset is the downscaled version of the Food121 dataset. All images are downscaled to a maximum of 224*224. This dataset is the combination of the Food101, Indian Food Classification and The-massive-Indian-Food-Dataset datasets. This Dataset aims to be a viable dataset for Image Classification of Foods with an added Indian context. This dataset has 121 classes with each class having 800 images in the train split and 200 images in… See the full description on the dataset page: https://huggingface.co/datasets/ItsNotRohit/Food121-224.imageimage-classification100K<n<1M2 likes95 downloads3y agoHugging Face16itsyoboieltr /pcbimage1K<n<10K4 likes89 downloads3y agoHugging Face17ITS-23-24 /draft_nbaimage1K<n<10K0 likes85 downloads2y agoHugging Face18albaz2000 /arabic-itsm-dataset Arabic ITSM Dataset A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release. Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.tabulartext-classification10K<n<100K0 likes81 downloads20d agoHugging Face19itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes80 downloads1mo agoHugging Face20itsgupta /proper-agents-data ProPer Agents — data Data for ProPer Agents: Proactivity Driven Personalized Agents for Advancing Knowledge Gap Navigation (ACL 2026). Paper · Adapters Three domains: code, medical, pwab (product recommendation). Layout {domain}/ raw/train.jsonl source examples raw/test.jsonl raw/{domain}_rga_{train,test}.jsonl RGA SFT data (Alpaca format) raw/{domain}_dga_{train,test}.jsonl DGA SFT data (Alpaca format)… See the full description on the dataset page: https://huggingface.co/datasets/itsgupta/proper-agents-data.texttext-generation1K<n<10K0 likes78 downloads2mo agoHugging Face21ItsMaxNorm /ultralivecell Cell Segmentation Masks Dataset This repository contains NumPy mask files generated using AutoSeg-SAM2 for cell segmentation experiments. Dataset Structure The dataset is organized by experiment type and image number: {experiment_type}/ └── image_{number}/ └── cached_masks.npy Experiments Included Cdx2_Gata6_Oct4 Images: image_2, image_3, image_4 Sox2_Sox17 Images: image_1, image_2, image_3 Usage To load the NumPy mask… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/ultralivecell.imagen<1K0 likes74 downloads1y agoHugging Face22AI4Manufacturing /ITSCgated ITSC Park-vector current loci of a three-phase induction motor, as a four-class stator winding task: normal / phase_A / phase_B / phase_C. 183 records. reasoning is empty here; the twin repo ITSC-annotated is identical except that field is filled. The reading Both coordinates are divided by the radius of the equal-area circle, so that circle is the 1.0 ring on every image and the tick numbers carry no current amplitude at all. Two steps, and both are drawn on… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/ITSC.imageimage-classificationn<1K0 likes74 downloads16h agoHugging Face23AI4Manufacturing /ITSC-perceptiongated Roles Roles: perception view of ITSC — annot is the source label (normal / phase_A / phase_B / phase_C), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the repo ships an induction motor's stator-current magnitude in four image encodings as four equal-sized configs — reshaped (consecutive samples arranged as the rows of a grayscale square), scalogram (a continuous-wavelet time-scale view), spectrogram (a short-time… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/ITSC-perception.imageimage-classificationn<1K0 likes73 downloads18h agoHugging Face24ITS23 /TACK_Tunnel_Data TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/ITS23/TACK_Tunnel_Data.image1K<n<10K0 likes65 downloads7mo agoHugging Face25itsbib /nepali-asr-processedtext100K<n<1M0 likes63 downloads2y agoHugging Face26aakash0017 /it-support-llmtext1K<n<10K3 likes62 downloads3y agoHugging Face27itseffi /epfl-enterprise-osai-adoption-research-data EPFL Enterprise Open-Source AI Adoption Research Dataset Dataset Summary This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption. Dataset Structure This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.tabulartext-classificationn<1K0 likes62 downloads1y agoHugging Face28itsZyn /ConvES30K ConvES10K-HQ-LLM — High-Quality Spanish Conversations (LLM-Generated) Recommended for AI training. This version replaces the template-based builds and is fully LLM-generated for true independence. Why this version Previous builds (30K and 10K-template) were template-based (gen.py + 178 situations): 13,156 distinct messages from 74,526 total → 82.3% duplicate messages, max x109 on narrative blocks (Mesa seis pegada a la mesa siete...) Coherence failures from… See the full description on the dataset page: https://huggingface.co/datasets/itsZyn/ConvES30K.texttext-generation10K<n<100K0 likes60 downloads29d agoHugging Face29ItsMaxNorm /privasis-reasoning-qa Privasis Reasoning-QA Open-ended reasoning question–answer pairs derived from the NVIDIA Privasis-Zero dataset. Two configs are provided: qa50k — 50,000 pairs sampled from the Privasis-Zero corpus split (record field). Main set. qa500 — 500 pairs from the hard_test split (original_record field). Original pilot. from datasets import load_dataset ds = load_dataset("ItsMaxNorm/privasis-reasoning-qa", "qa50k", split="train") Each item presents one question that requires… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/privasis-reasoning-qa.textquestion-answering10K<n<100K0 likes59 downloads3mo agoHugging Face30its5Q /bigger-ru-bookaudio10K<n<100K13 likes58 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.