CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CAiRE /ASCEND Dataset Card for ASCEND Dataset Summary ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/ASCEND.audioautomatic-speech-recognition10K<n<100K53 likes1.9k downloads2y agoHugging Face02beta3 /ASCII_Alphabet_Dataset_571_Fonts Dataset Description This dataset provides programmatically generated ASCII representations of the English alphabet rendered using 571 fonts from the PyFiglet library. Each letter (A–Z) is available in multiple typographic styles, resulting in a structured and high-variability dataset suitable for research, experimentation, and creative applications. The dataset was created to support tasks involving text-based pattern recognition, synthetic data generation, typography analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/beta3/ASCII_Alphabet_Dataset_571_Fonts.texttext-classification1 likes1.3k downloads7mo agoHugging Face03starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes967 downloads2y agoHugging Face04tuanphong /ascent_kb Dataset Card for Ascent KB Dataset Summary This dataset contains 8.9M commonsense assertions extracted by the Ascent pipeline developed at the Max Planck Institute for Informatics. The focus of this dataset is on everyday concepts such as elephant, car, laptop, etc. The current version of Ascent KB (v1.0.0) is approximately 19 times larger than ConceptNet (note that, in this comparison, non-commonsense knowledge in ConceptNet such as lexical relations is excluded). For… See the full description on the dataset page: https://huggingface.co/datasets/tuanphong/ascent_kb.textother10M<n<100M4 likes653 downloads3y agoHugging Face05ytzi /the-stack-dedup-python-filtered-non_asciiThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_non_ascii tabular10M<n<100M1 likes488 downloads2y agoHugging Face06apehex /ascii-art ASCII Art Description This is a text-to-image dataset, where the images are actually ASCII art. The ASCII arts come from various sources: asciiart: made by independent artists, listed on asciiart.eu copypasta: common twitch emotes, listed on twitchquotes.com graffiti: text samples styled using various ASCII art fonts with a tool images: conversion of a portion of the dataset DataCompDR-12M using a tool Metadata homepage:… See the full description on the dataset page: https://huggingface.co/datasets/apehex/ascii-art.text10K<n<100K5 likes319 downloads1y agoHugging Face07AscendKernelGen /Ascend-COT-v2-json AscendKernelGen/Ascend-COT-v2-json AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.texttext-generation10K<n<100K3 likes294 downloads5mo agoHugging Face08ASCCCCCCCC /amazon_zhthis is a datasets about amazon reviews text100K<n<1M2 likes260 downloads5y agoHugging Face09cyrilzhang /TinyStories2-ascii Dataset Card for "TinyStories2-ascii" TinyStoriesV2-GPT4-{train,validation}.txt from roneneldan/TinyStories ad-hoc Unicode -> ASCII normalization remove empty/incomplete stories text1M<n<10M1 likes260 downloads3y agoHugging Face10apehex /ascii-art-datacompdr-12m ASCII Art DataCompDR-12M Description This is a text-to-image dataset, where the images are actually ASCII art. The images and captions were sampled from DataCompDR-12M. The conversion was performed with the tool ascii-image-converter. Metadata homepage: https://github.com/apehex/scrapscii version: 0.1.0 Config Split Size Samples 'default' 'train' 4.1 GB 643072 'default' 'fixed' 552 MB 262144 The ASCII art in "fixed" all have a width of 64… See the full description on the dataset page: https://huggingface.co/datasets/apehex/ascii-art-datacompdr-12m.text100K<n<1M0 likes210 downloads1y agoHugging Face11WTFO /ascend_MIXED_cleaned_vadaudio1K<n<10K0 likes205 downloads4mo agoHugging Face12YuvrajSingh9886 /asciitermdraw-bench-public ASCIITermDraw-Bench — Public Examples 12 public example tasks from ASCIITermDraw-Bench, a benchmark for evaluating whether language models can generate and edit structured ASCII diagrams. The full benchmark has 80 private, held-out tasks used for actual scoring — those are not distributed here. This dataset is a separate, hand-authored set of 12 tasks (one easy, one medium, one hard per category) in the exact same format, so anyone can see what a task looks like and run the… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/asciitermdraw-bench-public.imageimage-to-textn<1K1 likes201 downloads5d agoHugging Face13Smith42 /ascl-code ASCL Astronomy Source Code The Astrophysics Source Code Library (ASCL) is a curated registry of source code used in astronomy and astrophysics research. This dataset contains source files extracted from ASCL-listed repositories, paired with catalog metadata. Dataset Structure Manifest (manifest.parquet) One row per ASCL catalog entry with the following fields: Field Description ascl_id ASCL identifier (e.g., [ascl:2306.019]) title Software title… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/ascl-code.texttext-generation100K<n<1M0 likes199 downloads6mo agoHugging Face14WTFO /ascend_ZH_cleaned_vadaudio10K<n<100K0 likes199 downloads4mo agoHugging Face15WTFO /ascend_EN_cleaned_vadaudio1K<n<10K0 likes193 downloads4mo agoHugging Face16Lots-of-LoRAs /task1148_maximum_ascii_value Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1148_maximum_ascii_value Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1148_maximum_ascii_value.texttext-generationn<1K0 likes180 downloads2y agoHugging Face17AscendKernelGen /Ascend-CoT-v3-json Ascend-CoT-v3-json Ascend-CoT-v3-json is an Ascend C / CANN supervised fine-tuning dataset for custom operator development. It contains cleaned CoT-style samples for Ascend C kernel implementation, tiling logic, CANN API usage, debugging, and operator-development reasoning. The release is organized into two final SFT subsets in one dataset repository. Related Artifacts Paper: AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-CoT-v3-json.texttext-generation100K<n<1M2 likes173 downloads4mo agoHugging Face18ASCCCCCCCC /amazon_zh_simpletext10K<n<100K1 likes159 downloads5y agoHugging Face19cyrilzhang /TinyStories-ascii TinyStories-{train,validation}.txt from roneneldan/TinyStories ad-hoc Unicode -> ASCII normalization remove empty/incomplete stories text1M<n<10M0 likes136 downloads3y agoHugging Face20webxos /3d-ascii-typeface-v1 3d_ascii_typeface_dataset Sampler https://github.com/webxos for more info. Generated by ASCII Dataset Studio v2.0 (by webXOS) Contains 155 ASCII art samples. A full typeface dataset in 3D ASCII Format. Dataset Structure file_name: Path to the ASCII text file in the data/ folder. text: The ASCII art content as a string. Generation Synthetically generated using multiple modes: image-to-ASCII, text-to-ASCII (fonts), geometric patterns, and video… See the full description on the dataset page: https://huggingface.co/datasets/webxos/3d-ascii-typeface-v1.textimage-to-text1K<n<10K1 likes136 downloads29d agoHugging Face21mondk /ASCII-art-videoYou can ask an AI to write code to view the video as fps (frame by frame). format: ▍ KHUNG 1/18 image_here ·················································· ▍ KHUNG 2/18 image_here ·················································· ▍ KHUNG 3/18 image_here ·················································· HUGGING FACE!!… See the full description on the dataset page: https://huggingface.co/datasets/mondk/ASCII-art-video.text10K<n<100K2 likes131 downloads22d agoHugging Face22AsciiMAster /pl-web-graph-2026-09-14 Polish Web Domain Observations 2026-09-14 A curated snapshot of a .pl-focused domain crawler: DNS observations, host availability metadata, discovered URL references, and the crawl frontier. No HTML, page text, or website classifications are included. Observations accumulated over months, so September 14 dates the export itself while each row carries its own observation time. Coverage is whatever one crawler reached, and liveness holds as of the recorded timestamp. Rendered… See the full description on the dataset page: https://huggingface.co/datasets/AsciiMAster/pl-web-graph-2026-09-14.tabular100M<n<1B0 likes123 downloads10d agoHugging Face23PeiyangLiu /ascp-context-attribution ASCP: Causal Context Attribution and Probe Benchmark Released artifacts for The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search. 📄 Paper: https://arxiv.org/abs/2608.23252 💻 Code: https://github.com/PeiYangLiu/ascp Retrieval-augmented generation is usually measured with relevance proxies — BM25, query–document cosine, output overlap — that score how related a passage looks, not whether the generator used it. This dataset ships… See the full description on the dataset page: https://huggingface.co/datasets/PeiyangLiu/ascp-context-attribution.tabularquestion-answering10K<n<100K0 likes117 downloads1mo agoHugging Face24mondk /ASCII-art-images-for-trainSource: https://data.caltech.edu/records/mzrjq-6wc02 I converted these into ASCII art. Link to download the original PNG files if you need them: https://data.caltech.edu/records/mzrjq-6wc02/files/caltech-101.zip?download=1 The txt files can be used to train a text-gen-only model that can "see" images in the form of ASCII art. ty! text100K<n<1M2 likes105 downloads23d agoHugging Face25Antix5 /vi-gym-causal-ascii Vi-Gym Causal ASCII Trajectories This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment. Warning This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership. Dataset Structure Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.texttext-generation100K<n<1M0 likes82 downloads7mo agoHugging Face26ASCIIEval /ASCIIEval ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art 📖 Arxiv | 🤗 ASCIIEval Dataset | 🤗 ASCIITune Dataset TABLE OF CONTENTS Introduction Data Leaderboards Leaderboard for Textual Input Leaderboard for Image Input Leaderboard for Average Cross-Modality Performance Citation Introduction Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and… See the full description on the dataset page: https://huggingface.co/datasets/ASCIIEval/ASCIIEval.textvisual-question-answering1K<n<10K0 likes81 downloads10mo agoHugging Face27georgechang8 /ASCEND_CLEAN Dataset Card for Dataset Name This dataset is derived from CAiRE/ASCEND. More information is available at https://huggingface.co/datasets/CAiRE/ASCEND. Removed 嗯 呃 um uh Resolved [UNK]'s using whisper-medium Usage Default utterances with cleaned transcripts from datasets import load_dataset data = load_dataset("georgechang8/ASCEND_CLEAN") # add split="train" for train set, etc. Concatenated 30s utterances with cleaned transcripts… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/ASCEND_CLEAN.audio10K<n<100K0 likes77 downloads2y agoHugging Face28leideng /nanochat-ascend-dataset nanochat-ascend-dataset Unified training and evaluation data bundle for nanochat-ascend. This repository is designed to make the nanochat-ascend training procedure easy to reproduce. Instead of asking users to collect multiple task and evaluation datasets and manually reconstruct the expected directory structure, this repository preserves the local filesystem layout expected by the training code. The intended usage is simple: place this repository at .cache/dataset download… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-dataset.texttext-generation10K<n<100K0 likes71 downloads6mo agoHugging Face29AsciiMAster /uke-mobile-network-permits-poland UKE Mobile Network Permits - Poland Radio permits issued by the Polish Office of Electronic Communications (Urząd Komunikacji Elektronicznej, UKE) for cellular base stations across all bands: GSM 900/1800, GSM-R, UMTS 900/2100, LTE 420/450/700/800/900/1800/2100/2600, 5G 700/900/1800/2100/2600/3600, and CDMA 420. The dataset is a single GeoParquet file in WGS84 (EPSG:4326) containing one row per permit per station per band. Source:… See the full description on the dataset page: https://huggingface.co/datasets/AsciiMAster/uke-mobile-network-permits-poland.tabulartabular-classification100K<n<1M1 likes63 downloads4mo agoHugging Face30ednalaxer /datacomp_small_clip1_30pct_asciichr_greater_than_4 Dataset Card for "datacomp_small_clip1_30pct_asciichr_greater_than_4" More Information needed image1M<n<10M0 likes56 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.