datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASCEND
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/ASCEND.ASCII_Alphabet_Dataset_571_Fonts
Dataset Description
This dataset provides programmatically generated ASCII representations of the English alphabet rendered using 571 fonts from the PyFiglet library. Each letter (A–Z) is available in multiple typographic styles, resulting in a structured and high-variability dataset suitable for research, experimentation, and creative applications.
The dataset was created to support tasks involving text-based pattern recognition, synthetic data generation, typography analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/beta3/ASCII_Alphabet_Dataset_571_Fonts.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.ascent_kb
Dataset Card for Ascent KB
Dataset Summary
This dataset contains 8.9M commonsense assertions extracted by the Ascent pipeline developed at the Max Planck Institute for Informatics.
The focus of this dataset is on everyday concepts such as elephant, car, laptop, etc.
The current version of Ascent KB (v1.0.0) is approximately 19 times larger than ConceptNet (note that, in this comparison, non-commonsense knowledge in ConceptNet such as lexical relations is excluded).
For… See the full description on the dataset page: https://huggingface.co/datasets/tuanphong/ascent_kb.the-stack-dedup-python-filtered-non_asciiThis is a dataset originated from bigcode/the-stack-dedup with some filters applied.
The filters filtered in this dataset are:
remove_non_ascii
ascii-art
ASCII Art
Description
This is a text-to-image dataset, where the images are actually ASCII art.
The ASCII arts come from various sources:
asciiart: made by independent artists, listed on asciiart.eu
copypasta: common twitch emotes, listed on twitchquotes.com
graffiti: text samples styled using various ASCII art fonts with a tool
images: conversion of a portion of the dataset DataCompDR-12M using a tool
Metadata
homepage:… See the full description on the dataset page: https://huggingface.co/datasets/apehex/ascii-art.Ascend-COT-v2-json
AscendKernelGen/Ascend-COT-v2-json
AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.amazon_zhthis is a datasets about amazon reviews
TinyStories2-ascii
Dataset Card for "TinyStories2-ascii"
TinyStoriesV2-GPT4-{train,validation}.txt from roneneldan/TinyStories
ad-hoc Unicode -> ASCII normalization
remove empty/incomplete stories
ascii-art-datacompdr-12m
ASCII Art DataCompDR-12M
Description
This is a text-to-image dataset, where the images are actually ASCII art.
The images and captions were sampled from DataCompDR-12M.
The conversion was performed with the tool ascii-image-converter.
Metadata
homepage: https://github.com/apehex/scrapscii
version: 0.1.0
Config
Split
Size
Samples
'default'
'train'
4.1 GB
643072
'default'
'fixed'
552 MB
262144
The ASCII art in "fixed" all have a width of 64… See the full description on the dataset page: https://huggingface.co/datasets/apehex/ascii-art-datacompdr-12m.ascend_MIXED_cleaned_vadasciitermdraw-bench-public
ASCIITermDraw-Bench — Public Examples
12 public example tasks from ASCIITermDraw-Bench, a benchmark for
evaluating whether language models can generate and edit structured ASCII
diagrams.
The full benchmark has 80 private, held-out tasks used for actual scoring —
those are not distributed here. This dataset is a separate, hand-authored
set of 12 tasks (one easy, one medium, one hard per category) in the
exact same format, so anyone can see what a task looks like and run the… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/asciitermdraw-bench-public.ascl-code
ASCL Astronomy Source Code
The Astrophysics Source Code Library (ASCL) is a curated registry of
source code used in astronomy and astrophysics research. This dataset contains source files
extracted from ASCL-listed repositories, paired with catalog metadata.
Dataset Structure
Manifest (manifest.parquet)
One row per ASCL catalog entry with the following fields:
Field
Description
ascl_id
ASCL identifier (e.g., [ascl:2306.019])
title
Software title… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/ascl-code.ascend_ZH_cleaned_vadascend_EN_cleaned_vadtask1148_maximum_ascii_value
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1148_maximum_ascii_value
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1148_maximum_ascii_value.Ascend-CoT-v3-json
Ascend-CoT-v3-json
Ascend-CoT-v3-json is an Ascend C / CANN supervised fine-tuning dataset for custom operator development. It contains cleaned CoT-style samples for Ascend C kernel implementation, tiling logic, CANN API usage, debugging, and operator-development reasoning.
The release is organized into two final SFT subsets in one dataset repository.
Related Artifacts
Paper: AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-CoT-v3-json.amazon_zh_simpleTinyStories-ascii
TinyStories-{train,validation}.txt from roneneldan/TinyStories
ad-hoc Unicode -> ASCII normalization
remove empty/incomplete stories
3d-ascii-typeface-v1
3d_ascii_typeface_dataset Sampler
https://github.com/webxos for more info.
Generated by ASCII Dataset Studio v2.0 (by webXOS)
Contains 155 ASCII art samples. A full typeface dataset in 3D ASCII Format.
Dataset Structure
file_name: Path to the ASCII text file in the data/ folder.
text: The ASCII art content as a string.
Generation
Synthetically generated using multiple modes: image-to-ASCII, text-to-ASCII (fonts), geometric patterns, and video… See the full description on the dataset page: https://huggingface.co/datasets/webxos/3d-ascii-typeface-v1.ASCII-art-videoYou can ask an AI to write code to view the video as fps (frame by frame).
format:
▍ KHUNG 1/18
image_here
··················································
▍ KHUNG 2/18
image_here
··················································
▍ KHUNG 3/18
image_here
··················································
HUGGING FACE!!… See the full description on the dataset page: https://huggingface.co/datasets/mondk/ASCII-art-video.pl-web-graph-2026-09-14
Polish Web Domain Observations 2026-09-14
A curated snapshot of a .pl-focused domain crawler: DNS observations, host availability metadata, discovered URL references, and the crawl frontier. No HTML, page text, or website classifications are included. Observations accumulated over months, so September 14 dates the export itself while each row carries its own observation time. Coverage is whatever one crawler reached, and liveness holds as of the recorded timestamp.
Rendered… See the full description on the dataset page: https://huggingface.co/datasets/AsciiMAster/pl-web-graph-2026-09-14.ascp-context-attribution
ASCP: Causal Context Attribution and Probe Benchmark
Released artifacts for The Laws of Context Allocation: Causal Measurement and
Closed-Loop Orchestration in Generative Search.
📄 Paper: https://arxiv.org/abs/2608.23252
💻 Code: https://github.com/PeiYangLiu/ascp
Retrieval-augmented generation is usually measured with relevance proxies —
BM25, query–document cosine, output overlap — that score how related a passage
looks, not whether the generator used it. This dataset ships… See the full description on the dataset page: https://huggingface.co/datasets/PeiyangLiu/ascp-context-attribution.ASCII-art-images-for-trainSource: https://data.caltech.edu/records/mzrjq-6wc02
I converted these into ASCII art.
Link to download the original PNG files if you need them: https://data.caltech.edu/records/mzrjq-6wc02/files/caltech-101.zip?download=1
The txt files can be used to train a text-gen-only model that can "see" images in the form of ASCII art.
ty!
vi-gym-causal-ascii
Vi-Gym Causal ASCII Trajectories
This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment.
Warning
This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership.
Dataset Structure
Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.ASCIIEval
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
📖 Arxiv |
🤗 ASCIIEval Dataset |
🤗 ASCIITune Dataset
TABLE OF CONTENTS
Introduction
Data
Leaderboards
Leaderboard for Textual Input
Leaderboard for Image Input
Leaderboard for Average Cross-Modality Performance
Citation
Introduction
Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and… See the full description on the dataset page: https://huggingface.co/datasets/ASCIIEval/ASCIIEval.ASCEND_CLEAN
Dataset Card for Dataset Name
This dataset is derived from CAiRE/ASCEND. More information is available at https://huggingface.co/datasets/CAiRE/ASCEND.
Removed 嗯 呃 um uh
Resolved [UNK]'s using whisper-medium
Usage
Default utterances with cleaned transcripts
from datasets import load_dataset
data = load_dataset("georgechang8/ASCEND_CLEAN") # add split="train" for train set, etc.
Concatenated 30s utterances with cleaned transcripts… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/ASCEND_CLEAN.nanochat-ascend-dataset
nanochat-ascend-dataset
Unified training and evaluation data bundle for nanochat-ascend.
This repository is designed to make the nanochat-ascend training procedure easy to reproduce. Instead of asking users to collect multiple task and evaluation datasets and manually reconstruct the expected directory structure, this repository preserves the local filesystem layout expected by the training code.
The intended usage is simple:
place this repository at .cache/dataset
download… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-dataset.uke-mobile-network-permits-poland
UKE Mobile Network Permits - Poland
Radio permits issued by the Polish Office of Electronic Communications (Urząd
Komunikacji Elektronicznej, UKE) for cellular base stations across all bands:
GSM 900/1800, GSM-R, UMTS 900/2100, LTE 420/450/700/800/900/1800/2100/2600,
5G 700/900/1800/2100/2600/3600, and CDMA 420.
The dataset is a single GeoParquet file in WGS84 (EPSG:4326) containing
one row per permit per station per band.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/AsciiMAster/uke-mobile-network-permits-poland.datacomp_small_clip1_30pct_asciichr_greater_than_4
Dataset Card for "datacomp_small_clip1_30pct_asciichr_greater_than_4"
More Information needed
