datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.Ascend-COT-v2-json
AscendKernelGen/Ascend-COT-v2-json
AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.ascl-code
ASCL Astronomy Source Code
The Astrophysics Source Code Library (ASCL) is a curated registry of
source code used in astronomy and astrophysics research. This dataset contains source files
extracted from ASCL-listed repositories, paired with catalog metadata.
Dataset Structure
Manifest (manifest.parquet)
One row per ASCL catalog entry with the following fields:
Field
Description
ascl_id
ASCL identifier (e.g., [ascl:2306.019])
title
Software title… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/ascl-code.asciitermdraw-bench-public
ASCIITermDraw-Bench — Public Examples
12 public example tasks from ASCIITermDraw-Bench, a benchmark for
evaluating whether language models can generate and edit structured ASCII
diagrams.
The full benchmark has 80 private, held-out tasks used for actual scoring —
those are not distributed here. This dataset is a separate, hand-authored
set of 12 tasks (one easy, one medium, one hard per category) in the
exact same format, so anyone can see what a task looks like and run the… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/asciitermdraw-bench-public.task1148_maximum_ascii_value
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1148_maximum_ascii_value
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1148_maximum_ascii_value.Ascend-CoT-v3-json
Ascend-CoT-v3-json
Ascend-CoT-v3-json is an Ascend C / CANN supervised fine-tuning dataset for custom operator development. It contains cleaned CoT-style samples for Ascend C kernel implementation, tiling logic, CANN API usage, debugging, and operator-development reasoning.
The release is organized into two final SFT subsets in one dataset repository.
Related Artifacts
Paper: AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-CoT-v3-json.vi-gym-causal-ascii
Vi-Gym Causal ASCII Trajectories
This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment.
Warning
This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership.
Dataset Structure
Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.ASCIIEval
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
📖 Arxiv |
🤗 ASCIIEval Dataset |
🤗 ASCIITune Dataset
TABLE OF CONTENTS
Introduction
Data
Leaderboards
Leaderboard for Textual Input
Leaderboard for Image Input
Leaderboard for Average Cross-Modality Performance
Citation
Introduction
Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and… See the full description on the dataset page: https://huggingface.co/datasets/ASCIIEval/ASCIIEval.nanochat-ascend-dataset
nanochat-ascend-dataset
Unified training and evaluation data bundle for nanochat-ascend.
This repository is designed to make the nanochat-ascend training procedure easy to reproduce. Instead of asking users to collect multiple task and evaluation datasets and manually reconstruct the expected directory structure, this repository preserves the local filesystem layout expected by the training code.
The intended usage is simple:
place this repository at .cache/dataset
download… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-dataset.THE.ASCII.ART.EMPORIUMhttps://asciiartist.com/sitefiles/respectartistscampaign.html
(published by Laura Brown aka ldb)
Respect ASCII Artists Campaign
Most ASCII artists will tag their ASCII creation with their initials.
This is not just about signing your art, it shows the original artist.
If someone else colours the art, or modifies it in any other way, the
artist initials need to be kept with it. Anyone modifying art can add
their initials (usually something like ldb/ you) and a note about what
they had… See the full description on the dataset page: https://huggingface.co/datasets/Csplk/THE.ASCII.ART.EMPORIUM.qwen38-27b-triton-ascend-rl-trajectories
Qwen3.8-27B Triton-Ascend RL Trajectories
This dataset contains 1,000 multi-turn trajectories for Triton-Ascend kernel generation. Every included trajectory passed compilation and correctness validation on one official npu-kernelbench workload. Qwen3.8-27B generated an initial solution and received evaluator feedback for up to five calls.
Dataset Viewer subsets
trajectories (default): one row per sample with only messages. The initial system and task user… See the full description on the dataset page: https://huggingface.co/datasets/Yukki1011/qwen38-27b-triton-ascend-rl-trajectories.Asclepius-Synthetic-Clinical-NotesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/Asclepius-Synthetic-Clinical-Notes.asc-finance-reasoning-20k
Finance Reasoning SFT Dataset — AutoScientist Challenge (20K)
A ~20,018-row chain-of-thought finance-reasoning dataset for supervised fine-tuning — the exact
training set behind the companion model (Mixtral-8x7B-Instruct LoRA, lead entry E24). Built as the
data-recipe submission to the AutoScientist Challenge
(Finance, Part 1, Adaption Labs, 2026) and released under CC-BY-4.0.
It is a 4,018-row curated seed (Adaptive Data quality 9.3/10, grade A) expanded to ~20K with… See the full description on the dataset page: https://huggingface.co/datasets/tejadhith/asc-finance-reasoning-20k.ASCII-Art
🎨 Pink Pixel ASCII Art Dataset ✨
Welcome to the official Pink Pixel ASCII Art Dataset! 💖 This dataset is a curated collection of 1,221 conversational pairs designed to teach AI models how to generate beautiful, creative, and terminal-friendly ASCII art.
🚀 Dataset Summary
This dataset follows the ChatML (OpenAI) format, featuring a system message, a user request, and an assistant's ASCII art response. It covers a wide range of subjects:
🐱 Animals: Cats, dogs, bunnies… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/ASCII-Art.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.distiset-ascii-art-a1
Dataset Card for distiset-ascii-art-a1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/DominguesAddem1974/distiset-ascii-art-a1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DominguesAddem1974/distiset-ascii-art-a1.nanochat-ascend-task
Introduction
This dataset contains the task dataset of nanochat-asecnd, excluding identity_conversations.jsonl and words_alpha.txt. It includes the following subsets:
ai2_arc
gsm8k
humaneval
mmlu
smol-smoltalk
Layout Tree
├── ai2_arc
│ ├── ARC-Challenge
│ │ ├── test-00000-of-00001.parquet
│ │ ├── train-00000-of-00001.parquet
│ │ └── validation-00000-of-00001.parquet
│ ├── ARC-Easy
│ │ ├── test-00000-of-00001.parquet
│ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-task.
