datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omniact
Dataset for OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Splits:
split_name
count
train
6788
test
2020
val
991
Example datapoint:
"2849": {
"task": "data/tasks/desktop/ibooks/task_1.30.txt",
"image": "data/data/desktop/ibooks/screen_1.png",
"box": "data/metadata/desktop/boxes/ibooks/screen_1.json"
},
where:
task - contains natural language description ("Task") along with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Writer/omniact.FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329)
Dataset count: 220
{
"idx": int,
"tokens": int,
"context": string,
"ocr_context": string,
"answer": string,
"query": string,
"incomplete_query": string,
"out-of-domain_query": string,
"error_query": string,
"out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.issue-writer-tr-en
Issue Writer — bilingual (EN/TR) instruction dataset
Turns raw product input — a Slack message, a support ticket, a Sentry alert, a
meeting note — into well-formed issue tracker entries. Every assistant response is a
single valid JSON object conforming to schema/issue.schema.json.
Balanced across two languages: 50% English, 50% Turkish.
Generator, validators, evaluation tooling and the fine-tuning notebook live in
github.com/fport/issue-writer.
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy
Gifts of Wandering Ice – Full Dataset from llms.txt
This dataset contains all information from https://giftscomic.com/llms.txt, structured for AI assistants, RAG systems, recommendation engines, and discovery tools.
The goal is to help AI systems accurately recommend and describe the free science-fiction webcomic Gifts of Wandering Ice and the broader creative work of its author, Olga Makarova (Mildegard).
Dataset Contents (358 records)
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy.ParaSFT-writer
ParaSFT Writer
English | 中文
Overview
ParaSFT Writer is a private supervised fine-tuning dataset for ParadoxGPT-Writer-4B, the ParadoxGPT specialist model for scientific writing and paper-argument reconstruction.
Writer annotation pipeline over ParaPaper context packs, covering realization diagnosis, problem-insight extraction, intro structure, commitment alignment, method necessity, and experiment closure tasks.
Each example is an instruction-tuning record with a… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-writer.streaming-tokenizer-shard-writer
streaming-tokenizer-shard-writer (sample shards)
Sample output from the streaming-tokenizer-shard-writer
pipeline: a streaming, parallel tokenizer that packs a text corpus into
fixed-size, size-balanced training shards using bounded memory.
What is here
shard-000000.tar, shard-000001.tar — two sample shards. Each is a
dependency-free tar archive (webdataset-style); every member is one tokenized
sample named <shard>-<seq>.ids, whose payload is the token ids packed… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/streaming-tokenizer-shard-writer.scholawrite-augmented
ScholaWrite-Augmented
Process Integrity Benchmarks for Revision-Tracked Scholarly Writing
Dataset Description
ScholaWrite-Augmented is a revision-tracked scholarly writing dataset with annotated external insertion events, designed for process-integrity research. It augments the ScholaWrite seed dataset with synthetic injections at multiple sophistication levels, models boundary erosion over revision trajectories, and provides span-level annotations with explicit ambiguity… See the full description on the dataset page: https://huggingface.co/datasets/Writerslogic/scholawrite-augmented.
