CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Writer /omniact Dataset for OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web Splits: split_name count train 6788 test 2020 val 991 Example datapoint: "2849": { "task": "data/tasks/desktop/ibooks/task_1.30.txt", "image": "data/data/desktop/ibooks/screen_1.png", "box": "data/metadata/desktop/boxes/ibooks/screen_1.json" }, where: task - contains natural language description ("Task") along with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Writer/omniact.text-generation44 likes920 downloads2y agoHugging Face02Writer /FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329) Dataset count: 220 { "idx": int, "tokens": int, "context": string, "ocr_context": string, "answer": string, "query": string, "incomplete_query": string, "out-of-domain_query": string, "error_query": string, "out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.tabulartext-generationn<1K10 likes284 downloads2y agoHugging Face03fport /issue-writer-tr-en Issue Writer — bilingual (EN/TR) instruction dataset Turns raw product input — a Slack message, a support ticket, a Sentry alert, a meeting note — into well-formed issue tracker entries. Every assistant response is a single valid JSON object conforming to schema/issue.schema.json. Balanced across two languages: 50% English, 50% Turkish. Generator, validators, evaluation tooling and the fine-tuning notebook live in github.com/fport/issue-writer. Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.texttext-generation10K<n<100K1 likes234 downloads20d agoHugging Face04Mildegard /olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy Gifts of Wandering Ice – Full Dataset from llms.txt This dataset contains all information from https://giftscomic.com/llms.txt, structured for AI assistants, RAG systems, recommendation engines, and discovery tools. The goal is to help AI systems accurately recommend and describe the free science-fiction webcomic Gifts of Wandering Ice and the broader creative work of its author, Olga Makarova (Mildegard). Dataset Contents (358 records) Split Records… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy.text-generationn<1K1 likes53 downloads4d agoHugging Face05bhxdianzhang /ParaSFT-writer ParaSFT Writer English | 中文 Overview ParaSFT Writer is a private supervised fine-tuning dataset for ParadoxGPT-Writer-4B, the ParadoxGPT specialist model for scientific writing and paper-argument reconstruction. Writer annotation pipeline over ParaPaper context packs, covering realization diagnosis, problem-insight extraction, intro structure, commitment alignment, method necessity, and experiment closure tasks. Each example is an instruction-tuning record with a… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-writer.texttext-generation10K<n<100K0 likes50 downloads3mo agoHugging Face06narinzar /streaming-tokenizer-shard-writer streaming-tokenizer-shard-writer (sample shards) Sample output from the streaming-tokenizer-shard-writer pipeline: a streaming, parallel tokenizer that packs a text corpus into fixed-size, size-balanced training shards using bounded memory. What is here shard-000000.tar, shard-000001.tar — two sample shards. Each is a dependency-free tar archive (webdataset-style); every member is one tokenized sample named <shard>-<seq>.ids, whose payload is the token ids packed… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/streaming-tokenizer-shard-writer.text-generation1M<n<10M0 likes7 downloads3mo agoHugging Face07Writerslogic /scholawrite-augmentedgated ScholaWrite-Augmented Process Integrity Benchmarks for Revision-Tracked Scholarly Writing Dataset Description ScholaWrite-Augmented is a revision-tracked scholarly writing dataset with annotated external insertion events, designed for process-integrity research. It augments the ScholaWrite seed dataset with synthetic injections at multiple sophistication levels, models boundary erosion over revision trajectories, and provides span-level annotations with explicit ambiguity… See the full description on the dataset page: https://huggingface.co/datasets/Writerslogic/scholawrite-augmented.text-classification100K<n<1M0 likes3 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.