CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Writer /omniact Dataset for OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web Splits: split_name count train 6788 test 2020 val 991 Example datapoint: "2849": { "task": "data/tasks/desktop/ibooks/task_1.30.txt", "image": "data/data/desktop/ibooks/screen_1.png", "box": "data/metadata/desktop/boxes/ibooks/screen_1.json" }, where: task - contains natural language description ("Task") along with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Writer/omniact.text-generation44 likes920 downloads2y agoHugging Face02m-a-p /COIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes. 🔔 Introduction COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-Writer.question-answering1K<n<10K27 likes426 downloads8mo agoHugging Face03open-llm-leaderboard-old /details_Writer__palmyra-med-20b Dataset Card for Evaluation run of Writer/palmyra-med-20b Dataset Summary Dataset automatically created during the evaluation run of model Writer/palmyra-med-20b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-med-20b.1 likes331 downloads3y agoHugging Face04Writer /FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329) Dataset count: 220 { "idx": int, "tokens": int, "context": string, "ocr_context": string, "answer": string, "query": string, "incomplete_query": string, "out-of-domain_query": string, "error_query": string, "out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.tabulartext-generationn<1K10 likes284 downloads2y agoHugging Face05fport /issue-writer-tr-en Issue Writer — bilingual (EN/TR) instruction dataset Turns raw product input — a Slack message, a support ticket, a Sentry alert, a meeting note — into well-formed issue tracker entries. Every assistant response is a single valid JSON object conforming to schema/issue.schema.json. Balanced across two languages: 50% English, 50% Turkish. Generator, validators, evaluation tooling and the fine-tuning notebook live in github.com/fport/issue-writer. Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.texttext-generation10K<n<100K1 likes234 downloads19d agoHugging Face06open-llm-leaderboard-old /details_Writer__palmyra-large Dataset Card for Evaluation run of Writer/palmyra-large Dataset Summary Dataset automatically created during the evaluation run of model Writer/palmyra-large on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-large.0 likes212 downloads3y agoHugging Face07open-llm-leaderboard-old /details_Writer__palmyra-20b-chat Dataset Card for Evaluation run of Writer/palmyra-20b-chat Dataset Summary Dataset automatically created during the evaluation run of model Writer/palmyra-20b-chat on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-20b-chat.0 likes127 downloads3y agoHugging Face0811-47 /Ghost-Writer-GOD1 likes92 downloads10mo agoHugging Face09Writer /IRT-mislabeled-items Potentially Mislabeled Items Detected by IRT Potential mislabeled benchmark items surfaced by the paper "Auditing LLM Benchmarks with Item Response Theory". Paper: https://arxiv.org/abs/2605.30504 Rows are included when either delta_li > 0 or the GPT-5.4 weak-reference label is mislabel or unsure. This is the union of items flagged by the unsupervised indicator and items flagged by the weak-reference labeler. For items flagged only by the weak-reference labeler but filtered out… See the full description on the dataset page: https://huggingface.co/datasets/Writer/IRT-mislabeled-items.tabular1K<n<10K0 likes80 downloads4mo agoHugging Face10maanas-writer /housing_qa_statutestext1M<n<10M0 likes61 downloads11mo agoHugging Face11ChaoticNeutrals /Thudm-Long_Writer-4.4k-ShareGPTOrginal Dataset from: https://huggingface.co/datasets/THUDM/LongWriter-6k Converted, deslopped, refusals removed, grammar corrected, min-hash deduplicated using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing @article{bai2024longwriter, title={LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs}, author={Yushi Bai and Jiajie Zhang and Xin Lv and Linzhi Zheng and Siqi Zhu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Thudm-Long_Writer-4.4k-ShareGPT.text1K<n<10K3 likes56 downloads2y agoHugging Face12Mildegard /olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy Gifts of Wandering Ice – Full Dataset from llms.txt This dataset contains all information from https://giftscomic.com/llms.txt, structured for AI assistants, RAG systems, recommendation engines, and discovery tools. The goal is to help AI systems accurately recommend and describe the free science-fiction webcomic Gifts of Wandering Ice and the broader creative work of its author, Olga Makarova (Mildegard). Dataset Contents (358 records) Split Records… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy.text-generationn<1K1 likes53 downloads3d agoHugging Face13bhxdianzhang /ParaSFT-writer ParaSFT Writer English | 中文 Overview ParaSFT Writer is a private supervised fine-tuning dataset for ParadoxGPT-Writer-4B, the ParadoxGPT specialist model for scientific writing and paper-argument reconstruction. Writer annotation pipeline over ParaPaper context packs, covering realization diagnosis, problem-insight extraction, intro structure, commitment alignment, method necessity, and experiment closure tasks. Each example is an instruction-tuning record with a… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-writer.texttext-generation10K<n<100K0 likes50 downloads3mo agoHugging Face14adiboy04 /tier2_writer0 likes46 downloads2mo agoHugging Face15open-llm-leaderboard /lars1234__Mistral-Small-24B-Instruct-2501-writer-detailsgated Dataset Card for Evaluation run of lars1234/Mistral-Small-24B-Instruct-2501-writer Dataset automatically created during the evaluation run of model lars1234/Mistral-Small-24B-Instruct-2501-writer The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lars1234__Mistral-Small-24B-Instruct-2501-writer-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face16maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-docfinqa-train-c8192-t4096-1000s-agnostictabular1K<n<10K0 likes32 downloads10mo agoHugging Face17open-llm-leaderboard-old /details_Writer__palmyra-base Dataset Card for Evaluation run of Writer/palmyra-base Dataset Summary Dataset automatically created during the evaluation run of model Writer/palmyra-base on the Open LLM Leaderboard. The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-base.0 likes30 downloads3y agoHugging Face18Adarsh203 /IAM_line_with_writerimage10K<n<100K0 likes30 downloads2y agoHugging Face19narrativescreativelabs /edge-telugu-writers-room The Edge Telugu Writers Room A self-evolving, local-first filmmaking + Telugu-LLM studio. This repo is the software spine of the architecture: Layers 1, 2, and 4 are fully built here. Layer 0 (your Mac's models) and Layer 3 (cloud neural scoring) are wired as interfaces you point at when ready. Working principles (baked into the code) No tool is an oracle. Every persona score is logged as a prediction, and calibrate.py weights personas by how well they predicted… See the full description on the dataset page: https://huggingface.co/datasets/narrativescreativelabs/edge-telugu-writers-room.0 likes30 downloads10d agoHugging Face20niklashcs /test_single_episode_writervideon<1K0 likes29 downloads10mo agoHugging Face21open-llm-leaderboard-old /details_Writer__InstructPalmyra-20b Dataset Card for Evaluation run of Writer/InstructPalmyra-20b Dataset Summary Dataset automatically created during the evaluation run of model Writer/InstructPalmyra-20b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__InstructPalmyra-20b.0 likes28 downloads3y agoHugging Face22houyuhuifei /ekg-writer EKG-Writer Structured long-form story generation with a five-layer Event Knowledge Graph (EKG). This repository is the research code and materials accompanying the paper "Controllable story generation with event knowledge graphs: an exploratory case study of human-AI long-form narrative writing" (submitted to Natural Language Processing, Cambridge University Press). What is EKG-Writer? EKG-Writer is a research prototype for structured AI story generation. It uses… See the full description on the dataset page: https://huggingface.co/datasets/houyuhuifei/ekg-writer.0 likes28 downloads8d agoHugging Face23thanhdath /grpo_sql_writer_bird_train_reference_sqls_addedtext1K<n<10K0 likes26 downloads8mo agoHugging Face24shelly-writer /triviaqa-unmemorizedtext10K<n<100K0 likes25 downloads1y agoHugging Face25LizstMaster8888 /Lost-TXT-For-Doyle-Writer0 likes25 downloads18d agoHugging Face26open-llm-leaderboard-old /details_Writer__camel-5b-hf Dataset Card for Evaluation run of Writer/camel-5b-hf Dataset Summary Dataset automatically created during the evaluation run of model Writer/camel-5b-hf on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__camel-5b-hf.0 likes24 downloads3y agoHugging Face27jaisaravana /writer1.0100K<n<1M0 likes24 downloads2y agoHugging Face28maanas-writer /mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c27000-t1024-10s-agnostictabularn<1K0 likes24 downloads11mo agoHugging Face29JunoLi622 /COIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes. 🔔 Introduction COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/JunoLi622/COIG-Writer.textquestion-answering1K<n<10K0 likes23 downloads11mo agoHugging Face30maanas-writer /barexam_qatext1K<n<10K0 likes23 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.