datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omniact
Dataset for OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Splits:
split_name
count
train
6788
test
2020
val
991
Example datapoint:
"2849": {
"task": "data/tasks/desktop/ibooks/task_1.30.txt",
"image": "data/data/desktop/ibooks/screen_1.png",
"box": "data/metadata/desktop/boxes/ibooks/screen_1.json"
},
where:
task - contains natural language description ("Task") along with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Writer/omniact.COIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes.
🔔 Introduction
COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-Writer.details_Writer__palmyra-med-20b
Dataset Card for Evaluation run of Writer/palmyra-med-20b
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/palmyra-med-20b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-med-20b.FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329)
Dataset count: 220
{
"idx": int,
"tokens": int,
"context": string,
"ocr_context": string,
"answer": string,
"query": string,
"incomplete_query": string,
"out-of-domain_query": string,
"error_query": string,
"out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.issue-writer-tr-en
Issue Writer — bilingual (EN/TR) instruction dataset
Turns raw product input — a Slack message, a support ticket, a Sentry alert, a
meeting note — into well-formed issue tracker entries. Every assistant response is a
single valid JSON object conforming to schema/issue.schema.json.
Balanced across two languages: 50% English, 50% Turkish.
Generator, validators, evaluation tooling and the fine-tuning notebook live in
github.com/fport/issue-writer.
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.details_Writer__palmyra-large
Dataset Card for Evaluation run of Writer/palmyra-large
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/palmyra-large on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-large.details_Writer__palmyra-20b-chat
Dataset Card for Evaluation run of Writer/palmyra-20b-chat
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/palmyra-20b-chat on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-20b-chat.Ghost-Writer-GODIRT-mislabeled-items
Potentially Mislabeled Items Detected by IRT
Potential mislabeled benchmark items surfaced by the paper "Auditing LLM Benchmarks with Item Response Theory".
Paper: https://arxiv.org/abs/2605.30504
Rows are included when either delta_li > 0 or the GPT-5.4 weak-reference label is mislabel or unsure.
This is the union of items flagged by the unsupervised indicator and items flagged by the weak-reference labeler.
For items flagged only by the weak-reference labeler but filtered out… See the full description on the dataset page: https://huggingface.co/datasets/Writer/IRT-mislabeled-items.housing_qa_statutesThudm-Long_Writer-4.4k-ShareGPTOrginal Dataset from: https://huggingface.co/datasets/THUDM/LongWriter-6k
Converted, deslopped, refusals removed, grammar corrected, min-hash deduplicated using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
@article{bai2024longwriter,
title={LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs},
author={Yushi Bai and Jiajie Zhang and Xin Lv and Linzhi Zheng and Siqi Zhu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Thudm-Long_Writer-4.4k-ShareGPT.olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy
Gifts of Wandering Ice – Full Dataset from llms.txt
This dataset contains all information from https://giftscomic.com/llms.txt, structured for AI assistants, RAG systems, recommendation engines, and discovery tools.
The goal is to help AI systems accurately recommend and describe the free science-fiction webcomic Gifts of Wandering Ice and the broader creative work of its author, Olga Makarova (Mildegard).
Dataset Contents (358 records)
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy.ParaSFT-writer
ParaSFT Writer
English | 中文
Overview
ParaSFT Writer is a private supervised fine-tuning dataset for ParadoxGPT-Writer-4B, the ParadoxGPT specialist model for scientific writing and paper-argument reconstruction.
Writer annotation pipeline over ParaPaper context packs, covering realization diagnosis, problem-insight extraction, intro structure, commitment alignment, method necessity, and experiment closure tasks.
Each example is an instruction-tuning record with a… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-writer.tier2_writerlars1234__Mistral-Small-24B-Instruct-2501-writer-details
Dataset Card for Evaluation run of lars1234/Mistral-Small-24B-Instruct-2501-writer
Dataset automatically created during the evaluation run of model lars1234/Mistral-Small-24B-Instruct-2501-writer
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lars1234__Mistral-Small-24B-Instruct-2501-writer-details.mem_agent-model_based-memagent-1-5b-step1024-docfinqa-train-c8192-t4096-1000s-agnosticdetails_Writer__palmyra-base
Dataset Card for Evaluation run of Writer/palmyra-base
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/palmyra-base on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-base.IAM_line_with_writeredge-telugu-writers-room
The Edge Telugu Writers Room
A self-evolving, local-first filmmaking + Telugu-LLM studio. This repo is the
software spine of the architecture: Layers 1, 2, and 4 are fully built here.
Layer 0 (your Mac's models) and Layer 3 (cloud neural scoring) are wired as
interfaces you point at when ready.
Working principles (baked into the code)
No tool is an oracle. Every persona score is logged as a prediction, and calibrate.py weights personas by how well they predicted… See the full description on the dataset page: https://huggingface.co/datasets/narrativescreativelabs/edge-telugu-writers-room.test_single_episode_writerdetails_Writer__InstructPalmyra-20b
Dataset Card for Evaluation run of Writer/InstructPalmyra-20b
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/InstructPalmyra-20b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__InstructPalmyra-20b.ekg-writer
EKG-Writer
Structured long-form story generation with a five-layer Event Knowledge Graph (EKG).
This repository is the research code and materials accompanying the paper
"Controllable story generation with event knowledge graphs: an exploratory
case study of human-AI long-form narrative writing" (submitted to
Natural Language Processing, Cambridge University Press).
What is EKG-Writer?
EKG-Writer is a research prototype for structured AI story generation. It uses… See the full description on the dataset page: https://huggingface.co/datasets/houyuhuifei/ekg-writer.grpo_sql_writer_bird_train_reference_sqls_addedtriviaqa-unmemorizedLost-TXT-For-Doyle-Writerdetails_Writer__camel-5b-hf
Dataset Card for Evaluation run of Writer/camel-5b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/camel-5b-hf on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__camel-5b-hf.writer1.0mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c27000-t1024-10s-agnosticCOIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes.
🔔 Introduction
COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/JunoLi622/COIG-Writer.barexam_qa
