datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329)
Dataset count: 220
{
"idx": int,
"tokens": int,
"context": string,
"ocr_context": string,
"answer": string,
"query": string,
"incomplete_query": string,
"out-of-domain_query": string,
"error_query": string,
"out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.issue-writer-tr-en
Issue Writer — bilingual (EN/TR) instruction dataset
Turns raw product input — a Slack message, a support ticket, a Sentry alert, a
meeting note — into well-formed issue tracker entries. Every assistant response is a
single valid JSON object conforming to schema/issue.schema.json.
Balanced across two languages: 50% English, 50% Turkish.
Generator, validators, evaluation tooling and the fine-tuning notebook live in
github.com/fport/issue-writer.
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.IRT-mislabeled-items
Potentially Mislabeled Items Detected by IRT
Potential mislabeled benchmark items surfaced by the paper "Auditing LLM Benchmarks with Item Response Theory".
Paper: https://arxiv.org/abs/2605.30504
Rows are included when either delta_li > 0 or the GPT-5.4 weak-reference label is mislabel or unsure.
This is the union of items flagged by the unsupervised indicator and items flagged by the weak-reference labeler.
For items flagged only by the weak-reference labeler but filtered out… See the full description on the dataset page: https://huggingface.co/datasets/Writer/IRT-mislabeled-items.housing_qa_statutesThudm-Long_Writer-4.4k-ShareGPTOrginal Dataset from: https://huggingface.co/datasets/THUDM/LongWriter-6k
Converted, deslopped, refusals removed, grammar corrected, min-hash deduplicated using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
@article{bai2024longwriter,
title={LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs},
author={Yushi Bai and Jiajie Zhang and Xin Lv and Linzhi Zheng and Siqi Zhu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Thudm-Long_Writer-4.4k-ShareGPT.ParaSFT-writer
ParaSFT Writer
English | 中文
Overview
ParaSFT Writer is a private supervised fine-tuning dataset for ParadoxGPT-Writer-4B, the ParadoxGPT specialist model for scientific writing and paper-argument reconstruction.
Writer annotation pipeline over ParaPaper context packs, covering realization diagnosis, problem-insight extraction, intro structure, commitment alignment, method necessity, and experiment closure tasks.
Each example is an instruction-tuning record with a… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-writer.lars1234__Mistral-Small-24B-Instruct-2501-writer-details
Dataset Card for Evaluation run of lars1234/Mistral-Small-24B-Instruct-2501-writer
Dataset automatically created during the evaluation run of model lars1234/Mistral-Small-24B-Instruct-2501-writer
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lars1234__Mistral-Small-24B-Instruct-2501-writer-details.mem_agent-model_based-memagent-1-5b-step1024-docfinqa-train-c8192-t4096-1000s-agnosticIAM_line_with_writergrpo_sql_writer_bird_train_reference_sqls_addedtriviaqa-unmemorizedmem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c27000-t1024-10s-agnosticCOIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes.
🔔 Introduction
COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/JunoLi622/COIG-Writer.barexam_qatraining-data-blog-writer_v05-09-2023
Dataset Card for "training-data-blog-writer_v05-09-2023"
More Information needed
writing-in-the-margins-multihopragOrion-Co-Writer-51Kmemagent_hotpotqa_train_32ktraining-data-blog-writer_v03-09-2023
Dataset Card for "training-data-blog-writer_v03-09-2023"
More Information needed
mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c2048-t1024-10s-agnostic-nocontextmem_agent-model_based-memagent-1-5b-step1024-triviaqa-val-c27000-t2048-1000s-agnosticmem_agent-model_based-llama-3-3-70b-i-ruler-qa-test-c27000-t1024-10s-agnosticSynthetic-SlackDay-Writer-v1-SlackMessages-Enhanced
Dataset Card for Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/SalimMS/Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/SalimMS/Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced.bertscore-llama-3-3-70b-i-triviaqa-llama-memorization-val-c4096-t2048-1000s-agnosticmem_agent-bertscore-rl-memoryagent-14b-docfinqa-train-c4096-t4096-1000s-agnosticmem_agent-model_based-memagent-1-5b-step1024-infbench-longbook-qa-test-c8192-t4096-1000s-agnostiidentify-writers-country
目的
医学論文のabstractから、論文を書いた著者の国を推定する。
母語によって書く英語に特徴が出るのではないかと思った。
理想
著者の特徴を出すために一定以上の長さをもつabstractにする。
labelの偏りをなくす。
機械翻訳や生成AIの影響をなくすため、昔の論文にする。
国
日本、中国、ロシア、サウジアラビア
言語の系統と構造から離れているものを選んだ。
論文の検索条件
日本:1956~2010年で東京大学と京都大学に所属する研究者が出したもの。
中国:1945~2010年で北京大学、清華大学、復旦大学、上海交通大学
ロシア:1973~2021でモスクワ大学、サンクトペテルブルグ大学、ノボシビルスク大学
サウジアラビア:1981~2019でキングサウード大学、キングアブドゥルアズィーズ大学、キングファイサル大学
データセット作成条件
abstractが140words以上… See the full description on the dataset page: https://huggingface.co/datasets/takahashi111/identify-writers-country.mem_agent-model_based-rl-memoryagent-14b-infbench-longbook-qa-test-c31000-t4096-1000s-agnosticmem_agent-model_based-qwen3-1-5b-oldgrpo-2086-infbench-longbook-choice-test-c27000-t4096-10s-agnalston-writer-cpt
