CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Writer /FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329) Dataset count: 220 { "idx": int, "tokens": int, "context": string, "ocr_context": string, "answer": string, "query": string, "incomplete_query": string, "out-of-domain_query": string, "error_query": string, "out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.tabulartext-generationn<1K10 likes284 downloads2y agoHugging Face02fport /issue-writer-tr-en Issue Writer — bilingual (EN/TR) instruction dataset Turns raw product input — a Slack message, a support ticket, a Sentry alert, a meeting note — into well-formed issue tracker entries. Every assistant response is a single valid JSON object conforming to schema/issue.schema.json. Balanced across two languages: 50% English, 50% Turkish. Generator, validators, evaluation tooling and the fine-tuning notebook live in github.com/fport/issue-writer. Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.texttext-generation10K<n<100K1 likes234 downloads19d agoHugging Face03Writer /IRT-mislabeled-items Potentially Mislabeled Items Detected by IRT Potential mislabeled benchmark items surfaced by the paper "Auditing LLM Benchmarks with Item Response Theory". Paper: https://arxiv.org/abs/2605.30504 Rows are included when either delta_li > 0 or the GPT-5.4 weak-reference label is mislabel or unsure. This is the union of items flagged by the unsupervised indicator and items flagged by the weak-reference labeler. For items flagged only by the weak-reference labeler but filtered out… See the full description on the dataset page: https://huggingface.co/datasets/Writer/IRT-mislabeled-items.tabular1K<n<10K0 likes80 downloads4mo agoHugging Face04maanas-writer /housing_qa_statutestext1M<n<10M0 likes61 downloads11mo agoHugging Face05ChaoticNeutrals /Thudm-Long_Writer-4.4k-ShareGPTOrginal Dataset from: https://huggingface.co/datasets/THUDM/LongWriter-6k Converted, deslopped, refusals removed, grammar corrected, min-hash deduplicated using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing @article{bai2024longwriter, title={LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs}, author={Yushi Bai and Jiajie Zhang and Xin Lv and Linzhi Zheng and Siqi Zhu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Thudm-Long_Writer-4.4k-ShareGPT.text1K<n<10K3 likes56 downloads2y agoHugging Face06bhxdianzhang /ParaSFT-writer ParaSFT Writer English | 中文 Overview ParaSFT Writer is a private supervised fine-tuning dataset for ParadoxGPT-Writer-4B, the ParadoxGPT specialist model for scientific writing and paper-argument reconstruction. Writer annotation pipeline over ParaPaper context packs, covering realization diagnosis, problem-insight extraction, intro structure, commitment alignment, method necessity, and experiment closure tasks. Each example is an instruction-tuning record with a… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-writer.texttext-generation10K<n<100K0 likes50 downloads3mo agoHugging Face07open-llm-leaderboard /lars1234__Mistral-Small-24B-Instruct-2501-writer-detailsgated Dataset Card for Evaluation run of lars1234/Mistral-Small-24B-Instruct-2501-writer Dataset automatically created during the evaluation run of model lars1234/Mistral-Small-24B-Instruct-2501-writer The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lars1234__Mistral-Small-24B-Instruct-2501-writer-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face08maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-docfinqa-train-c8192-t4096-1000s-agnostictabular1K<n<10K0 likes32 downloads10mo agoHugging Face09Adarsh203 /IAM_line_with_writerimage10K<n<100K0 likes30 downloads2y agoHugging Face10thanhdath /grpo_sql_writer_bird_train_reference_sqls_addedtext1K<n<10K0 likes26 downloads8mo agoHugging Face11shelly-writer /triviaqa-unmemorizedtext10K<n<100K0 likes25 downloads1y agoHugging Face12maanas-writer /mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c27000-t1024-10s-agnostictabularn<1K0 likes24 downloads11mo agoHugging Face13JunoLi622 /COIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes. 🔔 Introduction COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/JunoLi622/COIG-Writer.textquestion-answering1K<n<10K0 likes23 downloads11mo agoHugging Face14maanas-writer /barexam_qatext1K<n<10K0 likes23 downloads11mo agoHugging Face15talentlabs /training-data-blog-writer_v05-09-2023 Dataset Card for "training-data-blog-writer_v05-09-2023" More Information needed text10K<n<100K0 likes21 downloads3y agoHugging Face16Writer /writing-in-the-margins-multihopragtext1K<n<10K0 likes21 downloads2y agoHugging Face17Delta-Vector /Orion-Co-Writer-51Ktext10K<n<100K3 likes20 downloads1y agoHugging Face18maanas-writer /memagent_hotpotqa_train_32ktext10K<n<100K0 likes19 downloads11mo agoHugging Face19talentlabs /training-data-blog-writer_v03-09-2023 Dataset Card for "training-data-blog-writer_v03-09-2023" More Information needed text1K<n<10K0 likes18 downloads3y agoHugging Face20maanas-writer /mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c2048-t1024-10s-agnostic-nocontexttabularn<1K0 likes18 downloads11mo agoHugging Face21maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-triviaqa-val-c27000-t2048-1000s-agnostictabular1K<n<10K0 likes18 downloads10mo agoHugging Face22maanas-writer /mem_agent-model_based-llama-3-3-70b-i-ruler-qa-test-c27000-t1024-10s-agnostictabularn<1K0 likes16 downloads11mo agoHugging Face23SalimMS /Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced Dataset Card for Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/SalimMS/Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/SalimMS/Synthetic-SlackDay-Writer-v1-SlackMessages-Enhanced.text1K<n<10K0 likes15 downloads1y agoHugging Face24maanas-writer /bertscore-llama-3-3-70b-i-triviaqa-llama-memorization-val-c4096-t2048-1000s-agnostictabular1K<n<10K0 likes15 downloads11mo agoHugging Face25maanas-writer /mem_agent-bertscore-rl-memoryagent-14b-docfinqa-train-c4096-t4096-1000s-agnostictabular1K<n<10K0 likes15 downloads11mo agoHugging Face26maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-infbench-longbook-qa-test-c8192-t4096-1000s-agnostitabularn<1K0 likes15 downloads10mo agoHugging Face27takahashi111 /identify-writers-country 目的 医学論文のabstractから、論文を書いた著者の国を推定する。 母語によって書く英語に特徴が出るのではないかと思った。 理想 著者の特徴を出すために一定以上の長さをもつabstractにする。 labelの偏りをなくす。 機械翻訳や生成AIの影響をなくすため、昔の論文にする。 国 日本、中国、ロシア、サウジアラビア 言語の系統と構造から離れているものを選んだ。 論文の検索条件 日本:1956~2010年で東京大学と京都大学に所属する研究者が出したもの。 中国:1945~2010年で北京大学、清華大学、復旦大学、上海交通大学 ロシア:1973~2021でモスクワ大学、サンクトペテルブルグ大学、ノボシビルスク大学 サウジアラビア:1981~2019でキングサウード大学、キングアブドゥルアズィーズ大学、キングファイサル大学 データセット作成条件 abstractが140words以上… See the full description on the dataset page: https://huggingface.co/datasets/takahashi111/identify-writers-country.texttext-classification10K<n<100K0 likes14 downloads1y agoHugging Face28maanas-writer /mem_agent-model_based-rl-memoryagent-14b-infbench-longbook-qa-test-c31000-t4096-1000s-agnostictabularn<1K0 likes14 downloads11mo agoHugging Face29maanas-writer /mem_agent-model_based-qwen3-1-5b-oldgrpo-2086-infbench-longbook-choice-test-c27000-t4096-10s-agntabularn<1K0 likes14 downloads10mo agoHugging Face30alst10 /alston-writer-cpttextn<1K0 likes14 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.