CoolFace
21 results

writer

Writer /omniact Dataset for OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web Splits: split_name count train 6788 test 2020 val 991 Example datapoint: "2849": { "task": "data/tasks/desktop/ibooks/task_1.30.txt", "image": "data/data/desktop/ibooks/screen_1.png", "box": "data/metadata/desktop/boxes/ibooks/screen_1.json" }, where: task - contains natural language description ("Task") along with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Writer/omniact.text-generation44 likes920 downloads2y agoHugging Facem-a-p /COIG-WriterThis repository contains the dataset and supplementary materials for the paper COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes. 🔔 Introduction COIG-Writer is a large-scale Chinese creative writing dataset that connects final literary works with their underlying reasoning processes.Each sample includes a reverse-engineered writing prompt, a step-by-step reasoning trace, and the final article.This design allows researchers to explore… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-Writer.question-answering1K<n<10K27 likes426 downloads8mo agoHugging Faceopen-llm-leaderboard-old /details_Writer__palmyra-med-20b Dataset Card for Evaluation run of Writer/palmyra-med-20b Dataset Summary Dataset automatically created during the evaluation run of model Writer/palmyra-med-20b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-med-20b.1 likes331 downloads3y agoHugging FaceWriter /FailSafeQABenchmark data introduced in the paper: Expect the Unexpected: FailSafeQA Long Context for Finance (https://arxiv.org/abs/2502.06329) Dataset count: 220 { "idx": int, "tokens": int, "context": string, "ocr_context": string, "answer": string, "query": string, "incomplete_query": string, "out-of-domain_query": string, "error_query": string, "out-of-scope_query":… See the full description on the dataset page: https://huggingface.co/datasets/Writer/FailSafeQA.tabulartext-generationn<1K10 likes284 downloads2y agoHugging Facefport /issue-writer-tr-en Issue Writer — bilingual (EN/TR) instruction dataset Turns raw product input — a Slack message, a support ticket, a Sentry alert, a meeting note — into well-formed issue tracker entries. Every assistant response is a single valid JSON object conforming to schema/issue.schema.json. Balanced across two languages: 50% English, 50% Turkish. Generator, validators, evaluation tooling and the fine-tuning notebook live in github.com/fport/issue-writer. Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fport/issue-writer-tr-en.texttext-generation10K<n<100K1 likes234 downloads19d agoHugging Faceopen-llm-leaderboard-old /details_Writer__palmyra-large Dataset Card for Evaluation run of Writer/palmyra-large Dataset Summary Dataset automatically created during the evaluation run of model Writer/palmyra-large on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-large.0 likes212 downloads3y agoHugging Face