CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01quantcodeeval /task_data QuantCodeEval A benchmark for evaluating LLM coding agents on quantitative-strategy code reproduction from finance research papers. Status: Anonymous artifact for the 30-task benchmark. Release mirrors The release is mirrored at two anonymous locations: Hugging Face Datasets — complete anonymous release: https://huggingface.co/datasets/quantcodeeval/task_data anonymous.4open.science — browseable mirror: https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.tabulartext-generationn<1K2 likes2k downloads2mo agoHugging Face02Huyisbeee /ViKm-Translation-Task ViKm-Trans A high-quality synthetic Vietnamese–Khmer parallel corpus. Overview ViKm-Trans is a synthetic parallel corpus for Vietnamese ↔ Khmer machine translation. Due to the scarcity of publicly available Vietnamese–Khmer parallel data, we propose a synthetic data generation framework that leverages abundant Vietnamese monolingual corpora together with large language models to construct high-quality parallel sentence pairs. The dataset was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/Huyisbeee/ViKm-Translation-Task.texttranslation10K<n<100K0 likes47 downloads3mo agoHugging Face03taide /TAIDE-14-tasks Dataset Card for TAIDE-14-tasks Dataset Summary The "TAIDE-14-tasks" dataset, derived from the TAIDE project, encompasses 14 prevalent text generation tasks. This dataset features a collection of 140 prompts tailored for assessing Traditional Chinese Large Language Models (LLM). GPT-4 meticulously crafted these prompts using the provided task, domain, and keywords from the instructions, with further validation by human experts. Each data entry not only contains the main… See the full description on the dataset page: https://huggingface.co/datasets/taide/TAIDE-14-tasks.texttext-generationn<1K30 likes37 downloads3y agoHugging Face04hadeelbkh /tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumtexttext-generation1K<n<10K2 likes30 downloads1y agoHugging Face05lgy0404 /mobileforge-generated-tasks MobileForge Generated Tasks This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation. Dataset summary File Rows Apps Size Description generated_tasks_26020301-all.csv 3,249 20 1.93 MB Consolidated AndroidWorld-side MobileForge task pool. The task pool is generated from real target-app… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-generated-tasks.tabulartext-generation1K<n<10K0 likes19 downloads3mo agoHugging Face061-800-SHARED-TASKS /telugu-summarization-generation Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/telugu-summarization-generation.texttext-generation100K<n<1M0 likes11 downloads2y agoHugging Face07asishley /task-specs-small About Tasks specs texttext-generationn<1K0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.