CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01prose-ms /wtm-bench WTM-BENCH (Workbook Time Machine) WTM-BENCH is a benchmark for evaluating LLM agents on realistic, multi-artifact spreadsheet automation tasks. Each task pairs a starting Excel workbook with a natural-language request; the agent must drive the workbook to a target state through a multi-turn tool-calling loop, writing and executing real code each turn. Code, runner, grader, and reproduction rollouts: https://github.com/prose-ms/wtm-bench (see the benchmark.py harness).… See the full description on the dataset page: https://huggingface.co/datasets/prose-ms/wtm-bench.textother1K<n<10K0 likes1.9k downloads1mo agoHugging Face02anonymous-md /ProseOnlyRepair_linguistic_MQ0 likes1k downloads3mo agoHugging Face03anonymous-md /ProseOnlyRepair_linguistic_LQ0 likes819 downloads3mo agoHugging Face04anonymous-md /ProseOnlyRepair_linguistic_OCRrepair_LQ0 likes654 downloads3mo agoHugging Face05anonymous-md /ProseOnlyRepair_linguistic_OCRrepair_MQ10 likes553 downloads3mo agoHugging Face06anonymous-md /ProseOnlyRepair_linguistic_OCRrepair_MQ0 likes404 downloads3mo agoHugging Face07anonymous-md /ProseOnlyRepair_linguistic_MQ10 likes403 downloads3mo agoHugging Face08anonymous-md /ProseOnlyRepair_linguistic_OCRrepair_HQ0 likes365 downloads3mo agoHugging Face09JingminSun /proseThis is data for : PROSE-PDE paper: Towards a Foundation Model for Partial Differential Equations: Multi-Operator Learning and Extrapolation. LeMON: Learning to Learn Multi-Operator Networks. PROSE-SymPy: Time-Series Forecasting and Refinement within a Multimodal PDE Foundation Model. Citation If you find our paper and code useful, please consider citing: @article{sun2024towards, title = {Towards a foundation model for partial differential equations: Multioperator learning… See the full description on the dataset page: https://huggingface.co/datasets/JingminSun/prose.0 likes314 downloads7mo agoHugging Face10anonymous-md /ProseOnlyRepair_linguistic_HQ0 likes224 downloads3mo agoHugging Face11PocketDoc /Dans-Prosemaxx-Opus-Writingtextn<1K1 likes189 downloads2y agoHugging Face12letrinhan /vn-provinces-criminal-cases-prosecuted Vietnam criminal cases prosecuted Vietnam criminal cases prosecuted. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (189 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (18 rows) data/regions.csv data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-criminal-cases-prosecuted.tabularn<1K0 likes98 downloads4d agoHugging Face13PocketDoc /Dans-Prosemaxx-Gutenbergtextn<1K1 likes81 downloads2y agoHugging Face14anonymous-md /ProseOnlyRepair_LQtext100M<n<1B0 likes75 downloads4mo agoHugging Face15PocketDoc /Dans-Prosemaxx-Adventuretabularn<1K4 likes74 downloads2y agoHugging Face16averoo /ls_prose_classictext10K<n<100K0 likes66 downloads2y agoHugging Face17MichaelAnthony /lemonseed-prose lemonseed-prose LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments). Format JSON Lines (.jsonl), one example per line. Provenance & License Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0. texttext-generation1K<n<10K0 likes56 downloads29d agoHugging Face18nevmenandr /accentual-syllabic-verse-in-russian-prose Overview Dataset contains 462 texts of Russian fiction prose of the 19th century, in which accents are marked. Based on this markup, in the texts fragments were found that can be read as verse. original: Как же, ма'тушка! Изве'стно, се'льский во'здух о'чень здоро'в, в кни'гах пи'шут и все говоря'т! Accentual-syllabic fragment marked with italic (4 steps trochee): Как же, ма'тушка! Изве'стно, се'льский во'здух о'чень здоро'в, в кни'гах пи'шут и все говоря'т! Every ' marks the the… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/accentual-syllabic-verse-in-russian-prose.text1M<n<10M1 likes53 downloads2y agoHugging Face19open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v12-Prose-DS-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.tabular10K<n<100K0 likes50 downloads2y agoHugging Face20KRLabsOrg /lettucedetect-prose-hallucination LettuceDetect Prose Hallucination Dataset Token-level hallucination annotations on LLM answers grounded in prose context, drawn from two public RAG hallucination resources and mapped into one unified taxonomy. This is the prose counterpart to the structured-context (code, tool output, documents) collection — together they let a single detector be trained across modalities. Two sources sit side by side, distinguished by the dataset field: dataset Spans Source psiloqa… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/lettucedetect-prose-hallucination.texttoken-classification10K<n<100K1 likes48 downloads4mo agoHugging Face21hmcgovern /de-en-prosetabular1M<n<10M0 likes47 downloads2y agoHugging Face22zerofata /Anime-AMA-ProseDataset of anime / game characters and VTubers being asked various questions referencing their wiki page. Prompts were generated by GLM 4.6 Scenes were generated by GLM 4.6 Response Plan was generated by GLM 4.6 Initial response was generated by GLM 4.6 Rewrite (if enough overused words were found) was done by Gemma 3 27b. The rewrites are generally lower quality than the response provided by GLM, but they offer some different prose / word choices if preferred. Additionally, some rewrites… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Anime-AMA-Prose.text1K<n<10K7 likes46 downloads11mo agoHugging Face23PJMixers-Dev /PocketDoc_Dans-Prosemaxx-Cowriter-XL-8192-shrunk-l3import json from tqdm import tqdm from transformers import AutoTokenizer import re import pandas as pd def load_json_or_jsonl(file_path): try: with open(file_path, "r") as file: try: # Try loading the entire file as JSON data = json.load(file) returndata except json.JSONDecodeError: # If loading as JSON fails, try loading as JSON Lines file.seek(0) # Reset file pointer to the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/PocketDoc_Dans-Prosemaxx-Cowriter-XL-8192-shrunk-l3.text10K<n<100K0 likes45 downloads2y agoHugging Face24proserve /medical-instruct-mixertext100K<n<1M0 likes43 downloads3y agoHugging Face25PocketDoc /Dans-Prosemaxx-Gryphe-GPT4o-WritingPromptstextn<1K0 likes43 downloads2y agoHugging Face26open-llm-leaderboard /sometimesanotion__Qwen-14B-ProseStock-v4-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4 Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face27open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v13-Prose-DS-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v13-Prose-DS Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v13-Prose-DS The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face28wolfvswhale /prose-cadence-stats Prose cadence statistics Measurements of 38 stylometric features across 5,402 documents, split by authorship (human or machine) and by register (informal, formal, multi-paragraph). There is no text in this dataset. Every row is a set of numbers plus a stable reference to the document it was measured from. That is deliberate, and both reasons matter. The sources carry incompatible licenses, so republishing a merged text corpus would be a mess. Measurements are facts about text… See the full description on the dataset page: https://huggingface.co/datasets/wolfvswhale/prose-cadence-stats.tabulartext-classification1K<n<10K0 likes40 downloads2mo agoHugging Face29zarahall /prose-steering-results-n32tabularn<1K0 likes38 downloads1y agoHugging Face30artindnr /khayyam-challenge-prose-terra Khayyam Challenge - Prose (Terra) AI-generated Persian prose descriptions of the 20 classical poems in the Khayyam Challenge benchmark, produced by the model internally labeled Terra. Split low / medium / long by poem length (low: 10, medium: 7, long: 3). Each record contains everything in the poems repo (id, poet, title, form, verse_count, theme, text, ...) plus a conversion object with the generated prose, so this repo is self-contained -- no join required: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-terra.tabularn<1K1 likes37 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.