datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wtm-bench
WTM-BENCH (Workbook Time Machine)
WTM-BENCH is a benchmark for evaluating LLM agents on realistic,
multi-artifact spreadsheet automation tasks. Each task pairs a starting
Excel workbook with a natural-language request; the agent must drive the
workbook to a target state through a multi-turn tool-calling loop, writing and
executing real code each turn.
Code, runner, grader, and reproduction rollouts:
https://github.com/prose-ms/wtm-bench (see the benchmark.py harness).… See the full description on the dataset page: https://huggingface.co/datasets/prose-ms/wtm-bench.ProseOnlyRepair_linguistic_MQProseOnlyRepair_linguistic_LQProseOnlyRepair_linguistic_OCRrepair_LQProseOnlyRepair_linguistic_OCRrepair_MQ1ProseOnlyRepair_linguistic_OCRrepair_MQProseOnlyRepair_linguistic_MQ1ProseOnlyRepair_linguistic_OCRrepair_HQproseThis is data for :
PROSE-PDE paper: Towards a Foundation Model for Partial Differential Equations: Multi-Operator Learning and Extrapolation.
LeMON: Learning to Learn Multi-Operator Networks.
PROSE-SymPy: Time-Series Forecasting and Refinement within a Multimodal PDE Foundation Model.
Citation
If you find our paper and code useful, please consider citing:
@article{sun2024towards,
title = {Towards a foundation model for partial differential equations: Multioperator learning… See the full description on the dataset page: https://huggingface.co/datasets/JingminSun/prose.ProseOnlyRepair_linguistic_HQDans-Prosemaxx-Opus-Writingvn-provinces-criminal-cases-prosecuted
Vietnam criminal cases prosecuted
Vietnam criminal cases prosecuted. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (189 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (18 rows)
data/regions.csv
data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-criminal-cases-prosecuted.Dans-Prosemaxx-GutenbergProseOnlyRepair_LQDans-Prosemaxx-Adventurels_prose_classiclemonseed-prose
lemonseed-prose
LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments).
Format
JSON Lines (.jsonl), one example per line.
Provenance & License
Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0.
accentual-syllabic-verse-in-russian-prose
Overview
Dataset contains 462 texts of Russian fiction prose of the 19th century, in which accents are marked. Based on this markup, in the texts fragments were found that can be read as verse.
original:
Как же, ма'тушка! Изве'стно, се'льский во'здух о'чень здоро'в, в кни'гах пи'шут и все говоря'т!
Accentual-syllabic fragment marked with italic (4 steps trochee):
Как же, ма'тушка! Изве'стно, се'льский во'здух о'чень здоро'в, в кни'гах пи'шут и все говоря'т!
Every ' marks the the… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/accentual-syllabic-verse-in-russian-prose.sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.lettucedetect-prose-hallucination
LettuceDetect Prose Hallucination Dataset
Token-level hallucination annotations on LLM answers grounded in prose
context, drawn from two public RAG hallucination resources and mapped into one
unified taxonomy. This is the prose counterpart to the structured-context
(code, tool output, documents)
collection — together they let a single detector be trained across modalities.
Two sources sit side by side, distinguished by the dataset field:
dataset
Spans
Source
psiloqa… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/lettucedetect-prose-hallucination.de-en-proseAnime-AMA-ProseDataset of anime / game characters and VTubers being asked various questions referencing their wiki page.
Prompts were generated by GLM 4.6
Scenes were generated by GLM 4.6
Response Plan was generated by GLM 4.6
Initial response was generated by GLM 4.6
Rewrite (if enough overused words were found) was done by Gemma 3 27b.
The rewrites are generally lower quality than the response provided by GLM, but they offer some different prose / word choices if preferred. Additionally, some rewrites… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Anime-AMA-Prose.PocketDoc_Dans-Prosemaxx-Cowriter-XL-8192-shrunk-l3import json
from tqdm import tqdm
from transformers import AutoTokenizer
import re
import pandas as pd
def load_json_or_jsonl(file_path):
try:
with open(file_path, "r") as file:
try:
# Try loading the entire file as JSON
data = json.load(file)
returndata
except json.JSONDecodeError:
# If loading as JSON fails, try loading as JSON Lines
file.seek(0) # Reset file pointer to the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/PocketDoc_Dans-Prosemaxx-Cowriter-XL-8192-shrunk-l3.medical-instruct-mixerDans-Prosemaxx-Gryphe-GPT4o-WritingPromptssometimesanotion__Qwen-14B-ProseStock-v4-details
Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v13-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v13-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details.prose-cadence-stats
Prose cadence statistics
Measurements of 38 stylometric features across 5,402 documents, split by authorship (human or machine) and by register (informal, formal, multi-paragraph).
There is no text in this dataset. Every row is a set of numbers plus a stable reference to the document it was measured from. That is deliberate, and both reasons matter.
The sources carry incompatible licenses, so republishing a merged text corpus would be a mess. Measurements are facts about text… See the full description on the dataset page: https://huggingface.co/datasets/wolfvswhale/prose-cadence-stats.prose-steering-results-n32khayyam-challenge-prose-terra
Khayyam Challenge - Prose (Terra)
AI-generated Persian prose descriptions of the 20 classical poems in the
Khayyam Challenge benchmark,
produced by the model internally labeled Terra. Split low /
medium / long by poem length (low: 10, medium: 7, long: 3).
Each record contains everything in the
poems repo (id,
poet, title, form, verse_count, theme, text, ...) plus a
conversion object with the generated prose, so this repo is self-contained
-- no join required:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-terra.
