datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vn-provinces-criminal-cases-prosecuted
Vietnam criminal cases prosecuted
Vietnam criminal cases prosecuted. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (189 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (18 rows)
data/regions.csv
data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-criminal-cases-prosecuted.Dans-Prosemaxx-Adventuresometimesanotion__Qwenvergence-14B-v12-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.de-en-prosesometimesanotion__Qwen-14B-ProseStock-v4-details
Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v13-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v13-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details.prose-cadence-stats
Prose cadence statistics
Measurements of 38 stylometric features across 5,402 documents, split by authorship (human or machine) and by register (informal, formal, multi-paragraph).
There is no text in this dataset. Every row is a set of numbers plus a stable reference to the document it was measured from. That is deliberate, and both reasons matter.
The sources carry incompatible licenses, so republishing a merged text corpus would be a mess. Measurements are facts about text… See the full description on the dataset page: https://huggingface.co/datasets/wolfvswhale/prose-cadence-stats.prose-steering-results-n32khayyam-challenge-prose-terra
Khayyam Challenge - Prose (Terra)
AI-generated Persian prose descriptions of the 20 classical poems in the
Khayyam Challenge benchmark,
produced by the model internally labeled Terra. Split low /
medium / long by poem length (low: 10, medium: 7, long: 3).
Each record contains everything in the
poems repo (id,
poet, title, form, verse_count, theme, text, ...) plus a
conversion object with the generated prose, so this repo is self-contained
-- no join required:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-terra.khayyam-challenge-prose-luna
Khayyam Challenge - Prose (Luna)
AI-generated Persian prose descriptions of the 20 classical poems in the
Khayyam Challenge benchmark,
produced by the model internally labeled Luna. Split low /
medium / long by poem length (low: 10, medium: 7, long: 3).
Each record contains everything in the
poems repo (id,
poet, title, form, verse_count, theme, text, ...) plus a
conversion object with the generated prose, so this repo is self-contained
-- no join required:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-luna.arena-prose-100-49-models
Arena Prose: 100 prompts × 50 models
A paired exploratory AI-text-detection corpus: 5,000 successful generated responses from 50 models, each answering the same 100 English prose prompts. Generation was performed through OpenRouter in September 2026 with optional reasoning disabled and mandatory reasoning set to low. This is an independent local benchmark inspired by Pangram 4 §5.2, not an official Pangram dataset or exact replication.
Loading
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/woog/arena-prose-100-49-models.prose-steering-results-n35sft-bm-prose
khursanirevo/sft-bm-prose
Bahasa Melayu prose-format text (long-form lessons + textbook-style, ~232k rows).
Splits
split
rows
train
221,170
validation
11,635
Stratified 95/5 by source/category (seed=42).
Source files
data/midtrain/synth_hf_prose.jsonl
data/midtrain/synth_bm_50m.jsonl
Schema
Each row is a JSON object. See the loader script for field details.
Provenance
Generated as part of MaLLaM 2026… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/sft-bm-prose.sometimesanotion__lamarck-14b-prose-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/lamarck-14b-prose-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/lamarck-14b-prose-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__lamarck-14b-prose-model_stock-details.sometimesanotion__Qwenvergence-14B-v15-Prose-MS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v15-Prose-MS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v15-Prose-MS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v15-Prose-MS-details.sometimesanotion__Qwenvergence-14B-v3-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v3-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v3-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v3-Prose-details.sometimesanotion__Qwenvergence-14B-v6-Prose-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v6-Prose-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v6-Prose-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v6-Prose-model_stock-details.sometimesanotion__Qwenvergence-14B-v6-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v6-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v6-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v6-Prose-details.prose-steering-results-n34prose-steering-results-n325prose-steering-results-n326sometimesanotion__Qwentinuum-14B-v6-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwentinuum-14B-v6-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwentinuum-14B-v6-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwentinuum-14B-v6-Prose-details.prose-steering-results-n323prose-steering-results-n324sometimesanotion__Qwenvergence-14B-v2-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v2-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v2-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v2-Prose-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details.sometimesanotion__Qwenvergence-14B-v12-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-details.sometimesanotion__Qwen2.5-7B-Gordion-v0.1-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-7B-Gordion-v0.1-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-7B-Gordion-v0.1-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-7B-Gordion-v0.1-Prose-details.prose-steering-results-n36prose-steering-results-n322
