datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Jev
Open-Jev: typed decision datasets
Open-Jev turns a state and a question into a typed decision: a yes/no probability, a distribution over choices, independent label probabilities, or a discrete numeric/ordinal decision. This repository publishes twelve separate, frozen data configs from the Open-Jev project, together with original manifests, exact raw records, source code and reconstruction instructions.
These are controlled, mostly synthetic tasks and reference labels. They are… See the full description on the dataset page: https://huggingface.co/datasets/ZefanCai/Open-Jev.jev-stage2-image-beans-pilot
Beans: one natural question per image
Open the corrected preview.
natural_v4 is the recommended and default preview: 100 original images, 100 rows, one three-way condition-class Choice question per image. All targets come directly from the source labels column (34 angular leaf spot, 33 bean rust, 33 healthy). Original image bytes and source annotations are unchanged.
Example question: “Which source-defined condition class describes the bean leaf?” Options: angular_leaf_spot… See the full description on the dataset page: https://huggingface.co/datasets/FaroukMoc2/jev-stage2-image-beans-pilot.jev-decisions-v1
Jev Decisions v1
12M canonical agent-decision records for tool selection, routing, value prediction, completion, and local agent control.
Jev Decisions v1 is a derived, decision-oriented corpus built from public agent trajectory datasets. It canonicalizes heterogeneous trajectories into a shared learning interface:
state + available candidate decisions -> target / outcome / eligibility
mini-Jev is a related open decision-model project. Its currently downloadable v1 baseline was… See the full description on the dataset page: https://huggingface.co/datasets/samatv256/jev-decisions-v1.tasksource-jev-typed-decisions
tasksource-jev-typed-decisions
2.5 million typed decisions (choices, ratings and probabilities) from 670 sources.
Why use it
Real supervision. Labels, ratings, and annotator votes come from
established datasets, not a teacher model. Every row names its source.
Breadth. Over 300 dataset families: NLI and reasoning, QA and
commonsense, sentiment, intent and topic, toxicity and safety, preference
pairs, fact checking, entity tagging, and dozens of languages. GLUE… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.Open-Jev-v1.1
Open-Jev v1.1 data
This release publishes 326,619 typed decision records, including 147,139 training records, from the frozen community-hard-mix-v2-final mixture used for the new Open-Jev 27B training stage. It is a redistributable projection, not the exact complete training or evaluation dataset: 2,053 Wikispeedia records are omitted because a separate redistribution license for the archived graph/path data has not been verified.
Project website · Code · Previous data release ·… See the full description on the dataset page: https://huggingface.co/datasets/ZefanCai/Open-Jev-v1.1.JevEmbed-Data
JevEmbed-Data
This dataset contains 1,601,157 training and 66,482 test questions for JevEmbed. Here, test is the source corpus's validation split. The ten Parquet files each hold at most 200,000 rows.
Fields
Each row is one decision. request_json and answers_json are JSON strings; parse them with json.loads.
Field
Meaning
id
Unique question ID.
group
Case ID. Related questions stay in the same split.
request_json
Input: state plus a decision… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/JevEmbed-Data.gtow-llama-sft-v3
GTO Wizard — Heads-Up NL Hold'em 200BB — SFT dataset (v3)
Supervised fine-tuning data for heads-up No-Limit Texas Hold'em, 200 big
blinds deep. Each row is a single decision point: a natural-language
description of the game state, paired with the game-theory-optimal action
GTO Wizard chose in that spot.
Intended for instruction-tuning a chat LLM to play HU 200BB poker (see the
pokerbench agent it was built for).
Schema
Two flat columns:
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/jevonmao/gtow-llama-sft-v3.aave_matchedjev-japanese-judgmentnext-jev-phase1-accepted
NextJev Phase 1 accepted
This dataset contains all 137,960 independently verified Phase 1 accepted records.
train: 124,341 records compatible with the current three-way NextJev trainer.
validation: 2,377 held-out records, split by evidence group with seed 42.
excluded: 11,242 accepted records retained for audit but excluded from current NextJev training because their labels cannot be converted losslessly.
Images are embedded in the Parquet image column using the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/next-jev-phase1-accepted.next-jev-stage1-multi-model-cot
Next JEV Stage 1 multi-model CoT
One row is one prompt with its paired model responses. images is a list of
embedded Hugging Face Image features, ordered with image_sha256s. record_json
preserves the exact training row for reproducible materialization. This is a
CoT-only Stage 1 export, with no answer or classification target required.
The original source, split counts and JSONL SHA-256 values are in
reports/manifest.json; Parquet SHA-256 values are in reports/native.json.
The… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/next-jev-stage1-multi-model-cot.je-veux-un-pack-de-computer-vision-pour-detecter-les-objets
Je veux un pack de computer vision pour detecter les objets,…
Je veux un pack de computer vision pour detecter les objets, dans des salles de bains pour le segment home avec des depth map
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/je-veux-un-pack-de-computer-vision-pour-detecter-les-objets.next-jev-stage2-merged-verified-20260923
Next JEV Stage 2 verified merge
This dataset contains all 414,839 accepted records from the verified
Phase 1, 2 and 3 generation batches. Generation batch names are distinct from model
training stages. Phase 3 generation was still incomplete at this snapshot.
train: 378,903 records compatible with the current three-way NextJev trainer.
validation: 7,560 held-out records, split by evidence group with seed 42.
excluded: 28,376 accepted records retained for audit but excluded from… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/next-jev-stage2-merged-verified-20260923.jevm-news
jevm news screening dataset
Chinese financial news flash, each item labelled by how the A-share market actually reacted to it.
This is the news line training set of jevm: one row per news item, with
every feature computed strictly as of the publish time τ. The labels come only from minute-bar price and
turnover reactions and from objective propagation evidence. No labels come from model scores.
中文说明见下方
At a glance
Split
Period
Rows
Head-B labelled
30-min… See the full description on the dataset page: https://huggingface.co/datasets/TriadParty/jevm-news.jev-1min-stock-barsJevanhielle_Zyhamont_outCheckedfinetuning_demoturkish_jev_noul
turkish_jev_noul
Türkçe evet/hayır (noul) karar soruları. Format, Laya modelinin kullandığı LocalLLaMA/typed-decisions formatıyla aynı: bir metin (state), bu metin hakkında tipli bir soru (questions) ve cevabın olasılık dağılımı (gold).
Turkish yes/no (noul) decision questions in the typed-decisions format used by Laya, built from Turkish TrGLUE, Belebele and a Turkish toxic-language set with hand-reviewed question templates.
Bağlantı notu: Bu veri seti bağımsızdır. TypeSafe ya… See the full description on the dataset page: https://huggingface.co/datasets/ahmetege/turkish_jev_noul.jva-missions-report-raw
Dataset Card for "jva-missions-report-raw"
More Information needed
tape-testpokemon-with-pokedex-descriptions
Dataset Card for "pokemon-with-pokedex-descriptions"
More Information needed
tape-pocknihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_all
AudioSet Pipeline Output
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
1,023
Працягласць
3 гадз 26 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_all.2026-03-03_pika-test-bracket-insertiontape-bin-picking-rewardknihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_output_original
AudioSet Pipeline Output — арыгінальнае аўдыё
Мова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
knihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_output
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid — унікальны ідэнтыфікатар
Ліцэнзія… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_output_original.polaroid-button-presspolaroid-testJevanhielle_Zyhamont0_10_Checkedtape-poc-reward
