JonesLin/next-jev-phase1-accepted
NextJev Phase 1 accepted This dataset contains all 137,960 independently verified Phase 1 accepted records. train: 124,341 records compatible with the current three-way NextJev trainer. validation: 2,377 held-out records, split by evidence group with seed 42. excluded: 11,242 accepted records retained for audit but excluded from current NextJev training because their labels cannot be converted losslessly. Images are embedded in the Parquet image column using the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/next-jev-phase1-accepted.
NextJev Phase 1 accepted
This dataset contains all 137,960 independently verified Phase 1 accepted records.
train: 124,341 records compatible with the current three-way NextJev trainer.validation: 2,377 held-out records, split by evidence group with seed 42.excluded: 11,242 accepted records retained for audit but excluded from current NextJev training because their labels cannot be converted losslessly.
Images are embedded in the Parquet image column using the Hugging Face Image representation. The repository intentionally contains a few merged Parquet files instead of thousands of individual image files. materialize_next_jev.py converts the two training splits into the flat JSONL plus deduplicated local images expected by the current NextJev loader.
The training rows use id, premise, hypothesis, image_path, teacher_rationale, split_group, source, and gold_label. Teacher rationales are targets only; they are not input prompts. VQA rows follow the existing NextJev conversion policy and become entailment statements. Binary NLI and low-consensus VQAv2 rows remain in excluded with an explicit reason.
The source export was accepted/review partitioned and independently verified before conversion. See manifest.json for counts, hashes, selection methods, source datasets, and exclusion reasons.
