datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shotplan
ShotPlan Training Dataset
Multi-shot video training data for ShotPlan: Cinematic Video Generation with Learnable Planning Token.
💻 Code: https://github.com/Pensioner-11/ShotPlan
🤖 Models: ShotPlan-Wan2.1-T2V-14B · ShotPlan-Wan2.2-T2V-A14B-HighNoise
Contents
Path
Description
data/train_meta_16fps.json
6,404 training samples (metadata + captions)
data/videos/V*_16fps.mp4
549 source videos, re-encoded to 16 fps
Each sample is an 80-frame (5 s @… See the full description on the dataset page: https://huggingface.co/datasets/Pensioner/shotplan.lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-tulu-human-v2
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-tulu-human-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private.lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-NuminaCoT
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-NuminaCoT
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private.pentabrid-reproducibility
Pentabrid 27B: reproducibility package
Everything required to recompute the results of a controlled evaluation of fine-tuning
configurations for medical question answering. Openly available with no access
restrictions.
Contents
Path
Description
per_item/medxpertqa_*.jsonl
Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.lm-eval-results-penfever-Amber-7B-000-tulu-v2-private
Dataset Card for Evaluation run of penfever/Amber-7B-000-tulu-v2
Dataset automatically created during the evaluation run of model penfever/Amber-7B-000-tulu-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Amber-7B-000-tulu-v2-private.PenguinScrolls
PenguinScrolls: A User-Aligned Fine-Grained Benchmark for Long-Context Language Model Evaluation
Introduction
PenguinScrolls (企鹅卷轴) is a comprehensive benchmark designed to evaluate and enhance the long-text processing capabilities of large language models (LLMs).
Current benchmarks for evaluating long-context language models often rely on synthetic tasks that fail to adequately reflect real user needs, leading to a weak correlation between benchmark scores and actual… See the full description on the dataset page: https://huggingface.co/datasets/Penguin-Scrolls/PenguinScrolls.lm-eval-results-penfever-Mistral-7B-tulu-v2-private
Dataset Card for Evaluation run of penfever/Mistral-7B-tulu-v2
Dataset automatically created during the evaluation run of model penfever/Mistral-7B-tulu-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Mistral-7B-tulu-v2-private.gsm8kcommtokenHoangHa__Pensez-Llama3.1-8B-details
Dataset Card for Evaluation run of HoangHa/Pensez-Llama3.1-8B
Dataset automatically created during the evaluation run of model HoangHa/Pensez-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HoangHa__Pensez-Llama3.1-8B-details.DreadPoor__Minus_Penus-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Minus_Penus-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Minus_Penus-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Minus_Penus-8B-Model_Stock-details.kb-yurisprudensi-pencurianpencilboxDreadPoor__tests_pending-do_not_use_yet-details
Dataset Card for Evaluation run of DreadPoor/tests_pending-do_not_use_yet
Dataset automatically created during the evaluation run of model DreadPoor/tests_pending-do_not_use_yet
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__tests_pending-do_not_use_yet-details.pen-into-box
pen-into-box
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Pendulum-v17578578testpmon
