datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
loopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.Open-Sora-Plan-v1.3.0We have open-sourced our dataset of 32,555 pairs, which includes Chinese data. The dataset is available here. The details can be found here.
In fact, it is a JSON file with the following structure. More details can be found here.
[
{
"instruction": "Refine the sentence: \"A newly married couple sharing a piece of there wedding cake.\" to contain subject description, action, scene description. (Optional: camera language, light and shadow, atmosphere) and conceive some additional actions… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.3.0.Sorawiz__Gemma-Creative-9B-Base-details
Dataset Card for Evaluation run of Sorawiz/Gemma-Creative-9B-Base
Dataset automatically created during the evaluation run of model Sorawiz/Gemma-Creative-9B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sorawiz__Gemma-Creative-9B-Base-details.Sorawiz__Gemma-9B-Base-details
Dataset Card for Evaluation run of Sorawiz/Gemma-9B-Base
Dataset automatically created during the evaluation run of model Sorawiz/Gemma-9B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sorawiz__Gemma-9B-Base-details.
