datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shot2story
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Please download the multi-shot videos from OneDrive or HuggingFace.
We are excited to release a new video-text benchmark for multi-shot video understanding. This release contains a 134k version of our dataset. It includes detailed long summaries (human annotated + GPTV generated) for 134k videos and shot captions (human annotated) for 188k video shots.
Annotation Format
Our 134k multi-shot… See the full description on the dataset page: https://huggingface.co/datasets/mhan/shot2story.shotplan
ShotPlan Training Dataset
Multi-shot video training data for ShotPlan: Cinematic Video Generation with Learnable Planning Token.
💻 Code: https://github.com/Pensioner-11/ShotPlan
🤖 Models: ShotPlan-Wan2.1-T2V-14B · ShotPlan-Wan2.2-T2V-A14B-HighNoise
Contents
Path
Description
data/train_meta_16fps.json
6,404 training samples (metadata + captions)
data/videos/V*_16fps.mp4
549 source videos, re-encoded to 16 fps
Each sample is an 80-frame (5 s @… See the full description on the dataset page: https://huggingface.co/datasets/Pensioner/shotplan.Shot2Story-20K
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
We have a more recent release of 134K version here. Please have a look.
For video data downloading, please have a look at this issue.
We are excited to release a new video-text benchmark for multi-shot video understanding. This release contains a 134k version of our dataset. It includes detailed long summaries (human annotated + GPTV generated) for 134k videos and shot captions (human annotated) for… See the full description on the dataset page: https://huggingface.co/datasets/mhan/Shot2Story-20K.no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.shot2story
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Please download the multi-shot videos from OneDrive or HuggingFace.
We are excited to release a new video-text benchmark for multi-shot video understanding. This release contains a 134k version of our dataset. It includes detailed long summaries (human annotated + GPTV generated) for 134k videos and shot captions (human annotated) for 188k video shots.
Annotation Format
Our 134k multi-shot… See the full description on the dataset page: https://huggingface.co/datasets/huankguan2/shot2story.msmarco-atomic-id-3shot-v4_128k_few_shot
msmarco-atomic-id-3shot-v4_128k
MSMARCO few-shot evaluation dataset for in-context learning generative retrieval,
atomic-id variant.
Identical construction to
Lala8383/msmarco-item-id-3shot-v4_128k_few_shot,
except every document's Identifier (and the answer target) is an arbitrary
unique integer (Tay et al. DSI "Atomic Docid") instead of the natural-language
document title. The id carries no semantics, so a retriever can only answer by
matching the query to a document in… See the full description on the dataset page: https://huggingface.co/datasets/Lala8383/msmarco-atomic-id-3shot-v4_128k_few_shot.shotpath-grpo-trajectory-audit-20260729
ShotPath GRPO trajectory audit
This audit reconstructs the committed 260-step trajectory by keeping the last logged occurrence of each pair after time-limit rollbacks. It includes compact statistics for every committed group and detailed candidate/judge records plus pre/post images for stratified and contrast samples.
capybara-sharegpt
capybara-sharegpt
LDJnr/Capybara converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information. All credit goes to the original creator.
nq-item-id-llm-bullets-hardneg-few_shot-v4theory-of-mind-dpoThis is grimulkan/theory-of-mind with "rejected" responses generated using mistralai/Mistral-7B-Instruct-v0.2, and the file formatted for use in DPO training.
The code used to generate the dataset can be found in this repository: https://github.com/DocShotgun/LLM-datagen
cricket-shot
CricketShotClassification Dataset
Dataset Description
This dataset is designed for video classification of cricket shots. It contains labeled videos of ten different cricket shots, making it suitable for training and evaluating machine learning models for cricket action recognition.
Dataset Structure
The dataset contains videos of ten cricket shots:
Shot Name
Label
Class ID
Cover Drive
cover
0
Defense Shot
defense
1
Flick Shot
flick
2… See the full description on the dataset page: https://huggingface.co/datasets/lokeshreddy700232/cricket-shot.Chinese_Few-shot_NERcricket-shot-video-dataset
CricketShotClassification Dataset
Dataset Description
This dataset is designed for video classification of cricket shots. It contains labeled videos of ten different cricket shots, making it suitable for training and evaluating machine learning models for cricket action recognition.
Dataset Structure
The dataset contains videos of ten cricket shots:
Shot Name
Label
Class ID
Cover Drive
cover
0
Defense Shot
defense
1
Flick Shot
flick
2… See the full description on the dataset page: https://huggingface.co/datasets/rnrahate007/cricket-shot-video-dataset.rds-sels-bbh-shots-top326k
RDS+ Selected BBH shots 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using BBH few-shot samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-bbh-shots-top326k.cricket-shot1
CricketShotClassification Dataset
Dataset Description
This dataset is designed for video classification of cricket shots. It contains labeled videos of ten different cricket shots, making it suitable for training and evaluating machine learning models for cricket action recognition.
Dataset Structure
The dataset contains videos of ten cricket shots:
Shot Name
Label
Class ID
Cover Drive
cover
0
Defense Shot
defense
1
Flick Shot
flick
2
Hook Shot… See the full description on the dataset page: https://huggingface.co/datasets/nimra868/cricket-shot1.Doctor-Shotgun_theory-of-mind-dpo-PreferenceShareGPTrds-sels-gsm8k-shots-top326k
RDS+ Selected GSM8k shots 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using GSM8k few-shot samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-gsm8k-shots-top326k.cricket-shot
CricketShotClassification Dataset
Dataset Description
This dataset is designed for video classification of cricket shots. It contains labeled videos of ten different cricket shots, making it suitable for training and evaluating machine learning models for cricket action recognition.
Dataset Structure
The dataset contains videos of ten cricket shots:
Shot Name
Label
Class ID
Cover Drive
cover
0
Defense Shot
defense
1
Flick Shot
flick
2
Hook Shot… See the full description on the dataset page: https://huggingface.co/datasets/aroramoksh11/cricket-shot.cricket-shot
CricketShotClassification Dataset
Dataset Description
This dataset is designed for video classification of cricket shots. It contains labeled videos of ten different cricket shots, making it suitable for training and evaluating machine learning models for cricket action recognition.
Dataset Structure
The dataset contains videos of ten cricket shots:
Shot Name
Label
Class ID
Cover Drive
cover
0
Defense Shot
defense
1
Flick Shot
flick
2… See the full description on the dataset page: https://huggingface.co/datasets/vicky521/cricket-shot.zero_shot_classification_testkalo-opus-misc-kto-combined5-shot-test-ID-CoTrds-sels-mmlu-shots-top326k
RDS+ Selected MMLU shots 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using MMLU few-shot samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-mmlu-shots-top326k.prepared_data_file_zero_shot_prompting_5000Q_evidence_selected_plus_verbalizationprepared_data_file_zero_shot_prompting_5000Q_evidence_selected_plus_verbalization
camera-shot-templates
Camera Shot Templates
镜头语言模板库:镜头类型、运镜、灯光、情绪、Kling 示例 Prompt。
1-shot-cft-datakalo-22k-norefusal-kto-combinedkalo-3k-filtered-kto-combinedzero-shot-classification-large-testzero_shot_pubMed
