Liu-Junhua/TwiFF-Bench
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning 🧠 Method We present TwiFF, a unified model fine-tuned on a high-quality dynamic visual Chain-of-Thought (VCoT) dataset comprising 2.7 million samples. In dynamic multimodal question-answering tasks involving instructional, predictive, and camera, TwiFF iteratively generates future event frames alongside textual reasoning, thereby… See the full description on the dataset page: https://huggingface.co/datasets/Liu-Junhua/TwiFF-Bench.
0174
