datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BS-Objaverse
BS-Objaverse 660K Dataset Card
Dataset details
Dataset type:
BS-Objaverse 660k Dataset is a set of GPT4-Vision-powered multi-modal captions data.
It is constructed to enhance modality alignment and fine-grained visual concept perception for describing detailed information about the shape, texture of Objaverse 3D object.
obj_descript_gpt_10k.json is generated by GPT4-Vision.
objaverse660k_mvllava7b.json is generated by our MV-LLaVA trained on GPT4-Vision-generated data.… See the full description on the dataset page: https://huggingface.co/datasets/Zery/BS-Objaverse.3D-Object-ConflictE2E_real_object
E2E Real Object Direction
A video-based benchmark for evaluating VideoLLMs' directional reasoning and object recognition on real-world objects.
Conditions
Condition
Question
Answer
Purpose
direction_only
"In which direction is the object moving?"
"Up"
Baseline direction recognition
direction_obj_in_q
"In which direction is the car moving?"
"Up"
Does naming the object help?
direction_obj_in_a
"In which direction is the object moving?"
"The car is moving up"… See the full description on the dataset page: https://huggingface.co/datasets/KHUjongseo/E2E_real_object.cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5
CybersecJSONSchemaBench CloudTrail Objective Natural Hard v5
This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export.
Families
actor_recon_to_change: 22… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5.Task_saliency_objectscybersec-jsonschemabench-cloudtrail-objective-hard-v3
CybersecJSONSchemaBench CloudTrail Objective Hard v3
This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export.
Families
apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.multilingual-crossmodal-conflict-3D_Objects
Multilingual Cross-Modal Conflict — 3D Objects
A multilingual counterfactual MCQ dataset built from rendered 3D object scenes.
Each row contains a rendered 3D scene image, two captions (original vs counterfactual), and a multiple-choice question probing whether a VLM follows the image or the misleading text.
Languages
Language
Code
Rows
English
en
150
Hindi
hi
150
Telugu
te
150
Bahasa Indonesia
id
150
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/multilingual-crossmodal-conflict-3D_Objects.Task_relevant_objectsStack2Graph_VD_objective-c
Objective-C StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the Objective-C shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_objective-c.Object_detection_dataset
