datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gigahands-vitra-mano
GigaHands → VITRA Stage-1, official-MANO annotations
VITRA Stage-1 hand annotations for GigaHands, with all joint positions taken from GigaHands'
official MANO fit instead of mixing in triangulated keypoints. Annotations only — no videos
(get those from GigaHands; the mapping is described in §5).
episodes
13,247 (train 11,904 / test 1,343)
frames
3,395,733
camera
brics-odroid-001_cam0 (static rig; one constant extrinsic per scene)
source
GigaHands params/ +… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/gigahands-vitra-mano.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-4.7-reasoning-8.7k.ManoloPueblo__LLM_MERGE_CC2-details
Dataset Card for Evaluation run of ManoloPueblo/LLM_MERGE_CC2
Dataset automatically created during the evaluation run of model ManoloPueblo/LLM_MERGE_CC2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ManoloPueblo__LLM_MERGE_CC2-details.ManoloPueblo__LLM_MERGE_CC3-details
Dataset Card for Evaluation run of ManoloPueblo/LLM_MERGE_CC3
Dataset automatically created during the evaluation run of model ManoloPueblo/LLM_MERGE_CC3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ManoloPueblo__LLM_MERGE_CC3-details.ManoloPueblo__ContentCuisine_1-7B-slerp-details
Dataset Card for Evaluation run of ManoloPueblo/ContentCuisine_1-7B-slerp
Dataset automatically created during the evaluation run of model ManoloPueblo/ContentCuisine_1-7B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ManoloPueblo__ContentCuisine_1-7B-slerp-details.blender_duplicates
Dataset Card for Dataset Name
Contains reduced description of issues reported at https://projects.blender.org/blender/blender/issues and points to duplicate issues in order to categorize similarity.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Each report has been shortened by removing frequently repeated texts such as System Information, Blender Version… See the full description on the dataset page: https://huggingface.co/datasets/mano-wii/blender_duplicates.guanaco-pico-100-samplesThis dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main
This subset of the data only contains 100 samples of the highest-rated paths in the conversation tree.
This dataset was used to train Guanaco with QLoRA.
For further information, please see the original dataset.
License: Apache 2.0
elira-npc-datasetNL2SQLclaude-opus-4.6-10000xThis is a high-fidelity reasoning dataset synthesized using Claude Opus 4.6. The dataset is designed to capture the model's internal "Chain of Thought" and reasoning traces, specifically focusing on mathematical accuracy and structured logical deduction.
The dataset is intended for Supervised Fine-Tuning (SFT) and Distillation, allowing smaller open-source models to inherit the sophisticated reasoning patterns of Claude Opus 4.6.
Dataset Description
This collection combines high-difficulty… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-10000x.Manon_Lescaut_datasetempathetic-dialogue-datasettestdataset1Virtual-labuday_infoUAV_V0
