datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HyperThink-X-Nvidia-Opencode-Reasoning-200K
🔮 HyperThink
HyperThink is a premium, best-in-class dataset series capturing deep reasoning interactions between users and an advanced Reasoning AI system. Designed for training and evaluating next-gen language models on complex multi-step tasks, the dataset spans a wide range of prompts and guided thinking outputs.
🚀 Dataset Tiers
HyperThink is available in three expertly curated versions, allowing flexible scaling based on compute resources and training goals:… See the full description on the dataset page: https://huggingface.co/datasets/Sashvat/HyperThink-X-Nvidia-Opencode-Reasoning-200K.Hyperphantasia
A Benchmark for Evaluating the
Mental Visualization Capabilities of Multimodal LLMs
Mohammad Shahab Sepehri
Berk Tinaz
Zalan Fabian
Mahdi Soltanolkotabi
Github Repository
Hyperphantasia is a synthetic Visual Question Answering (VQA) benchmark dataset that probes the mental visualization capabilities of Multimodal Large Language Models (MLLMs) from a vision perspective. We reveal that state-of-the-art models struggle with simple tasks that require visual… See the full description on the dataset page: https://huggingface.co/datasets/shahab7899/Hyperphantasia.hyperion-v2.0
Hyperion v2.0
Introduction
Hyperion is a comprehensive question answering and conversational dataset designed to promote advancements in AI research with a particular emphasis on reasoning and understanding in scientific domains such as science, medicine, mathematics, and computer science. It integrates data from a wide array of datasets, facilitating the development of models capable of handling complex inquiries and instructions.
Dataset Description
Hyperion… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v2.0.hyperion-v3.0Hyperion-3.0 has significantly improved performance over its predecessors.
"I found that having more code datasets than general purpose datasets ironically decreases performance in both coding and general tasks."
Data sources:
OpenOrca/SlimOrca
cognitivecomputations/dolphin (300k examples)
microsoft/orca-math-word-problems-200k (60k examples)
glaiveai/glaive-code-assistant
Vezora/Tested-22k-Python-Alpaca
Unnatural Instructions
BI55/MedText
LDJnr/Pure-Dove
Various domain-specific datasets by… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v3.0.2SyntheticDatasetSmallhypernymy_pairs
Ukrainian Hypernymy Pairs Dataset
Background
Hypernymy is the super-subordinate or ISA semantic relation that links more general terms to more specific ones. For example, rose is a hyponym of flower, a hypernym of rose. Words that are hyponyms of the same hypernym are called co-hyponyms, for instance, rose and tulip. Hyponymy relation is transitive and asymmetric.
Hypernymy is also differentiated by:
Types — common nouns: armchair is a type (hyponym) of chair;
Instances… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/hypernymy_pairs.Hinglishqna-llm-tuneThis dataset is hinglish qna ,which involves abusive language from customers in financial domain. It also involves emotional cues such as , etc.
