datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-12m-mbart-50-multilingualconceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file.
We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs.
conceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following:
train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each)
val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.conceptual-12m-multilingual-marian-128This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following:
train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.conceptnet-full-en-essentials
Conceptnet Full EN (essentials)
Dataset Summary:
This dataset is a compact and simplified version of ConceptNet, emphasizing English concepts and their sources. It retains the essential information about the relations in a format that is straightforward and user-friendly. Designed for efficiency and ease of use, this dataset is particularly suitable for scenarios with computational constraints. While the original ConceptNet database exceeds 20GB in size, this streamlined… See the full description on the dataset page: https://huggingface.co/datasets/openworld-domains/conceptnet-full-en-essentials.friend-bench
Can a model — or a human — tell how two people are related from a 20-second clip of how they interact?
🌐 Built on Seamless Interaction
FriendBench is a suite of benchmarks for social perception from thin-slice dyadic
interaction — inferring facts about two people's relationship from a brief clip of how they
interact, built on the Seamless Interaction
dataset. Each released set is a config of this repository.
🎧 Multi-modal — text, audio, and video for every clip
🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.conceptual-12m-multilingual-marian-eshuman-conceptsDataset Summary
This dataset provides digitized versions of classic human categorization benchmarks from seminal cognitive psychology studies by Rosch (1973, 1975) and McCloskey & Glucksberg (1978). These datasets capture human judgments about semantic categories and typicality, offering high-fidelity insights into how humans organize conceptual knowledge.
This dataset was released as part of the study "From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning"
(Shani et al.… See the full description on the dataset page: https://huggingface.co/datasets/CShani/human-concepts.openalex-concepts
OpenAlex L1 + L2 Concepts
What this is
A snapshot of OpenAlex's Level 1 (broad fields) and Level 2 (subfields) concepts: 21,739 records across 284 broad fields and 21,455 subfields.
Level
Count
Examples
1
284
Computer science, Physics, Biology, Sociology
2
21,455
Machine learning, Quantum mechanics, Convolutional neural network
Each record:
Column
Type
Description
id
string
OpenAlex concept ID (e.g., C121955636)
name
string
Concept… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/openalex-concepts.conceptnet_en2en_relations
Dataset Description
This is a subset of the conceptnet5 dataset.
I merely parsed and extracted out my required portion and uploaded here, since processing the huge complete dataset is complicated for many users.
Please refer to the original authors' repo for a complete version.
ConceptNet is a multilingual knowledge base, representing words and
phrases that people use and the common-sense relationships between
them. The knowledge in ConceptNet is collected from a variety of… See the full description on the dataset page: https://huggingface.co/datasets/appledora/conceptnet_en2en_relations.uyghur_conceptnet
ConceptNet Data for the Uyghur Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/uyghur_conceptnet.tibetan_conceptnet
ConceptNet Data for the Tibetan Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/tibetan_conceptnet.javanese_conceptnet
ConceptNet Data for the Javanese Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_conceptnet.technical-concept-simplifier-dataset
Technical Concept Simplifier Dataset
Overview
The Technical Concept Simplifier Dataset is a curated instruction-tuning dataset designed to help Large Language Models (LLMs) explain complex technical concepts in a clear, beginner-friendly, and educational manner.
This dataset was developed as part of an AI model adaptation and fine-tuning project focused on improving the ability of language models to simplify advanced computer science, software engineering, cloud… See the full description on the dataset page: https://huggingface.co/datasets/ujjawalbansal/technical-concept-simplifier-dataset.indonesian_conceptnet
ConceptNet Data for the Indonesian Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/indonesian_conceptnet.quechua_conceptnet
ConceptNet Data for the Quechua Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/quechua_conceptnet.maltese_conceptnet
ConceptNet Data for the Maltese Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/maltese_conceptnet.nepali_conceptnet
ConceptNet Data for the Nepali Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/nepali_conceptnet.mining_conceptssinhala_conceptnet
ConceptNet Data for the Sinhala Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/sinhala_conceptnet.bulgarian_conceptnet
ConceptNet Data for the Bulgarian Language
Dataset Description:
This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub.
Data Structure:
The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/bulgarian_conceptnet.ConceptNetFiltersali_tagsconcept-to-root-dictionary
🌿 Concept-to-Root Dictionary
A mapping of universal concepts to Arabic triliteral roots for semantic compression
📖 Overview
This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models.
What are Arabic Roots?
Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots:
Root
Core Meaning
Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.concept-sentimentconceptNetGenspam_concept_awarelimit_up_conceptconcept_note_grader_1hate_concept_aware
