datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
subCat-human
SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language
A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories.
Dataset Creation
The dataset was created to study how Italian L1 speakers generate exemplars for common… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-human.subCat-llm
SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language
A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories.
This repository contains the generations obtained by prompting a series of LLMs to replicate the human experiment. You can… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-llm.ERCOT_AUS_HOUerc8004-simulated-agents
ERC-8004 Simulated Agents — labeled synthetic dataset (6,000 agents)
⚠️ This dataset is fully synthetic. No public labeled dataset of malicious ERC-8004 agents exists (the standard reached mainnet in 2026 and exposes no trust label), so this dataset simulates the feature distributions the three ERC-8004 registries would expose, for training/evaluating trustworthiness models. For real on-chain data see the companion Base mainnet census.
Composition
6,000 agents, 1… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-simulated-agents.
