datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synapse-set-10k
🧠 SynapseSet-10K
SynapseSet-10K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation.
This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems.
🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-10k.synapse-set-50k
🧠 SynapseSet-50K
SynapseSet-50K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation.
This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems.
🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-50k.synapse-set-100k
🧠 SynapseSet-100K
SynapseSet-100K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation.
This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems.
🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-100k.Synapse-PT-75M-Dataset
🧠 Synapse PT-BR 75M Dataset
Dataset em Português do Brasil desenvolvido pela Comunidade Synapse-BR para treinamento e ajuste fino de modelos de linguagem (LLMs).
🇧🇷 Sobre
O Synapse PT-BR 75M Dataset é um conjunto de dados em Português do Brasil cuidadosamente selecionado, organizado e processado para treinamento de modelos de linguagem.
O projeto faz parte da Comunidade Synapse-BR, que busca desenvolver recursos abertos para fortalecer o ecossistema de… See the full description on the dataset page: https://huggingface.co/datasets/Comunidade-Synapse-BR/Synapse-PT-75M-Dataset.synapse-nomes-brasileiros🇧🇷 Synapse Nomes brasileiros
Dataset de nomes brasileiros em linguagem natural desenvolvido pela Comunidade Synapse BR.
O dataset contém exemplos de interações em português brasileiro envolvendo sugestões e geração de nomes, desenvolvido para aplicações de Processamento de Linguagem Natural (PLN) e treinamento de modelos de linguagem.
📊 Estrutura
Cada exemplo contém:
"prompt" — solicitação ou contexto fornecido ao modelo.
"completion" — resposta correspondente em português brasileiro.
📤… See the full description on the dataset page: https://huggingface.co/datasets/Comunidade-Synapse-BR/synapse-nomes-brasileiros.synapsellm-v0-2synapsellm-v0-1Synapse-ASCII_image_v01Prof-Synapse_SynthDataecu_constitution_cotSYNAPSE-Bench
SYNAPSE-Bench
SYNAPSE-Bench is a multimodal benchmark for evaluating cooperative planning,
action selection, and communication in embodied multi-agent tasks. Each example
describes one agent's local decision context in a household-like environment,
including visual observations, structured belief state, teammate model,
coordination dependencies, interaction history, a local scene graph, and the
reference next-step response.
The released data is formatted for vision-language and… See the full description on the dataset page: https://huggingface.co/datasets/randomusername784358/SYNAPSE-Bench.
