datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GO-MO
GO-MO: A large-scale graph-augmented traffic dataset for data-driven spatio-temporal traffic analysis
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/GO-MO.go-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.Selena-Gomez-With-Lyrics-And-Spotify-Audio-FeaturesGO_MF
GO-MF Dataset
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF.GO_MF_ESMFold
GO-MF Dataset with ESMFold Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_ESMFold.GO_MF_AlphaFold2
GO-MF Dataset with AlphaFold2 Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_AlphaFold2.gomodel-go-expert-v4
GoModel Go Expert v4 Dataset
Description
A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer
with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with:
Structured messages format (not pre-rendered ChatML text)
Go AST-extracted code from real repositories using go/parser
Go 1.26 feature coverage (February 2026 release)
Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.go-mf
Overview
Gene Ontology (GO) is a database of gene-level functional annotations. This specific dataset collects the molecular function and activities of specific genes. The GO exists as a heirarchy, and we subset to GO terms at most 3 levels away from the molecular function root.
This dataset is redistributed as part of mRNABench: https://github.com/morrislab/mRNABench
Data Format
Description of data columns:
target: Multihot label indicating GO terms applicable to… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/go-mf.gomodel-go-expert-v5gomodel-go-expert-v7e2e-nlg-chatmlgomodel-go-expert-v6gomoku_vlm_ds
Gomoku VLM Dataset (LoRA finetuning)
This repository contains a synthetic, image-grounded instruction dataset for training and evaluating vision-language models (VLMs) on Gomoku (15×15).The dataset is designed for LoRA finetuning of image-text-to-text vision-language models on two complementary capabilities:
VisualTasks where the model must read the board image and produce a structured answer about the current position.This includes purely perceptual objectives (cell classification… See the full description on the dataset page: https://huggingface.co/datasets/eganscha/gomoku_vlm_ds.gomodel-tool-reinforced
GoModel Tool-Reinforced Dataset
A tool-reinforced instruction dataset for training Go coding assistants. It combines Go code tasks with examples that teach a model when and how to use repository and Go development tools.
Tools
read_file
search_code
list_functions
get_imports
write_file
go_vet
go_build
Dataset statistics
Split
File
Examples
Bytes
train
train_tool_reinforced.jsonl
1,732
3,726,144
eval
eval_tool_reinforced.jsonl
193
408… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-tool-reinforced.FinetuningTestafrica-senegal-production-de-gombo-dbc0117c
Production De Gombo | Africa (DHORT)
1 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 1 rows from DHORT, covering Production De Gombo. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Official statistics datasets help analysts inspect public data as published by… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-senegal-production-de-gombo-dbc0117c.gomodel-go-expert-v3
GoModel Go Expert v3
Dataset description
GoModel Go Expert v3 is an English instruction and completion dataset for training
Go coding assistants. It combines curated production Go, code-specific synthetic
tasks, and agentic tool trajectories. Every JSONL record contains a full
Qwen2.5-compatible ChatML conversation in its text field.
Key changes from v2
Tool calls now use Qwen2.5's native <tool_call> tags instead of bare JSON.
Tool definitions use… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v3.spotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.
Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/Gomly/spotify_audio_features.africa-senegal-rendement-gombo-ab251afb
Rendement Gombo | Africa (DHORT)
1 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 1 rows from DHORT, covering Rendement Gombo. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Official statistics datasets help analysts inspect public data as published by… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-senegal-rendement-gombo-ab251afb.gomes_substationCyberBrain-QA-Datasetgomodel-go-expert-v2
GoModel Go Expert v2
Dataset description
GoModel Go Expert v2 is an English instruction and completion dataset for training
Go coding assistants. It combines curated production Go with synthetic instruction
tasks and agentic tool trajectories. Version 2 is a new dataset and does not replace
the earlier GoModel repositories. Each JSONL record is already serialized as a full
Qwen-compatible ChatML conversation in its text field.
Data sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v2.selena-gomezThis dataset is designed to generate lyrics with HuggingArtists.GOMP_USDhumanoid-gombal-datallmtwinfukui-sakai-harue-gomigvrn_dcmnragas-test-datasetdatalchemy-dataset
