datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marl-gpt-datasets
MARL-GPT Datasets
Offline expert trajectories from “MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning”.
Environments
This dataset includes trajectories from the three evaluation domains used in MARL-GPT: SMACv2 (StarCraft multi-agent combat), Google Research Football (GRF), and POGEMA (partially observable multi-agent pathfinding on grids).
Format
Trajectories are stored sequentially (no shuffling). Use the done flag to split the stream into… See the full description on the dataset page: https://huggingface.co/datasets/nortem/marl-gpt-datasets.og-marl
@misc{formanek2024puttingdatacentreoffline,
title={Putting Data at the Centre of Offline Multi-Agent Reinforcement Learning},
author={Claude Formanek and Louise Beyers and Callum Rhys Tilbury and Jonathan P. Shock and Arnu Pretorius},
year={2024},
eprint={2409.12001},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2409.12001},
}
marlin-assetsmarl_checkpointopenwebtext-gemma-2-context-128
OpenWebTextCorpus tokenized for Gemma 2 with 128 context size
This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset
using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset.
This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it).
This dataset was created using SAELens, with the following settings:… See the full description on the dataset page: https://huggingface.co/datasets/Marlon154/openwebtext-gemma-2-context-128.marlinai2_arc-pt
marlosb/ai2-arc-pt
This dataset is a Portuguese translation of the original AI2 ARC (AI2 Reasoning Challenge) dataset released by AllenAI.
Original Dataset
Hugging Face: allenai/ai2_arc
Homepage: https://allenai.org/data/arc
Paper: Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Dataset Summary
AI2 ARC consists of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/ai2_arc-pt.sqp-tts-en
SQP TTS (English)
Synthesized speech for SQPsychConv_qwen-2.5, a synthetic CBT therapist-client
dialogue dataset (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/sqp-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/sqp-tts-en.mmlu-pt
marlosb/mmlu-pt
This dataset is a Portuguese translation of the original MMLU (Measuring Massive Multitask Language Understanding) dataset.
Original Dataset
Hugging Face: cais/mmlu
Repository: https://github.com/hendrycks/test
Paper: Measuring Massive Multitask Language Understanding
Dataset Summary
MMLU contains ~15,000 multiple-choice questions across 57 subjects (elementary mathematics, history, computer science, law, medicine, etc.), designed to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/mmlu-pt.BARE_grefumd_test_partitioning
RefCOCOg-UMD Test Partitioning
This dataset provides fine-grained evaluation splits for the RefCOCOg-UMD test set, designed for Referring Expression Comprehension (REC) tasks. Each sample is partitioned by difficulty level and referential type to enable more detailed performance analysis.
Data Structure
Each .pth file contains a list of tuples with the following structure:
import torch
data = torch.load("Easy.pth")
for item in data:
img_name, hw_dict, bbox, phrase… See the full description on the dataset page: https://huggingface.co/datasets/marloweee/BARE_grefumd_test_partitioning.lunghist700MARL-GPT-LMAPFsmol-smoltalk-pt
marlosb/smol-smoltalk-pt
This dataset is a Portuguese translation of the original Smol-SmolTalk dataset from HuggingFaceTB.
Original Dataset
Hugging Face: HuggingFaceTB/smol-smoltalk
Dataset Summary
This is a subset of SmolTalk dataset adapted for smol models with less than 1B parameters.
Compared to SmolTalk:
The conversations from Smol-Magpie-Ultra are shorter in this dataset
We include less task-specific data compared to SmolTalk (e.g., no function… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/smol-smoltalk-pt.stitch-s
stitch-s
Dataset splits:
train: original stitch-s data.
train_with_reference_ans: rows from train where input.reference_answer is not empty.
train_with_tool: rows from train_with_reference_ans where msg contains both <TOOL_CALL> and <TOOL_RESULT>.
moral-number-corpus
A Perspectivist Corpus of Numbers in Social Judgements
This is the dataset for A Perspectivist Corpus of Numbers in Social Judgements.
We constructed a corpus of moral and social judgements (questions are derived from the Commonsense Norm Bank) that asks people to fill in number ranges that do not change a given judgement.
Our corpus was crowdsourced from 30 annotators and contains 898 statements for a total of 3k annotations.
This work adds to available moral and social judgement… See the full description on the dataset page: https://huggingface.co/datasets/Marlon154/moral-number-corpus.GSM8K-pt
marlosb/gsm8k-pt
This dataset is a Portuguese translation of the original GSM8K (Grade School Math 8K) dataset released by OpenAI.
Original Dataset
Hugging Face: openai/gsm8k
GitHub: openai/grade-school-math
Paper: Training Verifiers to Solve Math Word Problems
Dataset Summary
GSM8K consists of 8.5K high-quality, linguistically diverse grade school math word problems created to support research on multi-step mathematical reasoning.Problems require 2–8 steps… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/GSM8K-pt.Preference-Conditioned-Heterogeneous-MARL-Microgrid
SEGAN OPSD-Derived Microgrid Multiyear Benchmark
This repository contains the processed multiyear microgrid benchmark used for the study
“Preference-Conditioned Heterogeneous Multi-Agent Reinforcement Learning for Safe Microgrid Energy Management.”
Files
microgrid_opsd_multiyear.csv — processed hourly benchmark data.
opsd_multiyear_metadata.json — provenance, selected OPSD nodes, source-column mapping, scaling notes, and processing metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Tristanchou/Preference-Conditioned-Heterogeneous-MARL-Microgrid.Imperativoswe-gym-marlmarlinAmazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset.
This dataset mainly includes reviews (ratings, text) and item metadata (desc-
riptions, category information, price, brand, and images). Compared to the pre-
vious versions, the 2023 version features larger size, newer reviews (up to Sep
2023), richer and cleaner meta data, and finer-grained timestamps (from day to
milli-second).marl-world-model-lerobotSTAR_MARLstarai02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "starai",
"total_episodes": 101,
"total_frames": 44639,
"total_tasks": 1,
"total_videos": 202,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Marlboro1998/starai02.stitch-s_2annomi-tts-en
AnnoMI TTS (English)
Synthesized speech for the AnnoMI motivational interviewing dialogues (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/annomi-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/annomi-tts-en.SegLLM_datasetdr-marl-papers
Distributionally-Robust RL & Cooperative MARL — top-tier paper index
A hand-reviewed index of 1,989 papers on distributionally-robust reinforcement learning
and cooperative multi-agent RL, drawn from a complete harvest of 75,819 accepted papers
at ICML, NeurIPS, ICLR, AAMAS and AISTATS (2010–2026).
Every candidate that survived the keyword filter was read and labelled one by one — 2,979
papers — rather than accepting an automatic classifier's output.
Buckets… See the full description on the dataset page: https://huggingface.co/datasets/Ngseo/dr-marl-papers.luculia_marlborough_violetevergarden
Dataset of Luculia Marlborough/ルクリア・モールバラ (Violet Evergarden)
This is the dataset of Luculia Marlborough/ルクリア・モールバラ (Violet Evergarden), containing 106 images and their tags.
The core tags of this character are long_hair, green_eyes, red_hair, freckles, braid, brown_hair, crown_braid, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/luculia_marlborough_violetevergarden.LC25000
