datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sea-commoncrawlsailormoon1990s
Bangumi Image Base of Sailor Moon (1990s)
This is the image base of bangumi Sailor Moon (1990s), we detected 132 characters, 14684 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sailormoon1990s.sea-syntheticsailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B.
sea-commoncrawl-high-qualitysea-pdf-textsea-internetaustin_sailor_dataset_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 240,
"total_frames": 353094,
"total_tasks": 4,
"total_videos": 480,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:240"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/austin_sailor_dataset_lerobot.austin_sailor_dataset_rawsailormoon2010s
Bangumi Image Base of Sailor Moon (2010s)
This is the image base of bangumi Sailor Moon (2010s), we detected 46 characters, 3463 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sailormoon2010s.sailor2-pretrain-data-stage2The pre-training dataset (stage2) for the Sailor2 models, including 1B, 8B and 20B.
austin_sailor_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 240,
"total_frames": 353094,
"total_tasks": 4,
"total_videos": 480,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:240"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/austin_sailor_dataset.community-datasetsea-ultrafeedbacksailor2-sft-stage1xcopa
XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
The Cross-lingual Choice of Plausible Alternatives dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across languages. The dataset is the translation and reannotation of the English COPA (Roemmele et al. 2011) and covers 11 languages from 11 families and several areas around the globe. The dataset is challenging as it requires both the command of world knowledge… See the full description on the dataset page: https://huggingface.co/datasets/sailor2/xcopa.isam_synthversedetails_sail__Sailor-0.5B
Dataset Card for Evaluation run of sail/Sailor-0.5B
Dataset automatically created during the evaluation run of model sail/Sailor-0.5B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sail__Sailor-0.5B.austin_sailor_dataset_augmented
austin_sailor_dataset_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7
FPS: 20
Episodes: 240
Frames: 353,094
Splits:
train: 0:240
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/austin_sailor_dataset_augmented.Vietnamese_RAG
Dataset Card for Dataset Name
Vi's RAG is an comprehensive Vietnamese dataset optimized for RAG Evaluation, build by ZD AI lab and release under Apache license 2.0.
Dataset Details
There are four datasets in this card :
Vietnamese version of Expert QA that we utilize the strong translation ability of GPT-4 for translation task
RAG ViQuAD which was carefully chosen from UIT-ViQuAD2.0 with additional context column filtered by title
Legal RAG and BKAI_RAG are long form RAG… See the full description on the dataset page: https://huggingface.co/datasets/sailor2/Vietnamese_RAG.details_sail__Sailor-0.5B-Chat
Dataset Card for Evaluation run of sail/Sailor-0.5B-Chat
Dataset automatically created during the evaluation run of model sail/Sailor-0.5B-Chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sail__Sailor-0.5B-Chat.drifting-vla-v2-austin_sailorsailorv2details_sail__Sailor-4B
Dataset Card for Evaluation run of sail/Sailor-4B
Dataset automatically created during the evaluation run of model sail/Sailor-4B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sail__Sailor-4B.details_sail__Sailor-1.8B-Chat
Dataset Card for Evaluation run of sail/Sailor-1.8B-Chat
Dataset automatically created during the evaluation run of model sail/Sailor-1.8B-Chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sail__Sailor-1.8B-Chat.thinking_austin_sailor_dataset_lerobot_output_qwen3vldetails_sail__Sailor-7B
Dataset Card for Evaluation run of sail/Sailor-7B
Dataset automatically created during the evaluation run of model sail/Sailor-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sail__Sailor-7B.SAILORThis repository contains the data presented in A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search.
Code: https://github.com/arnavkj1995/SAILOR
literotica-8k-story-collection
Literotica 8K Story Collection
This dataset contains 8,000 highly curated and filtered stories from the erotic literature domain. It has been cleaned and formatted for high-quality language model fine-tuning.
Used in the creation of: Sinbad-The-Sailor/Qwen3.5-4B-NSFW-ARA-Heretic-Literotica
details_sail__Sailor-4B-Chat
Dataset Card for Evaluation run of sail/Sailor-4B-Chat
Dataset automatically created during the evaluation run of model sail/Sailor-4B-Chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sail__Sailor-4B-Chat.
