datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playLongTimeScopeIf you use TimeScope please cite the following:
@misc{zohar2025apollo2,
title = {Apollo2: Exploring the Long-Video Frontier of Large Multimodal Models},
author = {Zohar, Orr and Wang, Xiaohan and Li, Rui and Marafioti, Andrés and Farré, Miquel and Noyan, Merve and von Werra, Leandro and Yeung-Levy, Serena and Wolf, Thomas},
year = {2025},
}
TimeScopeIf you use TimeScope please cite the following:
@misc{zohar2025apollo2,
title = {Apollo2: Exploring the Long-Video Frontier of Large Multimodal Models},
author = {Zohar, Orr and Wang, Xiaohan and Li, Rui and Marafioti, Andrés and Farré, Miquel and Noyan, Merve and von Werra, Leandro and Yeung-Levy, Serena and Wolf, Thomas},
year = {2025},
}
Apollo-VL-Massive-Dataset
🧬 Apollo-VL-Massive-Dataset
Apollo-VL-Massive-Dataset is a 161,562-row multimodal dataset engineered to train Vision-Language Models (VLMs) to perform structured visual reasoning using <think> tags before producing a final answer.
Designed by Pluto AI Labs, the dataset is optimized for knowledge distillation, behavioral cloning, visual reasoning, OCR, chart understanding, and multimodal instruction tuning of compact VLMs.
📊 Dataset Overview
Property… See the full description on the dataset page: https://huggingface.co/datasets/Pluto-AI-Labs/Apollo-VL-Massive-Dataset.ApolloCorpus-ja-askllm-v1
ApolloCorpus-ja-askllm-v1
データセット kunishou/ApolloCorpus-ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/ApolloCorpus-ja-askllm-v1.ReOpus-ApolloBooks-EN-NL-1M
ReOpus-ApolloBooks-1M
A high-quality English-Dutch (EN-NL) parallel corpus containing 1 million sentence pairs, constructed through strategic sampling and neural retranslation.
Overview
ReOpus-ApolloBooks-1M is a parallel translation corpus designed for training and evaluating English-Dutch machine translation systems. The corpus combines carefully sampled data from OPUS with neural retranslation using Qwen models, augmented with the Apollo Books parallel corpus.… See the full description on the dataset page: https://huggingface.co/datasets/OpenOranje/ReOpus-ApolloBooks-EN-NL-1M.apollo-11-diarized
Apollo 11 Mission Audio — Diarized Transcripts
Machine-generated transcripts with speaker diarization and timestamps for
103 tapes (175 hours) of Apollo 11 mission audio from the
Internet Archive's Apollo11Audio
collection (NASA recordings, public domain).
Generated in a single Hugging Face Job with
OpenMOSS-Team/MOSS-Transcribe-Diarize
(0.9B, Apache 2.0) — joint transcription + speaker attribution + timestamps in one
generation pass per clip.
Configs
segments… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/apollo-11-diarized.TimeScope-v0emgena_graphql_apollo_dataloader_mcp_teaser
🚀 API Architecture - GraphQL & Apollo Federation N+1 DataLoader Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
N+1 query resolver elimination, DataLoader batching deadlock resolution, and subgraph schema conflict triage… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_graphql_apollo_dataloader_mcp_teaser.apollo-preview-v0.3
apollo-preview-v0.2
Apollo is a RP/ERP dataset, which has been generated by amalgamating several data sources with roleplay, creative writing, and general instruction following data. This version is a preview version, with future versions planned to include our own data generation pipelines.
The data sources were generated using Claude 3 Opus, Claude 3.5 Sonnet, GPT-4o, GPT-4, and other open-sourced language models.
Data sources
This dataset was generated from the… See the full description on the dataset page: https://huggingface.co/datasets/QuasarResearch/apollo-preview-v0.3.apollo-preview-v0.4
apollo-preview-v0.4
Apollo is a RP/ERP dataset, which has been generated by amalgamating several data sources with roleplay, creative writing, and general instruction following data. This version is a preview version, with future versions planned to include our own data generation pipelines.
The data sources were generated using Claude 3 Opus, Claude 3.5 Sonnet, GPT-4o, GPT-4, and other open-sourced language models.
Data sources
This dataset was generated from the… See the full description on the dataset page: https://huggingface.co/datasets/QuasarResearch/apollo-preview-v0.4.apollo-llama3.3apollo-telephony-en
Apollo telephony — English (processed)
English Apollo air-to-ground / mission-control style telephony speech from the
Apollo FSC P4 ASR track-2 release (WAV + transcriptions.txt in apollo_data.zip).
Credits & license
Original corpus: Apollo FSC P4 ASR track 2.License: CC BY-4.0 (commercial use allowed with attribution).
Snapshot
Language: en
Samples: 48068 clips (~38.39 h)
Audio: OGG/Opus mono, native 8 kHz telephony
Construction: utterance-level… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/apollo-telephony-en.ApolloAuto-zyx-apollo
Dataset Card for "ApolloAuto-zyx-apollo"
More Information needed
apollo-preview-v0.2
apollo-preview-v0.2
Apollo is a RP/ERP dataset, which has been generated by amalgamating several data sources with roleplay, creative writing, and general instruction following data. This version is a preview version, with future versions planned to include our own data generation pipelines.
The data sources were generated using Claude 3 Opus, Claude 3.5 Sonnet, GPT-4o, GPT-4, and other open-sourced language models.
Data sources
This dataset was generated from the… See the full description on the dataset page: https://huggingface.co/datasets/QuasarResearch/apollo-preview-v0.2.apollo-llama3.3-insider-trading-generationsApolloMoE_Dataset_ThaiMedical LLMs For Much More Languages
ApolloMoEDataset-koreanfrom datasets import load_dataset
dataset = load_dataset("FreedomIntelligence/ApolloMoEDataset", data_files="ApolloMoEDataset.json", split="train")
filtered_dataset = dataset.filter(lambda ex: ex['type']!="general" and ex['language']=="ko")
At first glance, its quality is not bad.
apollo-llama3.3-ai-audit-a1-2-reasoningapollo-llama3.3-ai-audit-reasoningwiki_apollo_frac_20apollo-llama3.3-ai-audit-a1-2apollo-llama3.3-ai-liar-original-without-answersapollo-llama3.3-sandbagging-v2-wmdp-mmluphoebus-3.1-instructionsapollo1
Dataset Card for "apollo1"
More Information needed
apollo-llama3.3-alpaca-plainwiki_apollo_datasetfinepdfs_eng_Latn_labeledApolloRP-2.0-SFTreadme coming soon.
