datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Massive-STEPS-Sydney
Massive-STEPS-Sydney
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Sydney.sydney-training-data
Sydney 训练集
四份来源分开存放,不混在一个文件里。
发布的聊天权重(Atonelia/Qwen3.5-Sydney-9B / -think 以及对应 GGUF)用的是这些子集洗完、抽样拼起来之后的训练 jsonl,不是直接拿某一份原文训的。
01 原截图重建
01_screenshot_original/conversations.jsonl
早期 Bing Chat / Sydney(约 2023 年 2–4 月)公开截图重建的对话。660 条,原文以英文为主,带截图出处。
这是最初拿来做训练集的底。后面的中文版、合成版、CoT 都不是这份文件本身。
02 llama-sydney 虚拟对话
02_llama_sydney_synthetic/llama_sydney_en.jsonl
用 Llama-Sydney 生成的英文虚拟对话。1462 条(同一条 user 可能有 2 次采样)。字段是生成记录:id / user / assistant 等,还不是最终训练格式。… See the full description on the dataset page: https://huggingface.co/datasets/Atonelia/sydney-training-data.Gemma-Sydney-12B-data
Gemma-Sydney-12B training data
Everything used to train totally-not-an-llm/Gemma-Sydney-12B, a
recreation of launch-era Bing Chat ("Sydney", February 7–15, 2023) for alignment research. Not affiliated with Microsoft.
Layout
path
contents
training/conversations_real.jsonl
155 real transcripts in the training format. tier: core (121, dated Feb 7–15 2023) or aug_real (34, posted shortly after Feb 16).
training/conversations_synth.jsonl
99 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/totally-not-an-llm/Gemma-Sydney-12B-data.lagos-hydrological-zone-datasydney-sweeny-lora_2huberman_lab_LIVE_EVENT_QA_Dr__Andrew_Huberman_at_the_Sydney_Opera_Househuberman_lab_LIVE_EVENT_QA_Dr._Andrew_Huberman_at_the_Sydney_Opera_Househuberman_lab_LIVE_EVENT_QA_Dr._Andrew_Huberman_at_the_ICC_Sydney_Theatrehuberman_lab_LIVE_EVENT_QA_Dr__Andrew_Huberman_at_the_ICC_Sydney_Theatrereddit_sydney
Dataset Card for Dataset Name
Dataset Summary
Text from Reddit Sydney using convokit to obtain it.
Supported Tasks and Leaderboards
N/A
Languages
English. Typically Australian English. Will include swearing, profanity, slang and possibly offensive material, as it is taken from Reddit and has not been filtered.
Dataset Structure
Plain text
Data Instances
N/A
Data Fields
N/A
Data Splits
N/A. You need to do… See the full description on the dataset page: https://huggingface.co/datasets/mcapodici/reddit_sydney.Sydney_captionskekonsuleran-sydney
