datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GHIST-Plus-bundle
GHIST+ data and model bundle
This bundle contains the model artifacts, predictions, evaluation inputs,
comparison outputs, and plot-ready tables released with
GHIST+. Source code is provided in
that repository.
Download
hf download SydneyBioX/GHIST-Plus-bundle \
--repo-type dataset \
--local-dir bundle
The bundle is approximately 38 GB. Individual files can also be downloaded
from this page.
Use with the figure notebooks
These notebooks… See the full description on the dataset page: https://huggingface.co/datasets/SydneyBioX/GHIST-Plus-bundle.details_FPHam__Sydney_Overthinker_13b_HF
Dataset Card for Evaluation run of FPHam/Sydney_Overthinker_13b_HF
Dataset Summary
Dataset automatically created during the evaluation run of model FPHam/Sydney_Overthinker_13b_HF on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_FPHam__Sydney_Overthinker_13b_HF.Massive-STEPS-Sydney
Massive-STEPS-Sydney
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Sydney.details_FPHam__Free_Sydney_13b_HF
Dataset Card for Evaluation run of FPHam/Free_Sydney_13b_HF
Dataset Summary
Dataset automatically created during the evaluation run of model FPHam/Free_Sydney_13b_HF on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_FPHam__Free_Sydney_13b_HF.Greater-Sydney-Trees-BuildingsExplore our Greater Sydney Building Footprint and Tree Patch datasets using a PMTiles map here
aigis
AIGIS
AI annotation, segmentation, and conversion tools for GIS imagery
aigis is a comprehensive toolkit for aerial and satellite imagery acquisition, processing, annotation, and analysis using artificial intelligence. This repository contains three main components:
annotate: Tools for annotating aerial imagery data.
convert: Utilities for converting… See the full description on the dataset page: https://huggingface.co/datasets/SIH/Greater-Sydney-Trees-Buildings.sydney-training-data
Sydney 训练集
四份来源分开存放,不混在一个文件里。
发布的聊天权重(Atonelia/Qwen3.5-Sydney-9B / -think 以及对应 GGUF)用的是这些子集洗完、抽样拼起来之后的训练 jsonl,不是直接拿某一份原文训的。
01 原截图重建
01_screenshot_original/conversations.jsonl
早期 Bing Chat / Sydney(约 2023 年 2–4 月)公开截图重建的对话。660 条,原文以英文为主,带截图出处。
这是最初拿来做训练集的底。后面的中文版、合成版、CoT 都不是这份文件本身。
02 llama-sydney 虚拟对话
02_llama_sydney_synthetic/llama_sydney_en.jsonl
用 Llama-Sydney 生成的英文虚拟对话。1462 条(同一条 user 可能有 2 次采样)。字段是生成记录:id / user / assistant 等,还不是最终训练格式。… See the full description on the dataset page: https://huggingface.co/datasets/Atonelia/sydney-training-data.Sydney-CaptionsGemma-Sydney-12B-data
Gemma-Sydney-12B training data
Everything used to train totally-not-an-llm/Gemma-Sydney-12B, a
recreation of launch-era Bing Chat ("Sydney", February 7–15, 2023) for alignment research. Not affiliated with Microsoft.
Layout
path
contents
training/conversations_real.jsonl
155 real transcripts in the training format. tier: core (121, dated Feb 7–15 2023) or aug_real (34, posted shortly after Feb 16).
training/conversations_synth.jsonl
99 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/totally-not-an-llm/Gemma-Sydney-12B-data.cs224n-nlp-summarization
NLP Summarization Audio Video Data Notes
Dataset summary
A documented NLP Summarization data-preparation workflow for Audio Video records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the… See the full description on the dataset page: https://huggingface.co/datasets/sydneydvig/cs224n-nlp-summarization.Sydney-Sweeney-FLUX-LoRaweather-corpus
Weather Pointcloud Text Data Notes
Dataset summary
A documented Weather data-preparation workflow for Pointcloud Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/sydneyjone/weather-corpus.lagos-hydrological-zone-dataIndian_Medicinal_Plants_and_Leaf_Dataset_Large
Project Documentation: Identification of Medicinal Plants/Raw Materials through Image Processing Using Machine Learning
Ministry of AYUSH - Student Innovation Hackathon (SIH1343)
Project Overview:
Our project aims to address the challenge of accurately identifying medicinal plant species and raw materials through the use of a web application. Leveraging image processing and machine learning algorithms, the application allows users to upload or capture images of… See the full description on the dataset page: https://huggingface.co/datasets/sydney31/Indian_Medicinal_Plants_and_Leaf_Dataset_Large.sydney-sweeny-lora_2Sydney_LLaVA_0610huberman_lab_LIVE_EVENT_QA_Dr__Andrew_Huberman_at_the_Sydney_Opera_Househuberman_lab_LIVE_EVENT_QA_Dr._Andrew_Huberman_at_the_Sydney_Opera_Housedataset_077279393_education_audio_text
dataset_077279393_education_audio_text.py
Dataset Summary
A education dataset with audio text modality, stored in webdataset format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: randaugment
Splits & Sampling
Split strategy: temporal
Sampling: contrastive
Quality & Labeling
Quality filtering: adaptive
Labeling: self training
Files
dataset_077279393_education_audio_text.py —… See the full description on the dataset page: https://huggingface.co/datasets/Sydneyclark/dataset_077279393_education_audio_text.huberman_lab_LIVE_EVENT_QA_Dr._Andrew_Huberman_at_the_ICC_Sydney_Theatrehuberman_lab_LIVE_EVENT_QA_Dr__Andrew_Huberman_at_the_ICC_Sydney_Theatreimnet1k_silky_terrier_Sydney_silkyreddit_sydney
Dataset Card for Dataset Name
Dataset Summary
Text from Reddit Sydney using convokit to obtain it.
Supported Tasks and Leaderboards
N/A
Languages
English. Typically Australian English. Will include swearing, profanity, slang and possibly offensive material, as it is taken from Reddit and has not been filtered.
Dataset Structure
Plain text
Data Instances
N/A
Data Fields
N/A
Data Splits
N/A. You need to do… See the full description on the dataset page: https://huggingface.co/datasets/mcapodici/reddit_sydney.Alejandrosydneyrl-sydney-rolloutsSydney_LLaVA_1210This time only required images are in the zip and I hope errors in the dataset are gone too.
Not sure how to make it work with dataset viewer though.
Sydney_captionsSydney_CaptionSydneyGhost
