datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Swift-OpenX-Embodimentswift-reasoning-rollouts-deepscaler-ministral8b
DeepScaleR Reasoning Rollouts (Ministral-8B)
This dataset contains reasoning rollouts used to train the SWIFT reward head.
Paper page: https://huggingface.co/papers/2505.12225
GitHub: https://github.com/aster2024/SWIFT/
Generator model: mistralai/Ministral-8B-Instruct-2410 (https://huggingface.co/mistralai/Ministral-8B-Instruct-2410)
Dataset Description
This dataset contains 10000 samples corresponding to the Generalization Test setup.
Source: DeepScaleR.
Generator:… See the full description on the dataset page: https://huggingface.co/datasets/Aster2024/swift-reasoning-rollouts-deepscaler-ministral8b.dataVLM-swiftswift
Dataset Card for "swift"
More Information needed
image-pointing-1M-sft-swiftSwift-OpenX-Embodiment-wrist-imagesthe-stack-swift-clean
Dataset 1: TheStack - Swift - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Swift, a popular statically typed language.
Target Language: Swift
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Swift as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-swift-clean.mmu_swift_sne_ia
mmu_swift_sne_ia HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_swift_sne_ia.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_swift_sne_ia.swiftdataStack2Graph_KG_swift
Swift StackOverflow Knowledge Graph
Summary
This Hugging Face dataset repository contains the Swift shard of the Stack2Graph StackOverflow Knowledge Graph.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content.
Stack2Graph… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_swift.swift-17
SWIFT 17 — superseded by the n=26 census
SUPERSEDED. This 17-bank press census is superseded by live GET https://councilof.ai/api/swift (n=26 census: 3 LIVE · 9 COMMITTED · 14 DISCOVERED · n_measured=0). Schema notes supersedes: csoai.swift-17/0.1. Prefer csoai/gspc-swift-26 (census mirror) or the live API. Not clients. Not GPI. Not MEASURED.
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Historical 17 names from Swift… See the full description on the dataset page: https://huggingface.co/datasets/csoai/swift-17.gspc-swift-17
SWIFT 17 — superseded by the n=26 census
SUPERSEDED by GET https://councilof.ai/api/swift (n=26). Not clients. Not GPI. Not MEASURED (n_measured=0 on the live census). Prefer csoai/gspc-swift-26 or the live API.
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
17 DISCOVERED from Swift press 2026-07-09 was the prior tape. Live census is n=26 (THREE_STATE). This Hub page is a printer/alias, not a second engine. Not a… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swift-17.gspc-swift-50
SWIFT 50 — honesty shell, no public 50
Honest census: live GET https://councilof.ai/api/swift reports n=26 named banks (not 50). Swift cites 40+ in MVP construction; only 26 sourced to a real reachable dated press URL. Remaining ~14+ are not enumerated — no name invented to reach 50. Not clients. Not MEASURED.
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
This slug exists so inbound "swift-50" searches land on an honest… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swift-50.bigquery-swift-unfiltered
GitHub Swift Repositories
Dataset Description
Dataset Summary
This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license.
Source Data
Initial Data Collection and Normalization
The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.gspc-swift-26
SWIFT census n=26 — mirror of /api/swift
Census mirror of live GET https://councilof.ai/api/swift — n=26 (3 LIVE · 9 COMMITTED · 14 DISCOVERED · n_measured=0). Kind=reader · writes_board=false. Supersedes csoai.swift-17/0.1 / Hub csoai/gspc-swift-17 / csoai/swift-17.
Live SWIFT census: https://councilof.ai/api/swift
Live GSPC board: https://councilof.ai/api/gspc
Honest sourced census of 26 named banks. Not clients. Not GPI. Not MEASURED — no ISO 20022 / copybook / MT artifact… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swift-26.swiftswift-xrt-cleaned-events
Swift XRT cleaned events
This dataset contains the Swift XRT cleaned photon-counting-mode event file
sw00020000001xpcw4po_cl.evt.gz for public observation 00020000001, target
GRB041217. The source is
sw00020000001xpcw4po_cl.evt.gz,
checked 2026-08-30: 47,326 bytes, SHA-256
65c8bd98bd37a184ee584459a25c09e9c0af92b0cc8235eb5071d593a2d3e61c.
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/swift-xrt-cleaned-events.synthetic-swift-data-single-turn
Dataset Card for synthetic-swift-data-single-turn
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn.taylor_swift
Dataset Card for "taylor_swift"
More Information needed
Stack2Graph_VD_swift
Swift StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the Swift shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_swift.swiftsilkie-ms-swift
Dataset Card for "silkie-ms-swift"
More Information needed
iva-swift-codeint
IVA Swift GitHub Code Dataset
Dataset Description
This is the raw IVA Swift dataset extracted from GitHub.
It contains uncurated Swift files gathered with the purpose to train a code generation model.
The dataset consists of 753693 swift code files from GitHub totaling ~700MB of data.
The dataset was created from the public GitHub dataset on Google BiqQuery.
How to use it
To download the full dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint.voxconverse_swift
VoxConverse — Speaker Diarization in the Wild (MS-Swift Format)
This dataset is a reformatted version of VoxConverse for fine-tuning and evaluating multimodal large language models on speaker diarization, packaged in the MS-Swift Parquet format.
Note: The underlying audio is sourced from YouTube videos whose copyright remains with the original owners. This reformatted dataset is intended for research purposes only. For the original annotations and audio, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/AtwMaxime/voxconverse_swift.swiftplan-isaac-sim
SwiftPlan Isaac Sim Dataset
This dataset contains Isaac Sim observation images for frame-level high-level action selection in robotic task planning.
Each sample includes:
an RGB observation image,
a task instruction,
a frame-level high-level action label,
an action type,
an optional target object.
The dataset is designed for execution-time high-level decision making, where a model selects the next high-level action from the current observation and task instruction.… See the full description on the dataset page: https://huggingface.co/datasets/Kuoskyler/swiftplan-isaac-sim.LLaVA-Video-small-swift
Dataset Card LLaVA-Video-small-swift
Small subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.
ms-swiftLLaVA-Video-large-swift
Dataset Card LLaVA-Video-medium-swift
A subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.
SMBU-ECE-GRADE1-MATERIALiva-swift-codeint-clean-train
IVA Swift GitHub Code Dataset
Dataset Description
This is the curated train split of IVA Swift dataset extracted from GitHub.
It contains curated Swift files gathered with the purpose to train a code generation model.
The dataset consists of 320000 Swift code files from GitHub.
Here is the unsliced curated dataset and
here is the raw dataset.
How to use it
To download the full dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean-train.
