bts
Datasets
All datasets matching “bts”mmu_btsbot
mmu_btsbot HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_btsbot.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_btsbot.bts-flights-weather
US flights joined to observed airport weather (2019 to 2025)
Every US domestic flight joined to the weather actually observed at both
airports at the right UTC hour, from January 2019 to July 2025, with a leakage
safe schema, aviation flight categories, and a documented airport to station
crosswalk, in Parquet.
Coverage
January 2019 to July 2025. This is a complete historical panel with a fixed end
date, not a live feed, and it does not track current conditions.… See the full description on the dataset page: https://huggingface.co/datasets/Arimancy/bts-flights-weather.common-voice-26-mn
Common Voice 26.0 Mongolian (cleaned)
A quality-filtered, normalised subset of Mozilla Common Voice Corpus 26.0, Mongolian, prepared for training
Mongolian (Khalkha Cyrillic) text-to-speech with
oron-tts.
Built by oron-cleaner. Every threshold
was calibrated on this corpus, and every number and column on this page is read
from the shipped data rather than asserted.
from datasets import load_dataset
ds = load_dataset("btsee/common-voice-26-mn", split="train")
print(ds[0]["text"]… See the full description on the dataset page: https://huggingface.co/datasets/btsee/common-voice-26-mn.btsmbspeech-mn
MBSpeech Mongolian (cleaned)
A quality-filtered, normalised subset of MBSpeech Mongolian (Bible read-speech), prepared for training
Mongolian (Khalkha Cyrillic) text-to-speech with
oron-tts.
Built by oron-cleaner. Every threshold
was calibrated on this corpus, and every number and column on this page is read
from the shipped data rather than asserted.
from datasets import load_dataset
ds = load_dataset("btsee/mbspeech-mn-clean", split="train")
print(ds[0]["text"]… See the full description on the dataset page: https://huggingface.co/datasets/btsee/mbspeech-mn.WorldSpeech-mn
WorldSpeech Mongolian (cleaned)
A quality-filtered, normalised subset of disco-eth/WorldSpeech, config mn_mn, prepared for training
Mongolian (Khalkha Cyrillic) text-to-speech with
oron-tts.
Built by oron-cleaner. Every threshold
was calibrated on this corpus, and every number and column on this page is read
from the shipped data rather than asserted.
from datasets import load_dataset
ds = load_dataset("btsee/WorldSpeech-mn", split="train")
print(ds[0]["text"]… See the full description on the dataset page: https://huggingface.co/datasets/btsee/WorldSpeech-mn.
