datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toronto-bikeshareNetwork_Defense_Symmetric_Competitive102,400,000 timesteps, Multi-Agent Reinforcement Learning
Total Environment Steps= 10 parallel environments × 7,000 episodes ×2,048 steps= 102400000 Training Timesteps
-The Red Agent’s goal is to discover vulnerabilities, elevate privileges, compromise assets, and maintain persistence. Its action space can be modeled after phases of the
MITRE ATT&CK framework.
-The Blue Agent’s goal is to maintain system availability, reduce the attack surface, detect malicious… See the full description on the dataset page: https://huggingface.co/datasets/TorontoMetropolitanUniversity/Network_Defense_Symmetric_Competitive.toronto-tv-ukrainian
Toronto TV Ukrainian Speech Dataset
For educational purposes only.
All rights to the original video and audio content belong to Телебачення Торонто (YouTube channel).
Dataset Summary
A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip and its corresponding Ukrainian subtitle text, intended for use in automatic speech recognition (ASR) research and education.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/toronto-tv-ukrainian.toronto-slang-lexicon
Dataset Card for Toronto Slang / Multicultural Toronto English (MTE) Lexicon
Dataset Summary
A machine-readable lexicon of Toronto slang and Multicultural Toronto English (MTE) vocabulary: 287 entries (78 core), each with a headword, spelling variants, part of speech, short and long glosses, a pragmatic slot (how the word functions in an utterance), register, a Toronto-versus-London status, origin, an offensiveness rating, a core/peripheral flag, grammar and… See the full description on the dataset page: https://huggingface.co/datasets/devon7y/toronto-slang-lexicon.airbnb-ca-on-torontoopendata_toronto_building_permitsdeepscaler-verl-preprocessedssft1Kcode-v2-Sonnet4-5-High-Run1-temp1-max29000
Dataset card for ssft1Kcode-v2-Sonnet4-5-High-Run1-temp1-max29000
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"question": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition. Enclose your code within delimiters. The Chef likes to stay in touch with his staff. So, the Chef, the head server, and the… See the full description on the dataset page: https://huggingface.co/datasets/shengjia-toronto/ssft1Kcode-v2-Sonnet4-5-High-Run1-temp1-max29000.ssft1Kcode-v2-Sonnet4-5-High-Run2-temp1-max29000
Dataset card for ssft1Kcode-v2-Sonnet4-5-High-Run2-temp1-max29000
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"question": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition. Enclose your code within delimiters. The Chef likes to stay in touch with his staff. So, the Chef, the head server, and the… See the full description on the dataset page: https://huggingface.co/datasets/shengjia-toronto/ssft1Kcode-v2-Sonnet4-5-High-Run2-temp1-max29000.exploded-s1k-4mixed-reasoning-rawtoronto-city-council-transcriptsdeepscaler-verlaime24-verl-preprocessedssft1Kcode-v2-OSS-High-Run1-temp1-max32768
Dataset card for ssft1Kcode-v2-OSS-High-Run1-temp1-max32768
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"question": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition. Enclose your code within delimiters. The Chef likes to stay in touch with his staff. So, the Chef, the head server, and the… See the full description on the dataset page: https://huggingface.co/datasets/shengjia-toronto/ssft1Kcode-v2-OSS-High-Run1-temp1-max32768.ssft1Kcode_v2huberman_lab_LIVE_EVENT_QA_Dr._Andrew_Huberman_at_Meridian_Hall_in_Torontohuberman_lab_LIVE_EVENT_QA_Dr__Andrew_Huberman_at_Meridian_Hall_in_Torontodapo-math-17k-promptprocessedNYC_Toronto_ControlNetlatent-toronto-emotionThe Toronto Emotional Speech Set (TESS) is a dataset consisting of emotionally charged speech recordings, designed for emotion recognition tasks. This repository provides precomputed audio embeddings extracted using the Music2Latent model. These embeddings are derived from the TESS dataset, available at TESS dataset on Kaggle, enabling quick and efficient use for tasks like speech emotion recognition. The embeddings can be directly used for classification tasks, without the need for raw audio… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/latent-toronto-emotion.exploded-s1k-4mixed-reasoning-raw-tokenizeds4Kcode-R1-tokenizeds1k-4mixed-reasoning-N6-tokenizedtorontotoronto-housing-QAToronto_ControlNet_DatasetToronto_Collision_silver
