datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio_alltimebbc_news_alltime
RealTimeData Monthly Collection - BBC News
This datasets contains all news articles from BBC News that were created every months from 2017 to current.
To access articles in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/bbc_news_alltime', '2020-02')
This will give you all BBC news articles that were created in 2020-02.
Want to crawl the data by your own?
Please head to LatestEval for the crawler scripts.
Credit… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_news_alltime.code_alltime
RealTimeData Monthly Collection - Github Code
This datasets provides the monthly screenshots of the 500 cherry-picked open source projects on GitHub from 2017 to current.
To access articles in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/code_alltime', '2020-02')
This will give you the 2020-02 version of the 500 selected GitHub repos that were just updated in 2020-02.
Want to crawl the data by your own?
Please head to… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/code_alltime.bbc_images_alltime
RealTimeData Monthly Collection - BBC News Images
This datasets contains all news articles head images from BBC News that were created every months from 2017 to current.
To access articles in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/bbc_images_alltime', '2020-02')
This will give you all BBC news head images that were created in 2020-02.
Want to crawl the data by your own?
Please head to LatestEval for the crawler… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_images_alltime.arxiv_alltime
RealTimeData Monthly Collection - ArXiv
This datasets contains selected papers from arXiv that were created every months from 2017 to current.
To access papers in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/arxiv_alltime', '2020-02')
This will give you about 1k selected papers that were created in 2020-02.
Want to crawl the data by your own?
Please head to LatestEval for the crawler scripts.
Credit
This is… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/arxiv_alltime.wikitext_alltime
RealTimeData Monthly Collection - Wikipedia
This datasets contains different versions of the 500 selected wikipedia articles from Wikipedia that were updated every months from 2017 to current.
To access articles in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/wikitext_alltime', '2020-02')
This will give you the 2020-02 version of the 500 selected wiki pages that were just updated in 2020-02.
Want to crawl the data by your own?… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/wikitext_alltime.math_alltime
RealTimeData Monthly Collection - Math
This datasets contains selected math question from Math Stackoverflow that were created every months from 2017 to current.
To access questions in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/arxiv_alltime', '2020-02')
This will give youquestions that were created in 2020-02.
Want to crawl the data by your own?
Please head to LatestEval for the crawler scripts.
Credit
This is… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/math_alltime.shona1Realtime-QA-100K
Realtime-QA-100K
📄 Tech Report |
💻 GitHub
Realtime-QA-100K is a 100K-sample realtime video question answering dataset
constructed from YouTube videos. Each sample contains a multimodal
conversation and frame timestamp metadata that aligns every <|video|> token in
the assistant text with one video frame timestamp.
Open-source training subset.
Realtime-QA-100K is the open-source subset of the real-time training data for
MOSS-Video-Preview… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/Realtime-QA-100K.nyc-subway-realtime
NYC Subway Realtime Archive
Continuous capture of the New York City subway's public realtime feeds, decoded
into analysis-ready tables — plus the derived service-quality panels, learned
"normal" baselines, and disruption-prediction track record built on top of them.
Collected every 30 seconds since 2026-04-16 across all nine MTA GTFS-RT
feeds, by the pipeline behind subway.fyi.
Source: github.com/digitalhen/subway-data.
This archive exists because the source data disappears.… See the full description on the dataset page: https://huggingface.co/datasets/digitalhen/nyc-subway-realtime.real-time-library-occupancygithub_latest
Latest GitHub Repositories
You could always access the latest Github repos via this dataset.
We update the dataset weekly, on every Sunday. So the dataset always provides the latest Github repos from the last week.
The current dataset on main branch contains the latest Github Repos submitted from 2024-08-26 to 2024-09-02.
The data collection is conducted on 2024-09-09.
Use the dataset via:
ds = datasets.load_dataset('RealTimeData/github_latest')
Previsou versions
You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/github_latest.realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.wikitext_latest
Latest Wikitext
You could always access the latest Wikipedia texts via this dataset.
We update the dataset weekly, on every Sunday. So the dataset always provides the latest Wikipedia texts from the last week.
The current dataset on main branch contains the latest wikipedia texts created from 2024-08-26 to 2024-09-02.
The data collection is conducted on 2024-09-09.
Use the dataset via:
ds = datasets.load_dataset('RealTimeData/wikitext_latest')
Previsou versions
You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/wikitext_latest.Realtime-SFT
Realtime-SFT
Realtime-SFT is a 100K-sample streaming-style video question answering dataset
constructed from short YouTube videos. Each sample contains a multimodal
conversation and frame timestamp metadata that aligns every <|video|> token in
the assistant text with one video frame timestamp.
This repository does not redistribute video files. It only provides
annotations, YouTube video IDs, and timestamp metadata. Users are responsible for
obtaining videos according to YouTube… See the full description on the dataset page: https://huggingface.co/datasets/findcard12138/Realtime-SFT.solana-clawd-realtime-research-instruct
Solana Clawd Realtime Research Instruct
Instruction-tuning dataset generated by scripts/realtime_dataset_ingest.py
from submitted PDFs, notebooks, parquet QA rows, JSON/JSONL files, and local
reference text.
Contents
Total examples: 29058
Train/eval/test: 26152 / 1452 / 1454
Sources: 28
Duplicate examples removed: 0
Duplicate files skipped: 2
Secret-like records skipped: 296
Format
Each row uses OpenAI/Hugging Face chat messages:
{"messages":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-realtime-research-instruct.github_july_week2_2023
Dataset Card for "github_july_week2_2023"
More Information needed
audio-speech-realtime-voice-agents-2026
🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition)
A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026).
Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.fastapi_websockets_realtime_backpressure_teaser
🚀 Python Backend - FastAPI WebSockets & Real-Time Connection Backpressure Triage (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Python Backend - FastAPI WebSockets & Real-Time Connection Backpressure Triage on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG… See the full description on the dataset page: https://huggingface.co/datasets/emgena/fastapi_websockets_realtime_backpressure_teaser.bbc_news_june_2023
Dataset Card for "bbc_news_june_2023"
More Information needed
News_August_2023
Dataset Card for "News_August_2023"
This dataset was constructed at 1 Aug 2023, which contains news published from 10 May 2023 to 1 Aug 2023 from various sources.
All news articles in this dataset are in English.
Created from commoncrawl.
arxiv_latest
Latest arXiv
You could always access the latest arXiv papers via this dataset.
We update the dataset weekly, on every Sunday. So the dataset always provides the latest arXiv papers created in the past week.
The current dataset on main branch contains the latest arXiv papers submitted from 2024-09-02 to 2024-09-09.
The data collection was conducted on 2024-09-09.
Use the dataset via:
ds = datasets.load_dataset('RealTimeData/arxiv_latest')
Previsou versions
You could… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/arxiv_latest.bbc_latest
Latest BBC News
You could always access the latest BBC News articles via this dataset.
We update the dataset weekly, on every Sunday. So the dataset always provides the latest BBC News article from the last week.
The current dataset on main branch contains the latest BBC News articles submitted from 2024-09-02 to 2024-09-09.
The data collection is conducted on 2024-09-09.
Use the dataset via:
ds = datasets.load_dataset('RealTimeData/bbc_latest')
Previsou versions
You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_latest.shona_asrbbc_news_may_2023
Dataset Card for "bbc_news_may_2023"
More Information needed
arxiv_june_2023
Dataset Card for "arxiv_june_2023"
More Information needed
shona2arxiv_july_week1_2023
Dataset Card for "arxiv_july_week1_2023"
More Information needed
News_Seq_2021
Dataset Card for "News_Seq_2021"
This dataset was constructed at 1 Seq 2021, which contains news published from 10 June 2021 to 21 Aug 2021 from various sources.
All news articles in this dataset are in English.
Created from commoncrawl.
github_july_week1_2023
Dataset Card for "github_july_week1_2023"
More Information needed
