CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RealTimeData /audio_alltimetextn<1K0 likes2k downloads3y agoHugging Face02RealTimeData /bbc_news_alltime RealTimeData Monthly Collection - BBC News This datasets contains all news articles from BBC News that were created every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/bbc_news_alltime', '2020-02') This will give you all BBC news articles that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler scripts. Credit… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_news_alltime.image100K<n<1M51 likes2k downloads1y agoHugging Face03RealTimeData /code_alltime RealTimeData Monthly Collection - Github Code This datasets provides the monthly screenshots of the 500 cherry-picked open source projects on GitHub from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/code_alltime', '2020-02') This will give you the 2020-02 version of the 500 selected GitHub repos that were just updated in 2020-02. Want to crawl the data by your own? Please head to… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/code_alltime.text10K<n<100K2 likes1.9k downloads1y agoHugging Face04RealTimeData /bbc_images_alltime RealTimeData Monthly Collection - BBC News Images This datasets contains all news articles head images from BBC News that were created every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/bbc_images_alltime', '2020-02') This will give you all BBC news head images that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_images_alltime.image100K<n<1M2 likes1.7k downloads1y agoHugging Face05RealTimeData /arxiv_alltime RealTimeData Monthly Collection - ArXiv This datasets contains selected papers from arXiv that were created every months from 2017 to current. To access papers in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/arxiv_alltime', '2020-02') This will give you about 1k selected papers that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler scripts. Credit This is… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/arxiv_alltime.text10K<n<100K10 likes1.3k downloads1y agoHugging Face06RealTimeData /wikitext_alltime RealTimeData Monthly Collection - Wikipedia This datasets contains different versions of the 500 selected wikipedia articles from Wikipedia that were updated every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/wikitext_alltime', '2020-02') This will give you the 2020-02 version of the 500 selected wiki pages that were just updated in 2020-02. Want to crawl the data by your own?… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/wikitext_alltime.text10K<n<100K3 likes1.1k downloads1y agoHugging Face07RealTimeData /math_alltime RealTimeData Monthly Collection - Math This datasets contains selected math question from Math Stackoverflow that were created every months from 2017 to current. To access questions in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/arxiv_alltime', '2020-02') This will give youquestions that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler scripts. Credit This is… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/math_alltime.tabular10K<n<100K3 likes1k downloads1y agoHugging Face08realtime-speech /shona1audio10K<n<100K1 likes407 downloads2y agoHugging Face09OpenMOSS-Team /Realtime-QA-100K Realtime-QA-100K 📄 Tech Report &nbsp;|&nbsp; 💻 GitHub &nbsp; Realtime-QA-100K is a 100K-sample realtime video question answering dataset constructed from YouTube videos. Each sample contains a multimodal conversation and frame timestamp metadata that aligns every <|video|> token in the assistant text with one video frame timestamp. Open-source training subset. Realtime-QA-100K is the open-source subset of the real-time training data for MOSS-Video-Preview… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/Realtime-QA-100K.textvisual-question-answering100K<n<1M8 likes219 downloads4mo agoHugging Face10digitalhen /nyc-subway-realtimegated NYC Subway Realtime Archive Continuous capture of the New York City subway's public realtime feeds, decoded into analysis-ready tables — plus the derived service-quality panels, learned "normal" baselines, and disruption-prediction track record built on top of them. Collected every 30 seconds since 2026-04-16 across all nine MTA GTFS-RT feeds, by the pipeline behind subway.fyi. Source: github.com/digitalhen/subway-data. This archive exists because the source data disappears.… See the full description on the dataset page: https://huggingface.co/datasets/digitalhen/nyc-subway-realtime.tabular1B<n<10B1 likes117 downloads12d agoHugging Face11davnas /real-time-library-occupancytabular10K<n<100K0 likes93 downloads1y agoHugging Face12RealTimeData /github_latest Latest GitHub Repositories You could always access the latest Github repos via this dataset. We update the dataset weekly, on every Sunday. So the dataset always provides the latest Github repos from the last week. The current dataset on main branch contains the latest Github Repos submitted from 2024-08-26 to 2024-09-02. The data collection is conducted on 2024-09-09. Use the dataset via: ds = datasets.load_dataset('RealTimeData/github_latest') Previsou versions You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/github_latest.tabularn<1K5 likes90 downloads2y agoHugging Face13beatsprom /realtime-conversational-voice-agent-duplex-2026 🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026) This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.texttext-generationn<1K0 likes70 downloads22d agoHugging Face14RealTimeData /wikitext_latest Latest Wikitext You could always access the latest Wikipedia texts via this dataset. We update the dataset weekly, on every Sunday. So the dataset always provides the latest Wikipedia texts from the last week. The current dataset on main branch contains the latest wikipedia texts created from 2024-08-26 to 2024-09-02. The data collection is conducted on 2024-09-09. Use the dataset via: ds = datasets.load_dataset('RealTimeData/wikitext_latest') Previsou versions You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/wikitext_latest.textn<1K3 likes59 downloads2y agoHugging Face15findcard12138 /Realtime-SFT Realtime-SFT Realtime-SFT is a 100K-sample streaming-style video question answering dataset constructed from short YouTube videos. Each sample contains a multimodal conversation and frame timestamp metadata that aligns every <|video|> token in the assistant text with one video frame timestamp. This repository does not redistribute video files. It only provides annotations, YouTube video IDs, and timestamp metadata. Users are responsible for obtaining videos according to YouTube… See the full description on the dataset page: https://huggingface.co/datasets/findcard12138/Realtime-SFT.textvisual-question-answering100K<n<1M0 likes59 downloads4mo agoHugging Face16solanaclawd /solana-clawd-realtime-research-instruct Solana Clawd Realtime Research Instruct Instruction-tuning dataset generated by scripts/realtime_dataset_ingest.py from submitted PDFs, notebooks, parquet QA rows, JSON/JSONL files, and local reference text. Contents Total examples: 29058 Train/eval/test: 26152 / 1452 / 1454 Sources: 28 Duplicate examples removed: 0 Duplicate files skipped: 2 Secret-like records skipped: 296 Format Each row uses OpenAI/Hugging Face chat messages: {"messages":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-realtime-research-instruct.texttext-generation10K<n<100K0 likes50 downloads3mo agoHugging Face17RealTimeData /github_july_week2_2023 Dataset Card for "github_july_week2_2023" More Information needed tabularn<1K0 likes48 downloads3y agoHugging Face18beatsprom /audio-speech-realtime-voice-agents-2026 🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition) A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026). Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.tabularaudio-to-audion<1K0 likes46 downloads1mo agoHugging Face19emgena /fastapi_websockets_realtime_backpressure_teaser 🚀 Python Backend - FastAPI WebSockets & Real-Time Connection Backpressure Triage (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Python Backend - FastAPI WebSockets & Real-Time Connection Backpressure Triage on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 📦 What is Inside the Full Production Package: 500 Verified FAANG… See the full description on the dataset page: https://huggingface.co/datasets/emgena/fastapi_websockets_realtime_backpressure_teaser.texttext-generationn<1K0 likes39 downloads6d agoHugging Face20RealTimeData /bbc_news_june_2023 Dataset Card for "bbc_news_june_2023" More Information needed text1K<n<10K0 likes37 downloads3y agoHugging Face21RealTimeData /News_August_2023 Dataset Card for "News_August_2023" This dataset was constructed at 1 Aug 2023, which contains news published from 10 May 2023 to 1 Aug 2023 from various sources. All news articles in this dataset are in English. Created from commoncrawl. image1K<n<10K0 likes37 downloads3y agoHugging Face22RealTimeData /arxiv_latest Latest arXiv You could always access the latest arXiv papers via this dataset. We update the dataset weekly, on every Sunday. So the dataset always provides the latest arXiv papers created in the past week. The current dataset on main branch contains the latest arXiv papers submitted from 2024-09-02 to 2024-09-09. The data collection was conducted on 2024-09-09. Use the dataset via: ds = datasets.load_dataset('RealTimeData/arxiv_latest') Previsou versions You could… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/arxiv_latest.text1K<n<10K4 likes37 downloads2y agoHugging Face23RealTimeData /bbc_latest Latest BBC News You could always access the latest BBC News articles via this dataset. We update the dataset weekly, on every Sunday. So the dataset always provides the latest BBC News article from the last week. The current dataset on main branch contains the latest BBC News articles submitted from 2024-09-02 to 2024-09-09. The data collection is conducted on 2024-09-09. Use the dataset via: ds = datasets.load_dataset('RealTimeData/bbc_latest') Previsou versions You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_latest.textn<1K5 likes32 downloads2y agoHugging Face24realtime-speech /shona_asrtextn<1K0 likes20 downloads2y agoHugging Face25RealTimeData /bbc_news_may_2023 Dataset Card for "bbc_news_may_2023" More Information needed text1K<n<10K0 likes18 downloads3y agoHugging Face26RealTimeData /arxiv_june_2023 Dataset Card for "arxiv_june_2023" More Information needed text10K<n<100K0 likes18 downloads3y agoHugging Face27realtime-speech /shona2audio1K<n<10K1 likes18 downloads2y agoHugging Face28RealTimeData /arxiv_july_week1_2023 Dataset Card for "arxiv_july_week1_2023" More Information needed text1K<n<10K0 likes17 downloads3y agoHugging Face29RealTimeData /News_Seq_2021 Dataset Card for "News_Seq_2021" This dataset was constructed at 1 Seq 2021, which contains news published from 10 June 2021 to 21 Aug 2021 from various sources. All news articles in this dataset are in English. Created from commoncrawl. image1K<n<10K0 likes17 downloads3y agoHugging Face30RealTimeData /github_july_week1_2023 Dataset Card for "github_july_week1_2023" More Information needed textn<1K1 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.