CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nmac /lex_fridman_podcast Dataset Card for "lex_fridman_podcast" Dataset Summary This dataset contains transcripts from the Lex Fridman podcast (Episodes 1 to 325). The transcripts were generated using OpenAI Whisper (large model) and made publicly available at: https://karpathy.ai/lexicap/index.html. Languages English Dataset Structure The dataset contains around 803K entries, consisting of audio transcripts generated from episodes 1 to 325 of the Lex Fridman… See the full description on the dataset page: https://huggingface.co/datasets/nmac/lex_fridman_podcast.textautomatic-speech-recognition100K<n<1M9 likes96 downloads4y agoHugging Face02Gopher-Lab /OpenAI_PodcastSentiment_XTwitterScraper_Example 🔍 X-Twitter Scraper: Real-Time Tweet Search & Scrape Tool Search and scrape X-Twitter for posts by keyword, account, or trending topics.A simple, no-code tool to pull real-time, relevant content in LLM-ready JSON format — perfect for agents, RAG systems, or content workflows. 👉 Start Searching & Scraping on Hugging Face ✨ Features ⚡ Real-Time FetchStream the latest tweets as they’re posted — no delay. 🎯 Flexible SearchSearch by keywords, #hashtags, $cashtags… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/OpenAI_PodcastSentiment_XTwitterScraper_Example.tabulartext-classificationn<1K0 likes17 downloads1y agoHugging Face03ogbrandt /pjf-podcast-qa-sharegptUsed TheBloke/OpenHermes-2-Mistral-7B-GPTQ to convert chunks into QA pairs used for finetuning textn<1K0 likes9 downloads3y agoHugging Face04AriMattiPodcastTranscript /ari-matti-podcast-transcriptionstabular10K<n<100K0 likes6 downloads2y agoHugging Face05elijah0528 /talk_tuah_podcasts Talk Tuah 1 This file is the dataset containing every Talk Tuah podcast transcript. Talk-Tuah-1 is an 80 million parameter GPT trained on all of Hailey Welch's inspirational podcast 'Talk Tuah'. This SOTA frontier model is trained on 13 hours of 'Talk Tuah'. The rationale was the discourse in the 'Talk Tuah' podcast is the most enlightened media that any human has created. Therefore, it should outperform any other LLM on any benchmark. With sufficient training and additional compute… See the full description on the dataset page: https://huggingface.co/datasets/elijah0528/talk_tuah_podcasts.textn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.