CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu.tabulartext-generation1B<n<10B1.3k likes426k downloads1y agoHugging Face02Helsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes211k downloads5mo agoHugging Face03eduagarcia-temp /llm_pt_leaderboard_raw_results0 likes103k downloads1y agoHugging Face04eddmpython /dartlab-data DartLab 데이터 종목코드 하나로 읽는 한국 DART + 미국 SEC EDGAR 공시 데이터 Structured Korean (DART) and US (SEC EDGAR) disclosure data, ready as Parquet. 무엇인가요? DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다. 한국 전 상장사(약 2,700사)와 미국 주요 상장사(약 1,000사)의 재무제표, 사업보고서 본문, 정형 공시, 주가, 거시지표가 들어 있습니다. 이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음). 코드 없이도 바로 씁니다… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/dartlab-data.table-question-answering1M<n<10M14 likes90k downloads48m agoHugging Face05airtrain-ai /fineweb-edu-fortified Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it? Fineweb-Edu-Fortified is a dataset derived from Fineweb-Edu by applying exact-match deduplication across the whole dataset and producing an embedding for each row. The number of times the text from each row appears is also included as a count column. The embeddings were produced using TaylorAI/bge-micro Fineweb and… See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.tabulartext-generation100M<n<1B65 likes60k downloads2y agoHugging Face06opencsg /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1.text-generation10B<n<100B80 likes54k downloads8mo agoHugging Face07edinburgh-dawg /mmlu-redux-2.0 Dataset Card for MMLU-Redux-2.0 MMLU-Redux is a subset of 5,700 manually re-annotated questions across 57 MMLU subjects. News [2025.02.25] We corrected one annotation in Abstract Algebra subset, as noted in the Issue #2. [2025.02.08] We corrected one annotation in High School Mathematics subset, as noted in the PlatinumBench paper. [2025.01.23] MMLU-Redux is accepted to NAACL 2025! Dataset Details Dataset Description Each data point in… See the full description on the dataset page: https://huggingface.co/datasets/edinburgh-dawg/mmlu-redux-2.0.textquestion-answering1K<n<10K38 likes51k downloads2y agoHugging Face08edbeeching /gia-dataset-tokenized-2024-2 Dataset Card for "gia-dataset-tokenized-2024-2" More Information needed 100K<n<1M0 likes37k downloads3y agoHugging Face09fineweb-retrieval /fineweb-edu-indexThis dataset contains the embeddings for the full fineweb-edu, embedded with the Cohere Embed V3 model. You can search on this dataset with just 500MB of memory using DiskVectorIndex. Installation & Usage Get your free Cohere API key from cohere.com. You must set this API key as an environment variable: export COHERE_API_KEY=your_api_key Install the package: pip install DiskVectorIndex You can then search via: from DiskVectorIndex import DiskVectorIndex index =… See the full description on the dataset page: https://huggingface.co/datasets/fineweb-retrieval/fineweb-edu-index.0 likes29k downloads1y agoHugging Face10edwarddgao /open-apply-jobs Open-Apply Jobs A daily-refreshed open dataset of active job postings sourced directly from public ATS APIs (Greenhouse, Lever, Ashby). Every record can be traced back to the hiring company's own career board. Refresh: automated daily at 06:00 UTC Partitioning: Hive-partitioned Parquet (date=YYYY-MM-DD/source={ats}) Source code: https://github.com/edwarddgao/openapply Usage from datasets import load_dataset ds = load_dataset('edwarddgao/open-apply-jobs') #… See the full description on the dataset page: https://huggingface.co/datasets/edwarddgao/open-apply-jobs.tabulartext-classification10M<n<100M9 likes21k downloads16h agoHugging Face11HuggingFaceFW /fineweb-edu-score-2 📚 FineWeb-Edu-score-2 1.3 trillion tokens of the finest educational data the 🌐 web has to offer What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens (FineWeb-Edu) and 5.4T tokens of educational web pages filtered from 🍷 FineWeb dataset. This is the 5.4 trillion version. Note: this version uses a lower educational score threshold = 2, which results in more documents, but lower quality compared to the 1.3T version. For more details check the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2.tabulartext-generation10B<n<100B89 likes21k downloads1y agoHugging Face12opencsg /chinese-fineweb-edu This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.texttext-generation10M<n<100M117 likes18k downloads10mo agoHugging Face13opencsg /Fineweb-Edu-Chinese-V2.2 Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train) [[中文]] | [[English]] OpenCSG Community | 👾 GitHub | 📖 Technical Report Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain. This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.2.text-generation10B<n<100B83 likes17k downloads8mo agoHugging Face14tonyc54 /Total_Editing_Synthetic_Video_Albedo_Full0 likes17k downloads1y agoHugging Face15HuggingFaceFW /finepdfs-edu 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.tabulartext-generation10M<n<100M98 likes16k downloads11mo agoHugging Face16TeraflopAI /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.texttext-generation1M<n<10M47 likes16k downloads5mo agoHugging Face17EdinburghNLP /xsum Dataset Card for "xsum" Dataset Summary Extreme Summarization (XSum) Dataset. There are three features: document: Input news article. summary: One sentence summary of the article. id: BBC ID of the article. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 257.30 MB Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/EdinburghNLP/xsum.textsummarization100K<n<1M153 likes15k downloads9mo agoHugging Face18Kashu7100 /eden_objaverse3d1K<n<10K0 likes15k downloads6mo agoHugging Face19edcci /GenECGGenECG is an image-based ECG dataset which has been created from the PTB-XL dataset (https://physionet.org/content/ptb-xl/1.0.3/). The PTB-XL dataset is a signal-based ECG dataset comprising 21799 unique ECGs. GenECG is divided into the following subsets: -Dataset A: ECGs without imperfections (Dataset_A_ECGs_without_imperfections) - This subset includes 21799 ECG images that have been generated directly from the original PTB-XL recordings, free from any visual imperfections. -Dataset B: ECGs… See the full description on the dataset page: https://huggingface.co/datasets/edcci/GenECG.image3 likes14k downloads1y agoHugging Face20edinburghcstr /ami Dataset Card for AMI Dataset Description The AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals synchronized to a common timeline. These include close-talking and far-field microphones, individual and room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings, the participants also have unsynchronized pens available to them that record what is written. The meetings were… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/ami.audioautomatic-speech-recognition100K<n<1M96 likes12k downloads9mo agoHugging Face21JoTalbot /ua-edrsr ЄДРСР — судові рішення України (нормалізоване дзеркало) Автоматичне дзеркало офіційних публікацій Єдиного державного реєстру судових рішень на data.gov.ua. Пайплайн: JoTalbot/ukraine. Роки та обсяги Рік Записів SHA-256 архіву 2006 340171 4e19811feef9… 2007 1084514 49cf87b3a0e1… 2008 2185311 ce7dbc1b8da3… 2009 3539627 3be90c8824ab… 2010 5869727 b9cd6b5e2562… 2011 7128372 4b845ee3b4bd… 2012 6903131 ba145669d313… 2013 7704297 1794ceb7e3a5…… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-edrsr.text100M<n<1B0 likes11k downloads2h agoHugging Face22UCSC-VLAA /gpt-edit-simplerimage1M<n<10M13 likes11k downloads1y agoHugging Face23karpathy /fineweb-edu-100b-shuffletext10M<n<100M171 likes11k downloads1y agoHugging Face24alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K2 likes10k downloads2d agoHugging Face25EdisonScientific /labbench2gated LABBench2 LABBench2 is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the Language Agent Biology Benchmark (LAB-Bench), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts. LABBench2 provides a meaningful jump in difficulty over LAB-Bench (model-specific accuracy differences range from −26% to −46% across subtasks), underscoring continued room for improvement.… See the full description on the dataset page: https://huggingface.co/datasets/EdisonScientific/labbench2.textquestion-answering1K<n<10K60 likes9.7k downloads7mo agoHugging Face26c3po-ai /edgar-corpusThe dataset contains annual filings (10K) of all publicly traded firms from 1993-2020. The table data is stripped but all text is retained. This dataset allows easy access to the EDGAR-CORPUS dataset based on the paper EDGAR-CORPUS: Billions of Tokens Make The World Go Round (See References in README.md for details).textother100K<n<1M10 likes9.7k downloads3y agoHugging Face27eduagarcia-temp /llm_pt_leaderboard_requests0 likes8.7k downloads6mo agoHugging Face28lance-format /fineweb-edu FineWeb-Edu (Lance Format) A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance. Key features Cleaned passage text in the text column with the source url and title carried alongside. Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.tabulartext-retrieval1B<n<10B8 likes8.4k downloads4mo agoHugging Face29rytfh /edtyhgv6 likes8.1k downloads5mo agoHugging Face30ByteDance-Seed /EdgeBench Overview EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.texttext-generationn<1K84 likes7.5k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.