CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /meta-active-readingtext1B<n<10B37 likes5k downloads1y agoHugging Face02emozilla /sat-reading Dataset Card for "sat-reading" This dataset contains the passages and questions from the Reading part of ten publicly available SAT Practice Tests. For more information see the blog post Language Models vs. The SAT Reading Test. For each question, the reading passage from the section it is contained in is prefixed. Then, the question is prompted with Question #:, followed by the four possible answers. Each entry ends with Answer:. Questions which reference a diagram, chart, table… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/sat-reading.textn<1K3 likes1.1k downloads4y agoHugging Face03betterMateusz /SAT_Writting_Reading_Assessment_Question_Bank Dataset Card for SAT Reading and Writing Dataset This dataset card aims to be a base template for the SAT Reading and Writing Dataset, optimized for use with Hugging Face's datasets library. Dataset Details Dataset Description This dataset contains SAT Reading and Writing assessment questions sourced from the College Board's SAT Suite Question Bank, intended for use in training and evaluating Language Models like LLMs. Curated by: College Board License:… See the full description on the dataset page: https://huggingface.co/datasets/betterMateusz/SAT_Writting_Reading_Assessment_Question_Bank.textn<1K2 likes851 downloads3y agoHugging Face04ReadingTimeMachine /rtm-sgt-ocr-v1 Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic Data from the ar𝜒iv for OCR Post Correction of Historic Scientific Articles". Synthetic ground truth (SGT) sentences have been mined from the ar𝜒iv Bulk Downloads source documents, and Optical Character Recognition (OCR) sentences have been generated with the Tesseract OCR engine on the PDF pages generated from compiled source documents.… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/rtm-sgt-ocr-v1.texttext-classification1M<n<10M4 likes674 downloads1y agoHugging Face05manus4oHER /cia-declassified-reading-room CIA Declassified Reading Room HF Library Target account: manus4oHER This project is a streaming pipeline for building a Hugging Face dataset mirror of public CIA declassified Reading Room / CREST records without staging the full corpus on this laptop. The laptop stores only scripts, small manifests, and logs. Bulk crawling should run in Hugging Face Jobs, one bounded page range per job. Each job uploads its own shard and then exits. Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/manus4oHER/cia-declassified-reading-room.document10K<n<100K2 likes458 downloads3mo agoHugging Face06ChunkrAI /chunkr-reading-order-bench-oss Chunkr Reading Order Bench - Open Source Subset Open-source subset of the Chunkr Reading Order benchmark dataset, containing 733 professionally annotated documents with detailed reading order annotations across diverse document layouts. This dataset benchmarks reading order detection models on complex, real-world documents including financial reports, legal contracts, research papers, medical records, and more. Each document includes ground truth reading order sequences essential… See the full description on the dataset page: https://huggingface.co/datasets/ChunkrAI/chunkr-reading-order-bench-oss.imageobject-detectionn<1K1 likes379 downloads8mo agoHugging Face07zilongwang /ReadingBank ReadingBank ReadingBank is a benchmark dataset for reading order detection built with weak supervision from WORD documents, which contains 500K document images with a wide range of document types as well as the corresponding reading order information. Our paper "LayoutReader: Pre-training of Text and Layout for Reading Order Detection" has been accepted by EMNLP 2021. Refer to the official repo for more details: https://github.com/doc-analysis/ReadingBank tabular100K<n<1M4 likes222 downloads2y agoHugging Face08community-datasets /parsinlu_reading_comprehension Dataset Card for PersiNLU (Reading Comprehension) Dataset Summary A Persian reading comprehenion task (generating an answer, given a question and a context paragraph). The questions are mined using Google auto-complete, their answers and the corresponding evidence documents are manually annotated by native speakers. Supported Tasks and Leaderboards [More Information Needed] Languages The text dataset is in Persian (fa). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/parsinlu_reading_comprehension.textquestion-answering1K<n<10K3 likes153 downloads2y agoHugging Face09manu /french-bench-grammar-vocab-reading Dataset Card for "french-bench-grammar-vocab-reading" More Information needed textn<1K4 likes152 downloads1y agoHugging Face10goodcoffee /Meter_Readingimage1K<n<10K3 likes143 downloads2y agoHugging Face11NbAiLab /lunde_nor_nob_reading_optimisedTest only - not for training. First version - 0.1 of lunde_nor_nob_reading_optimised This dataset does not contain any audio data. Export Details Train samples: 10040932 Validation samples: 0 Test samples: 0 Dataset created using search datasets:lunde_nor_nob_reading_optimised. textautomatic-speech-recognition10M<n<100M0 likes112 downloads8mo agoHugging Face12Emulated-Inc /reading-comprehension-training-pool Reading comprehension training pool Public reading comprehension questions from six datasets, each a question about a passage with an answer that is a span of it, a number or a date, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 310728 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/reading-comprehension-training-pool.textquestion-answeringn<1K0 likes101 downloads13d agoHugging Face13DandinPower /chinese-reading-comprehensiontabular10K<n<100K0 likes95 downloads2y agoHugging Face14Symblai /reading-with-intent 📑 Paper    |    📑 Blog We introduce the Reading with Intent task and prompting method and accompanying datasets. The goal of this task is to have LLMs read beyond the surface level of text and integrate an understanding of the underlying sentiment of a text when reading it. The focus of this work is on sarcastic text. We've released: The code used creating the sarcastic datasets The sarcasm-poisoned dataset The reading with intent prompting method Citation… See the full description on the dataset page: https://huggingface.co/datasets/Symblai/reading-with-intent.text10K<n<100K0 likes83 downloads2y agoHugging Face15formospeech /hat_asr_sixian_reading_cleangated TRAIN Subset lang_group hours n_utts n_chars secs/utt chars/sec Hakka_Sixian 客語_四縣 203.03 106,493 2,137,244 6.86 2.92 Total - 203.03 106,493 2,137,244 6.86 2.92 audio100K<n<1M1 likes59 downloads9mo agoHugging Face16Emulated-Inc /reading-comprehension-qa-training-pool Reading comprehension question answering training pool Public question answering data from five datasets, every row a question, the passages it is answered from and every acceptable answer, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 311462 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/reading-comprehension-qa-training-pool.textquestion-answering100K<n<1M0 likes59 downloads14d agoHugging Face17thangvip /law-reading-comprehension-qa-filteredtext100K<n<1M0 likes55 downloads2y agoHugging Face18nlplabtdtu /law-reading-comprehension-qagatedtext10K<n<100K1 likes49 downloads2y agoHugging Face19lianghsun /Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading Dataset Card for Nexdata/Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading Dataset Summary This dataset is just a sample of Taiwan Mandarin Speech dataset(paid dataset) by mobile phone reading.The data collects 204 Taiwan residents with 450 sentences for each speaker. The recorded is rich in content, including economy, entertainment, news, spoken language, numbers, letters, etc., covering general scenes and human-computer interaction scenes. Manual transcription of text… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading.audion<1K0 likes48 downloads8mo agoHugging Face20formospeech /hat_asr_hailu_reading_cleangated TRAIN Subset lang_group hours n_utts n_chars secs/utt chars/sec Hakka_Hailu 客語_海陸 197.58 104,555 2,018,861 6.80 2.84 Total - 197.58 104,555 2,018,861 6.80 2.84 audio100K<n<1M1 likes47 downloads9mo agoHugging Face21Growspot /houseplant-light-needs-and-indoor-lux-readings Houseplant Light Needs and Indoor Lux Readings Two files, both CC BY 4.0. Archived with a DOI on Zenodo: 10.5281/zenodo.22023337. Note for loaders: both CSVs open with commented header lines (#) carrying the snapshot date, the licence and the band definitions. Skip them when reading. Files species-light-levels.csv — one row per houseplant species in the GrowSpot care library, with the light tier it belongs to. Columns: common_name, scientific_name… See the full description on the dataset page: https://huggingface.co/datasets/Growspot/houseplant-light-needs-and-indoor-lux-readings.tabular1K<n<10K0 likes42 downloads28d agoHugging Face22electricsheepafrica /africa-synth-education-early-grade-reading-proficiency-all Africa Synth Education Early Grade Reading Proficiency All | Africa (World Bank) Size category: 10K<n<100K - Formats: csv - Sector: education - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Education datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-early-grade-reading-proficiency-all.tabulartabular-classification10K<n<100K0 likes41 downloads1mo agoHugging Face23EmberTrail299 /right-reading-5bd6e9 right-reading-5bd6e9 Synthetic products test data: 39 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/EmberTrail299/right-reading-5bd6e9.tabularn<1K0 likes41 downloads15d agoHugging Face24241288-kltn /law-reading-comprehension-qatext10K<n<100K0 likes40 downloads2y agoHugging Face25formospeech /hat_asr_sixian_reading_clean_rgated hat_asr_sixian_reading_clean_r This dataset is an enhanced -R variant of formospeech/hat_asr_sixian_reading_clean. Summary Subset: Hakka_Sixian Dialect: 客語四縣 Train samples: 106493 Audio: enhanced 24 kHz WAV TRAIN Subset lang_group hours n_utts n_chars secs/utt chars/sec Hakka_Sixian 客語_四縣 203.03 106,493 5,597,645 6.86 7.66 Total - 203.03 106,493 5,597,645 6.86 7.66 Processing Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_sixian_reading_clean_r.audiotext-to-speech100K<n<1M1 likes40 downloads14d agoHugging Face26thangvip /law-reading-comprehension-qatext100K<n<1M0 likes33 downloads2y agoHugging Face27albertklorer /readingbank ReadingBank (HF conversion) Source paper: https://arxiv.org/abs/2108.11591 Original data: https://mail2sysueducn-my.sharepoint.com/:u:/g/personal/huangyp28_mail2_sysu_edu_cn/Efh3ZWjsA-xFrH2FSjyhSVoBMak6ypmbABWmJEmPwtKhhw?e=tbthMD Created with: https://github.com/albertklor/reading-bank Fields: file_name (name of the file): str page_number (index of the page number): int bounding_boxes (normalized bounding boxes in [x0, y0, x1, y1] format): list[list[int]] text… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/readingbank.text100K<n<1M1 likes32 downloads1y agoHugging Face28formospeech /hat_asr_sixian_reading_cmgated TRAIN Subset lang_name hours n_utts n_chars_in_utts secs/utt chars/sec n_sents n_chars_in_sents hak_sx Hakka_Sixian 204.60 105,513 2,152,732 6.98 2.92 0 0 Total - 204.60 105,513 2,152,732 6.98 2.92 0 0 audio100K<n<1M0 likes31 downloads2mo agoHugging Face29formospeech /hat_asr_hailu_reading_clean_rgated hat_asr_hailu_reading_clean_r This dataset is an enhanced -R variant of formospeech/hat_asr_hailu_reading_clean. Summary Subset: Hakka_Hailu Dialect: 客語海陸 Train samples: 104555 Audio: enhanced 24 kHz WAV TRAIN Subset lang_group hours n_utts n_chars secs/utt chars/sec Hakka_Hailu 客語_海陸 197.58 104,555 5,287,046 6.80 7.43 Total - 197.58 104,555 5,287,046 6.80 7.43 Processing Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_hailu_reading_clean_r.audiotext-to-speech100K<n<1M1 likes31 downloads14d agoHugging Face30AkabekoLabs /nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。 データセット統計 train: 2,418 サンプル validation: 302 サンプル test: 303 サンプル 総サンプル数: 3,023 ソース 生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/ サンプルデータ { "instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。", "input": "「究」の訓読みは?", "think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.tabulartext-generation1K<n<10K0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.