CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zwcolin /dot-distance-area Dot Distance / Area over Rich Backgrounds Cross-image spatial-aggregation data used in "Stateful Visual Encoders for Vision-Language Models" (the Cross-image Spatial Aggregation task). A red dot is overlaid on each of 2–5 screenshots (AgentNet backgrounds, downsampled to 384×216), and the model estimates a normalized geometric quantity across the images. Four sub-tasks: Sub-task dir Images / example Quantity dot_distance/ 2 normalized Euclidean distance… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/dot-distance-area.imageimage-to-text100K<n<1M0 likes1.6k downloads4mo agoHugging Face02NoeFlandre /schema-dot-org Geolocated text from the Web Data Commons schema.org GeoCoordinates subset 12,427,530 geolocated text records, extracted from the class-specific GeoCoordinates subset of the Web Data Commons schema.org data set series (release 2024-12). Each record pairs one coordinate pair published on a web page with the text published next to it on that same page. Each source stream is deduplicated by host-local runs: a coordinate-and-name pair is kept once per contiguous host run. A host… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/schema-dot-org.tabulartext-classification10M<n<100M0 likes116 downloads15d agoHugging Face03dots-studio /IMO-AnswerBench-Verified IMO AnswerBench Verified IMO AnswerBench Verified is a human-expert-verified derivative of OpenEvals/IMO-AnswerBench, originally curated by the Google DeepMind Superhuman Reasoning team. Every record in the 400-problem benchmark was reviewed individually. The review identified and corrected 13 records while preserving the benchmark's balanced coverage of four major mathematical areas. Dataset summary Total records: 400 Verification method: record-by-record human… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/IMO-AnswerBench-Verified.textquestion-answeringn<1K1 likes104 downloads1mo agoHugging Face04ammarxix /crawl-lirik-lagu-dot-nettext1K<n<10K0 likes70 downloads3y agoHugging Face05rodrigoramosrs /dotnet ⚡ SolarCurated-TechnicalDocs-QnA A Solar-Powered, Curated Dataset for Technical Reasoning and Instruction Tuning ☀️ 📘 Overview SolarCurated-TechnicalDocs-QnA is a large-scale, meticulously curated dataset containing ≈ 70,000 question–answer pairs, extracted and refined from the official .NET documentation repository. Built entirely through a solar-powered processing pipeline, this dataset demonstrates how high-quality, instruction-tuning data can be generated… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoramosrs/dotnet.text10K<n<100K2 likes63 downloads11mo agoHugging Face06referencesource /dot-hazmat-placarding-thresholds DOT hazmat placarding requirements by hazard class (any-quantity vs. 1,001 lb threshold) Canonical, always-current version: https://referencesource.org/dot-hazmat-placarding-thresholds/ Machine-readable: https://referencesource.org/dot-hazmat-placarding-thresholds/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-19 Stale after: 2027-02-15 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 23 A… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/dot-hazmat-placarding-thresholds.textn<1K0 likes60 downloads28d agoHugging Face07dongguanting /DotamathQA DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning Chengpeng Li, Guanting Dong, Mingfeng Xue, Ru Peng, Xiang Wang, Dayiheng Liu University of Science and Technology of China Qwen, Alibaba Inc. 📃 ArXiv Paper • 📚 Dataset If you find this work helpful for your research, please kindly cite it. @article{li2024dotamath, author = {Chengpeng Li and Guanting Dong and Mingfeng Xue and… See the full description on the dataset page: https://huggingface.co/datasets/dongguanting/DotamathQA.text100K<n<1M2 likes58 downloads2y agoHugging Face08usedot /dot-loom-conductor-v2 Dot Loom Conductor v2 11,100 synthetic traces for training budget-constrained multi-model routers. Trained adapter &nbsp;·&nbsp; Policy explorer &nbsp;·&nbsp; Generator and receipts Each example describes one task, three anonymous worker profiles, hard call, credit, and latency budgets, and the highest-utility feasible Lean, Balanced, or Strict execution plan. Property Value Examples 11,100 Train / validation / test 9,000 / 900 / 1,200 Policy balance 3,700… See the full description on the dataset page: https://huggingface.co/datasets/usedot/dot-loom-conductor-v2.texttext-generation10K<n<100K0 likes55 downloads2mo agoHugging Face09referencesource /dot-random-drug-alcohol-testing-rates-by-mode DOT minimum random drug and alcohol testing rates by transportation mode Canonical, always-current version: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/ Machine-readable: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-19 Stale after: 2027-02-15 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 7… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/dot-random-drug-alcohol-testing-rates-by-mode.textn<1K0 likes48 downloads28d agoHugging Face10dotdotdidi /fine_tuning_datraset_4_openaitextn<1K0 likes39 downloads3y agoHugging Face11Aiden07 /dota2_instruct_promptInstruction-answer dataset generated with GPT 3.5 Turbo using (html) data scrapped from fandom wiki. Data includes the following topics: Heroes Background lore Attributes / Stats Abilities Talents Runes Buildings Items Gameplay mechanics Creeps Pending enhancement: Data cleaning/preprocessing before fed into GPT 3.5 Turbo for instruction-answer set generation Strategy data of each hero, i.e. guide to using each hero Individual items' properties Types of creeps in details Types of runes… See the full description on the dataset page: https://huggingface.co/datasets/Aiden07/dota2_instruct_prompt.textquestion-answering1K<n<10K0 likes39 downloads2y agoHugging Face12dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes38 downloads8mo agoHugging Face13orekpk /DotaMatches-7.32etext1M<n<10M0 likes36 downloads3y agoHugging Face14ajaysri /steer_place_yellow_dot_red_arrow_example_ep201 Placement Yellow-Dot Red-Arrow Sanity Dataset This is a LeRobot-format one-episode sanity export derived from local HDF5 placement data. It uses source recording episode_00201.hdf5, one of the five episodes newer than the existing 197-episode placement export. Dataset size: episodes: 1 frames: 259 videos: 3 export fps: 100 frame stride from 100 Hz source: 1 source episode index: 201 Source goal label: full-resolution target: (564.0, 223.0) px 224x224 overlay target: (98.7… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/steer_place_yellow_dot_red_arrow_example_ep201.tabularroboticsn<1K0 likes32 downloads2mo agoHugging Face15AlphaChat-dotcom /alphaqa-cross-v03-sample AlphaQA-Cross v03 10,003 cross-document QA pairs (FREE SAMPLE) requiring reasoning across multiple Wikipedia articles. Unlike single-article QA datasets (SQuAD, NaturalQuestions), AlphaQA-Cross requires combining information from 2+ articles to answer each question. What Makes This Different Feature HotpotQA SQuAD 2.0 AlphaQA-Cross Cross-article Partial No Yes (all) Reasoning types 2 1 5 Generated by Crowd Crowd 35B LLM Quality consistency… See the full description on the dataset page: https://huggingface.co/datasets/AlphaChat-dotcom/alphaqa-cross-v03-sample.textquestion-answering10K<n<100K0 likes25 downloads3mo agoHugging Face16dots-studio /dots-imo2026From July 15 to 16, the 67th International Mathematical Olympiad (IMO 2026) was held in Shanghai. The dots team was invited by the IMO Organizing Committee and took part in IMO 2026 with an internal version of dots-note-3.0. This year's IMO brought together 666 contestants from 117 countries, setting records for both the number of participating countries and the number of contestants. Following official marking organized by the committee, dots-note-3.0 received the full 7 points on all six… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/dots-imo2026.tabularn<1K5 likes25 downloads2mo agoHugging Face17aemmeath /all_projects_per_file_dataset_dotatext1K<n<10K0 likes24 downloads3mo agoHugging Face18Jethro85 /dotrag Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Jethro85/dotrag.text10K<n<100K0 likes16 downloads7mo agoHugging Face19bluecopa /dotsocr-extractions-holdouttabularn<1K0 likes13 downloads9mo agoHugging Face20AlphaChat-dotcom /alphaqa-cross-v03gated AlphaQA-Cross v03 81,444 cross-document QA pairs requiring reasoning across multiple Wikipedia articles. Unlike single-article QA datasets (SQuAD, NaturalQuestions), AlphaQA-Cross requires combining information from 2+ articles to answer each question. What Makes This Different Feature HotpotQA SQuAD 2.0 AlphaQA-Cross Cross-article Partial No Yes (all) Reasoning types 2 1 5 Generated by Crowd Crowd 35B LLM Quality consistency Variable Variable… See the full description on the dataset page: https://huggingface.co/datasets/AlphaChat-dotcom/alphaqa-cross-v03.textquestion-answering10K<n<100K0 likes9 downloads3mo agoHugging Face21dots123 /theorem-engine-v3-training-datatext1K<n<10K0 likes8 downloads4mo agoHugging Face22rita1706 /dotsocr_bank_statement_1K_v2 dotsocr_bank_statement_1K Dataset Description This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab. DotsOCR Format Features Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right) Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_1K_v2.image1K<n<10K0 likes7 downloads1y agoHugging Face23dotesec /buraktextn<1K0 likes6 downloads2y agoHugging Face24albertvillanova /tmp-jsonl-dottextn<1K0 likes5 downloads2y agoHugging Face25DotStarFish /sadadatextn<1K0 likes5 downloads1y agoHugging Face26rita1706 /dotsocr_bank_statement_half dotsocr_bank_statement_half Dataset Description This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab. DotsOCR Format Features Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right) Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_half.imagen<1K0 likes4 downloads1y agoHugging Face27AlphaChat-dotcom /Wikipedia-Contradictions-2026gated WikiTruth: 184 Cross-Article Contradictions Found in Wikipedia by AI An AI system that read 86 billion tokens of Wikipedia (21 million article chunks) found 184 factual contradictions — cases where one Wikipedia article directly conflicts with another. These aren't formatting errors or vandalism. They're genuine knowledge conflicts that persist because no human editor reads every article. A system with 1M+ token context noticed what humans couldn't: facts stated in one article… See the full description on the dataset page: https://huggingface.co/datasets/AlphaChat-dotcom/Wikipedia-Contradictions-2026.tabularquestion-answeringn<1K0 likes4 downloads3mo agoHugging Face28ali-dot-com /Researcher_Writing_Style_FineTuningtextn<1K3 likes3 downloads2y agoHugging Face296shmqshy7q-dot /my-sft-datasettextn<1K0 likes2 downloads4mo agoHugging Face30Symato /DOT_excite_data_v0.0gated DOT Do One Thing Excite: Bài toán tìm nội dung trích dẫn trong văn bản Có rất ít dataset có liên quan tới citation, 2 datasets chúng tôi tìm thấy là https://huggingface.co/datasets/THUDM/LongCite-45k https://huggingface.co/datasets/THUDM/webglm-qa Chúng tôi dịch webgml-qa (chất lượng rất tốt) và phần answer của longcite sang tiếng Việt, rồi tạo citing data từ đó. longcite rất lớn nên chúng tôi chỉ chọn những sample có đội dài ctxlen <= ~24k. Ngoài phần dữ liệu… See the full description on the dataset page: https://huggingface.co/datasets/Symato/DOT_excite_data_v0.0.text10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.