datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dot-distance-area
Dot Distance / Area over Rich Backgrounds
Cross-image spatial-aggregation data used in "Stateful Visual Encoders for
Vision-Language Models" (the Cross-image Spatial Aggregation task). A red dot
is overlaid on each of 2–5 screenshots (AgentNet backgrounds, downsampled to
384×216), and the model estimates a normalized geometric quantity across the
images. Four sub-tasks:
Sub-task dir
Images / example
Quantity
dot_distance/
2
normalized Euclidean distance… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/dot-distance-area.schema-dot-org
Geolocated text from the Web Data Commons schema.org GeoCoordinates subset
12,427,530 geolocated text records, extracted from the class-specific
GeoCoordinates subset of the Web Data Commons schema.org data set series
(release 2024-12). Each record pairs one coordinate pair published on a web page
with the text published next to it on that same page.
Each source stream is deduplicated by host-local runs: a coordinate-and-name
pair is kept once per contiguous host run. A host… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/schema-dot-org.IMO-AnswerBench-Verified
IMO AnswerBench Verified
IMO AnswerBench Verified is a human-expert-verified derivative of OpenEvals/IMO-AnswerBench, originally curated by the Google DeepMind Superhuman Reasoning team. Every record in the 400-problem benchmark was reviewed individually. The review identified and corrected 13 records while preserving the benchmark's balanced coverage of four major mathematical areas.
Dataset summary
Total records: 400
Verification method: record-by-record human… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/IMO-AnswerBench-Verified.crawl-lirik-lagu-dot-netdotnet
⚡ SolarCurated-TechnicalDocs-QnA
A Solar-Powered, Curated Dataset for Technical Reasoning and Instruction Tuning ☀️
📘 Overview
SolarCurated-TechnicalDocs-QnA is a large-scale, meticulously curated dataset containing ≈ 70,000 question–answer pairs, extracted and refined from the official .NET documentation repository.
Built entirely through a solar-powered processing pipeline, this dataset demonstrates how high-quality, instruction-tuning data can be generated… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoramosrs/dotnet.dot-hazmat-placarding-thresholds
DOT hazmat placarding requirements by hazard class (any-quantity vs. 1,001 lb threshold)
Canonical, always-current version: https://referencesource.org/dot-hazmat-placarding-thresholds/
Machine-readable: https://referencesource.org/dot-hazmat-placarding-thresholds/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2027-02-15 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 23
A… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/dot-hazmat-placarding-thresholds.DotamathQA
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
Chengpeng Li, Guanting Dong, Mingfeng Xue, Ru Peng, Xiang Wang, Dayiheng Liu
University of Science and Technology of China
Qwen, Alibaba Inc.
📃 ArXiv Paper
•
📚 Dataset
If you find this work helpful for your research, please kindly cite it.
@article{li2024dotamath,
author = {Chengpeng Li and
Guanting Dong and
Mingfeng Xue and… See the full description on the dataset page: https://huggingface.co/datasets/dongguanting/DotamathQA.dot-loom-conductor-v2
Dot Loom Conductor v2
11,100 synthetic traces for training budget-constrained multi-model routers.
Trained adapter
·
Policy explorer
·
Generator and receipts
Each example describes one task, three anonymous worker profiles, hard call, credit, and latency
budgets, and the highest-utility feasible Lean, Balanced, or Strict execution plan.
Property
Value
Examples
11,100
Train / validation / test
9,000 / 900 / 1,200
Policy balance
3,700… See the full description on the dataset page: https://huggingface.co/datasets/usedot/dot-loom-conductor-v2.dot-random-drug-alcohol-testing-rates-by-mode
DOT minimum random drug and alcohol testing rates by transportation mode
Canonical, always-current version: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/
Machine-readable: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2027-02-15 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 7… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/dot-random-drug-alcohol-testing-rates-by-mode.fine_tuning_datraset_4_openaidota2_instruct_promptInstruction-answer dataset generated with GPT 3.5 Turbo using (html) data scrapped from fandom wiki. Data includes the following topics:
Heroes
Background lore
Attributes / Stats
Abilities
Talents
Runes
Buildings
Items
Gameplay mechanics
Creeps
Pending enhancement:
Data cleaning/preprocessing before fed into GPT 3.5 Turbo for instruction-answer set generation
Strategy data of each hero, i.e. guide to using each hero
Individual items' properties
Types of creeps in details
Types of runes… See the full description on the dataset page: https://huggingface.co/datasets/Aiden07/dota2_instruct_prompt.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.DotaMatches-7.32esteer_place_yellow_dot_red_arrow_example_ep201
Placement Yellow-Dot Red-Arrow Sanity Dataset
This is a LeRobot-format one-episode sanity export derived from local HDF5 placement data.
It uses source recording episode_00201.hdf5, one of the five episodes newer
than the existing 197-episode placement export.
Dataset size:
episodes: 1
frames: 259
videos: 3
export fps: 100
frame stride from 100 Hz source: 1
source episode index: 201
Source goal label:
full-resolution target: (564.0, 223.0) px
224x224 overlay target: (98.7… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/steer_place_yellow_dot_red_arrow_example_ep201.alphaqa-cross-v03-sample
AlphaQA-Cross v03
10,003 cross-document QA pairs (FREE SAMPLE) requiring reasoning across multiple Wikipedia articles.
Unlike single-article QA datasets (SQuAD, NaturalQuestions), AlphaQA-Cross requires combining information from 2+ articles to answer each question.
What Makes This Different
Feature
HotpotQA
SQuAD 2.0
AlphaQA-Cross
Cross-article
Partial
No
Yes (all)
Reasoning types
2
1
5
Generated by
Crowd
Crowd
35B LLM
Quality consistency… See the full description on the dataset page: https://huggingface.co/datasets/AlphaChat-dotcom/alphaqa-cross-v03-sample.dots-imo2026From July 15 to 16, the 67th International Mathematical Olympiad (IMO 2026) was held in Shanghai. The dots team was invited by the IMO Organizing Committee and took part in IMO 2026 with an internal version of dots-note-3.0.
This year's IMO brought together 666 contestants from 117 countries, setting records for both the number of participating countries and the number of contestants.
Following official marking organized by the committee, dots-note-3.0 received the full 7 points on all six… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/dots-imo2026.all_projects_per_file_dataset_dotadotrag
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Jethro85/dotrag.dotsocr-extractions-holdoutalphaqa-cross-v03
AlphaQA-Cross v03
81,444 cross-document QA pairs requiring reasoning across multiple Wikipedia articles.
Unlike single-article QA datasets (SQuAD, NaturalQuestions), AlphaQA-Cross requires combining information from 2+ articles to answer each question.
What Makes This Different
Feature
HotpotQA
SQuAD 2.0
AlphaQA-Cross
Cross-article
Partial
No
Yes (all)
Reasoning types
2
1
5
Generated by
Crowd
Crowd
35B LLM
Quality consistency
Variable
Variable… See the full description on the dataset page: https://huggingface.co/datasets/AlphaChat-dotcom/alphaqa-cross-v03.theorem-engine-v3-training-datadotsocr_bank_statement_1K_v2
dotsocr_bank_statement_1K
Dataset Description
This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab.
DotsOCR Format Features
Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right)
Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_1K_v2.buraktmp-jsonl-dotsadadadotsocr_bank_statement_half
dotsocr_bank_statement_half
Dataset Description
This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab.
DotsOCR Format Features
Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right)
Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_half.Wikipedia-Contradictions-2026
WikiTruth: 184 Cross-Article Contradictions Found in Wikipedia by AI
An AI system that read 86 billion tokens of Wikipedia (21 million article chunks) found 184 factual contradictions — cases where one Wikipedia article directly conflicts with another.
These aren't formatting errors or vandalism. They're genuine knowledge conflicts that persist because no human editor reads every article. A system with 1M+ token context noticed what humans couldn't: facts stated in one article… See the full description on the dataset page: https://huggingface.co/datasets/AlphaChat-dotcom/Wikipedia-Contradictions-2026.Researcher_Writing_Style_FineTuningmy-sft-datasetDOT_excite_data_v0.0
DOT
Do One Thing
Excite: Bài toán tìm nội dung trích dẫn trong văn bản
Có rất ít dataset có liên quan tới citation, 2 datasets chúng tôi tìm thấy là
https://huggingface.co/datasets/THUDM/LongCite-45k
https://huggingface.co/datasets/THUDM/webglm-qa
Chúng tôi dịch webgml-qa (chất lượng rất tốt) và phần answer của longcite sang tiếng Việt, rồi tạo citing data từ đó.
longcite rất lớn nên chúng tôi chỉ chọn những sample có đội dài ctxlen <= ~24k.
Ngoài phần dữ liệu… See the full description on the dataset page: https://huggingface.co/datasets/Symato/DOT_excite_data_v0.0.
