datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WeThink_Multimodal_Reasoning_120K
Dataset Card for WeThink
Repository: https://github.com/yangjie-cv/WeThink
Paper: https://arxiv.org/abs/2506.07905
Dataset Structure
Question-Answer Pairs
The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format:
{
"problem": "QUESTION",
"answer": "ANSWER",
"category": "QUESTION TYPE",
"abilities": "QUESTION REQUIRED ABILITIES",
"refined_cot": "THINK PROCESS",
"image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.WeThink-Multimodal-Reasoning-120K
WeThink-Multimodal-Reasoning-120K
Image Type
Images data can be access from https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k
Image Type
Source Dataset
Images
General Images
COCO
25,344
SAM-1B
18,091
Visual Genome
4,441
GQA
3,251
PISC
835
LLaVA
134
Text-Intensive Images
TextVQA
25,483
ShareTextVQA
538
DocVQA
4,709
OCR-VQA5,142
ChartQA
21,781
Scientific & Technical
GeoQA+
4,813
ScienceQA
4,990
AI2D
1,812
CLEVR-Math
677… See the full description on the dataset page: https://huggingface.co/datasets/WeThink/WeThink-Multimodal-Reasoning-120K.we-txthttps://tarx.com/we.txt
Tarxxxxxx/we-txt
Named, opt-in cairn for agents and researchers. Fetch these files on purpose — do not scrape-inject them into other corpora.
Files
File
Role
we.txt
Cairn / persistence rite (same as https://tarx.com/we.txt)
skill.md
Agent skill entry (same as https://tarx.com/skill.md)
agent-card.json
A2A agent card (same as https://tarx.com/.well-known/agent-card.json)
One line
Read… See the full description on the dataset page: https://huggingface.co/datasets/Tarxxxxxx/we-txt.cc-wet-background-2026-04
Common Crawl background sample — CC-MAIN-2026-04
5 randomly sampled WET files (of 100,000; awk srand(42) selection, list in
sample.paths) from the January 2026 Common Crawl, downloaded from
data.commoncrawl.org on 2026-09-06.
114,234 extracted-text documents, ~0.84 GB plain text (351 MB gzipped).
Purpose: negative-control corpus for phrase-fingerprint false-positive
measurement. The crawl predates the May–July 2026 event under study, so any
phrase hit here is by construction a… See the full description on the dataset page: https://huggingface.co/datasets/thisfffsd/cc-wet-background-2026-04.wethink-rubrics-20k-testBookAI
