datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
real-vs-ai-corpus
Real vs AI Corpus
A large-scale binary image classification dataset for training AI-image detectors.
Built by Zitacron from 17 public HuggingFace
sources, all streaming-merged with no intermediate local storage.
All constituent sources are CC BY 4.0, Apache 2.0, or MIT — fully
commercially usable. Models trained on this dataset may be used commercially
without restriction, provided attribution requirements below are met.
Usage
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Zitacron/real-vs-ai-corpus.entigraph-quality-corpus
EntiGraph Synthetic Corpus
The EntiGraph Synthetic Corpus is a dataset generated using the EntiGraph synthetic data augmentation algorithm, as described in the paper "Synthetic Continued Pretraining" by Yang et al. (2024).
The code for generating the dataset is avaliable at the Synthetic Continued Pretraining GitHub repo.
Dataset Summary
This dataset contains synthetic text generated by the EntiGraph algorithm, which extracts entities from source documents and… See the full description on the dataset page: https://huggingface.co/datasets/zitongyang/entigraph-quality-corpus.agent-runtime-recovery-bench
Agent Runtime Recovery Benchmark
This fully public dataset combines the current observation-restricted Agent
Runtime Recovery Benchmark with causally qualified native runtime cases for real
coding agents.
Contents
Source
Cases
Description
benchmark
769
Observation-restricted runtime recovery cases
real_agent_native
125
90 OpenHands and 35 mini-SWE-agent native runtime cases
Total
894
One unified public dataset
All rows are stored in one all… See the full description on the dataset page: https://huggingface.co/datasets/zitong1/agent-runtime-recovery-bench.NovGauge
NovGauge
NovGauge evaluates paper similarity along three dimensions: task, problem,
and method. This dataset contains final benchmark labels and bibliographic
metadata for pairwise classification and multi-paper grouping.
Dataset repository: ZitaGo/NovGauge.
Release contents
File
Records
Contents
data/positives.json
463
Paper pairs similar in at least one annotated dimension
data/negatives.json
156
Paper pairs dissimilar in the specified annotated… See the full description on the dataset page: https://huggingface.co/datasets/ZitaGo/NovGauge.lrs-wiki
Less Retarded Wiki
This is online at http://www.tastyfish.cz/lrs/main.html.
Wiki about less retarded software and related topics.
By contributing you agree to release your contribution under CC0 1.0, public domain (https://creativecommons.org/publicdomain/zero/1.0/). Please do not add anything copyrighted to this Wiki (such as copy pasted texts from elsewhere, images etc.).
Start reading at the main page.
test_zit_datasetLaMa_ZITS_cptinker-cookbook-2026-04-07zitiINSTARAW_zit_char_lora-datasetZITout_zitongzitong_lima_nostepsDMSPzitwastezitong_limazitong_lima_no_stepLEAP
LEAP Dataset
This dataset is used for training RoboFarseer(https://arxiv.org/abs/2509.25852), a Vision-Language Model (VLM) based robot task planner. The dataset is converted from human demonstration videos collected using UMI (Universal Manipulation Interface) gripper, and follows the standard Visual Question Answering (VQA) format.
Dataset Overview
The dataset contains three types of annotations for training different model capabilities:
Plan: Given the current scene… See the full description on the dataset page: https://huggingface.co/datasets/zitong86/LEAP.tinker-cookbook-2026-01-23cotmatheit_debug_zitongassignment4-preference-dataset-pairrmmizo_woman_zitentigraph-qasftassignment3-curated-datasetdas_datasetsassignment4-preference-dataset-llm-judgetestzitworkflowzit_filestwitter-ygzs88-2025.02.24-1894041106603454745-aH7b_ZItcVmFpy6a-part1
