synthetic_data
FLUX-SyntheticAnimeintfloat_multilingual-e5-large-instruct_FareedKhan_prime_synthetic_data_2k_2_4retrieval-mpnet-dot-finetuned-gpt-4o-mini-synthetic-datasetflax-sentence-embeddings_all_datasets_v4_MiniLM-L6_FareedKhan_prime_synthetic_data_2k_10_32TaylorAI_bge-micro-v2_FareedKhan_prime_synthetic_data_2k_10_32TaylorAI_bge-micro-v2_FareedKhan_prime_synthetic_data_2k_10_64bert-plus-L8-v1.0-syntheticSTS-4kBAAI_bge-m3_FareedKhan_prime_synthetic_data_2k_2_4
synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kturing-synthetic-radar-dataset
The Turing Synthetic Radar Dataset (TSRD)
Dataset Summary
The Turing Synthetic Radar Dataset is the first publicly available, comprehensively simulated pulse train dataset designed for radar pulse deinterleaving research. It provides a large-scale benchmark for developing and evaluating electronic warfare (EW) and signal intelligence (SIGINT) applications, enabling researchers to address the critical challenge of separating interleaved radar pulses from multiple… See the full description on the dataset page: https://huggingface.co/datasets/alan-turing-institute/turing-synthetic-radar-dataset.zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.VisRAG-Ret-Train-Synthetic-data
Dataset Description
This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries.
Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset.
Name
Source
Description
# Pages
Textbooks
https://openstax.org/
College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.Synthetic_Dataset_for_Stirrup_Rebar_SegmentationA Synthetic Dataset for Stirrup Rebar Segmentation
The dataset contains:
A synthetic training set of 12,000 images, and a synthetic validation set of 4,000.
A synthetic test set of 4,800 images (only top rebars are annotated).
A real-world test set of 233 images (only top rebars are annotated).
Diverse rebar specifications, stacking, lighting, distractors, and background conditions.
Before usage
mkdir -p train_syn/train2017
mv train_syn/train2017_sub{1,2,3}/*… See the full description on the dataset page: https://huggingface.co/datasets/tsrobcvai/Synthetic_Dataset_for_Stirrup_Rebar_Segmentation.ocr-synthetic-cheque-datatset
