datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/MRSHREY197/icrm-hitek-full-db-mixed.ICMR-HITEK-FULL-MIXED-DB
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/devil-69/ICMR-HITEK-FULL-MIXED-DB.baguetterMixedSignalsDatasetMixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.HISTAI-mixedDear researchers and engineers, you're accessing a dataset that would cost millions of dollars to build and took millions of nerves to negotiate favorable terms for its use. Your support, by liking the repositories and upvoting the collection, costs nothing but gives us valuable motivation to continue our contributions to the community. We reserve the right not to approve the request if you don't support our efforts. Thank you very much for collaboration!
HISTAI Dataset
Find… See the full description on the dataset page: https://huggingface.co/datasets/histai/HISTAI-mixed.icrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/sckeptic/icrm-hitek-full-db-mixed.seq2seq-mixed-pretraining-SmolLM2icrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/sauravsingh2111/icrm-hitek-full-db-mixed.stage2_mixed_curriculum_v1
Stage 2 Mixed Text/Phoneme TTS Dataset
This dataset contains mixed text/phoneme sequences for TTS training with curriculum learning.
Curriculum Learning
The probability of converting words to phonemes increases over the dataset:
Start: p = 0.3 (more text, less phonemes)
End: p = 1.0 (all phonemes)
Transition: Linear over 500,000 rows
Each row uses p(i) for ALL its words/spaces, then i increments for the next row.
Features
Column
Description
text… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/stage2_mixed_curriculum_v1.bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.mixed_msm_evalsicrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/Craige113/icrm-hitek-full-db-mixed.vehicle-mixed-traffic-detection
Visaitech Mixed-Traffic Vehicle Detection Dataset (v0.1)
Dashcam frames annotated for pedestrian / 2-wheeler / 3-wheeler / 4-wheeler
detection in South Asian mixed traffic, a class taxonomy general-purpose
COCO-trained detectors don't cover (COCO has no concept of an auto-rickshaw
or motorcycle-vs-bicycle-as-one-class "2-wheeler" grouping tuned for how
this traffic actually mixes on the road).
This is an early v0.1 release: 293 annotated frames from 6 source videos,
published… See the full description on the dataset page: https://huggingface.co/datasets/visaitech/vehicle-mixed-traffic-detection.MixBench25MixBench is a benchmark for evaluating mixed-modality retrieval. It contains queries and corpora from four datasets: MSCOCO, Google_WIT, VisualNews, and OVEN. Each subset provides: query, corpus, mixed_corpus, and qrel splits.diffbir-mixed-setsicrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/idkdd/icrm-hitek-full-db-mixed.german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.LIVR_mixed
LIVR_mixed (v2)
Curated multimodal data for training a Qwen2.5-VL self-reflection RL pipeline,
plus the held-out LIVR splits and three external benchmarks (BLINK +
PixMo-Count + VSP) used to evaluate it.
Layout
train/ # LIVR train (9 tasks × 1000)
metadata.jsonl
livr_v2_manifest.json
images/<task>/... # ~8.7 GB
livr_eval/ # LIVR's own held-out val + test (8 tasks; counting → pixmo_count_eval)
validation/… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/LIVR_mixed.Mixed-Signals-V2X
Mixed Signals V2X: Collaborative 3D Object Detection Dataset
Point clouds and 3D bounding-box labels for the Mixed Signals dataset, a
diverse, real-world dataset for heterogeneous LiDAR V2X collaboration
(ICCV 2025). Collected at a busy intersection with 3 connected vehicles and
a roadside unit (RSU) carrying two LiDARs, for 5 LiDAR sensors per synchronized
frame.
This repository accompanies a collaborative 3D object detection competition on
Codabench, built on the
Mixed Signals… See the full description on the dataset page: https://huggingface.co/datasets/sberrio/Mixed-Signals-V2X.icrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/princewebz/icrm-hitek-full-db-mixed.mixed_inst_multi_chat
Dataset Card for "mixed_inst_multi_chat"
More Information needed
mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo.
Used to teach a model to ignore languages that are not french
CEFR_Mixed_Dataset_1PutCab-Mixed-Train100-V4
PutCab Mixed V4 Train100
100 qualified demonstrations from a common 100-scene
protocol. Three 320×240 H.264 cameras at 50/3 FPS, 250 Hz physical control,
29794 observations. Construction/packaging Run: E108-R003.
The left arm opens the drawer and the right arm grasps, lifts, transfers and
releases the object. Sequential runs the two arm programs serially, with half
of each order. Concurrent starts both together. CTR keeps frozen normalized
Delta choices and records the exact… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/PutCab-Mixed-Train100-V4.AE_english_data_stage_3-0.2M-0.4M-mixedwikipedia-embed-en-2023-11prompt-swap-mixed12-5xlr-e1-mxfp4-mergedPutCab-Mixed-Train50-V4
PutCab Mixed V4 Train50
50 qualified demonstrations from a common 100-scene
protocol. Three 320×240 H.264 cameras at 50/3 FPS, 250 Hz physical control,
14653 observations. Construction/packaging Run: E108-R004.
The left arm opens the drawer and the right arm grasps, lifts, transfers and
releases the object. Sequential runs the two arm programs serially, with half
of each order. Concurrent starts both together. CTR keeps frozen normalized
Delta choices and records the exact… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/PutCab-Mixed-Train50-V4.prompt-swap-mixed12-5xlr-e2-mxfp4-merged
