sada
Datasets
All datasets matching “sada”CryoLithe-training-datasetThe training Dataset for CryoLithe Models
The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058
Icecream reconstructions were obtained by splitting across angles.
Whenever available, we also provide dose-fractionated tilt series.
Dataset format:
Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles.
Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.SADA22
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.sada-train-preprocessedsada-train-wav2vec2-xls-r-300m-ar-preprocessedsada_preprocessed_whisper_smallJitOPD-OpenR1-Math-220k-Teacher-Memory
JitOPD OpenR1-Math-220k Teacher Prefix-Logit Memory
This dataset contains sparse teacher next-token logits collected for JitOPD
retrieval-augmented decoding. The source prompts are the default configuration
of open-r1/OpenR1-Math-220k,
and the teacher is
Qwen/Qwen2.5-Math-7B-Instruct.
Only teacher trajectories whose final boxed answer passes both a numeric
signature prefilter and Math-Verify are retained. This release contains raw
teacher prefix/logit memory and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/sadadasdasdas/JitOPD-OpenR1-Math-220k-Teacher-Memory.
