datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
attention-uq-800q-colab
Attention/UQ 800-question Colab bundle
A deterministic 200-question subset for each of MultiModalQA, WebQA, HotpotQA, and TAT-QA. See manifest.json for exact upstream sources, hashes, counts, and the explicitly constructed WebQA distractor setting.
mmqa-first100-colab
MultiModalQA first-100 Colab subset
This repository contains the first 100 examples of the official MultiModalQA dev split and only their referenced text, table, and image assets. It is a reproducibility artifact for Untitled34_attention_uq_100q_benchmark.ipynb.
The original dataset is from allenai/multimodalqa. See manifest.json for counts and source hashes.
Ultimate-Offensive-Red-Team_claude_mythos_distilled_25k_colabcolab-training-demo-sft
colab-training demo SFT dataset
500 synthetic two-digit addition pairs in messages (chat) format.
Generated for validating the colab_training QLoRA pipeline; after training,
ask the adapter "What is 34 + 58?" and expect "34 + 58 = 92".
slm-rl-colab-dataslm-rl-colab-dataslm-rl-colab-dataslm-rl-colab-dataslm-rl-colab-datacolabaya_dataset-colab_copyJust a copy of the aya dataset so I can use it on colab
ColabFoldQaraami-Colab-Writecolabel-copyright-substitution-riskfp_run_colabtreinoPrimeiro teste
juciJsoncolab_uploadgdrive-sbpn-fresh-diarization-colab-l4-20260813-benchmarkSFT_run_colab_1-beam-cd-50colabel-ai-task-classificationwav2vec2-base-timit-demo-google-colab_3
