datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audioset-opus-fbanks
AudioSet Opus — INT8 log-mel filterbanks
Precomputed Kaldi log-mel filterbank features for AudioSet, quantized to INT8.
Derived from danjacobellis/audioset_opus_24kbps
and danjacobellis/audioset_opus_24kbps_balanced,
which were already deduplicated by content hash.
The point of this dataset is to remove audio decoding and filterbank computation
from the training loop. In a masked-autoencoder training step at 64×1008
geometry, the fbank front end costs ~41% of wall-clock at 1M… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset-opus-fbanks.mala-opus-dedup-2410-reLIDopus_lid_filtered
What this dataset is
This dataset is a filtered version of the MaLA-LM/mala-opus-dedup-2410 dataset after having all texts with incorrect language codes filtered out of it.
To do this, we use the cis-lmu/glotlid language ID model, predict the top probability language for both texts (source and target) and discard a row in which either text has a different language code to its original label.
How this dataset was made
from tqdm.auto import tqdm
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/opus_lid_filtered.mala-opus-dedup-2410-lid-filteredv3-2k-traj-claude-opus-4.7v4-4k-traj-claude-opus-4.7opus-5langs-1Mswebench-verified-claude-opus-4.7swebench-multilingual-claude-opus-4.7kernelbook-opus4.8-multiturn-traces
KernelBook → Triton: Multi-Turn Generation Traces (Opus 4.8)
Multi-turn agentic traces of Claude Opus 4.8 converting PyTorch modules into
Triton GPU kernels. Each row is one problem from
GPUMODE/KernelBook: the model
writes a kernel, runs it on a GPU against the reference, reads the
correctness + speedup feedback, and iterates — so every trace is a grounded,
tool-using optimization loop, not a single-shot completion.
How it was generated
Model: claude-opus-4-8… See the full description on the dataset page: https://huggingface.co/datasets/ppbhatt500/kernelbook-opus4.8-multiturn-traces.basharena_action_only_opus46_largeqgqa-gpqa-migrate-20260219-141149-claude-opus-4-6en-az-opus-filtered-parallel-corpus
Filtered EN-AZ OPUS Parallel Corpus
English–Azerbaijani parallel sentences pooled from OPUS corpora and filtered with a
two-stage quality-estimation pipeline.
Pipeline
LaBSE cross-lingual cosine similarity (kept the well-aligned pairs).
COMET-Kiwi (Unbabel/wmt22-cometkiwi-da) reference-free QE on the survivors.
Exact-pair deduplication.
Effective minimums in this release: LaBSE ≥ 0.900, COMET-Kiwi ≥ 0.900.
Columns
en_text — English (source)… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/en-az-opus-filtered-parallel-corpus.opus-es-monologues
OPUS Spanish Monologues
A dataset that captures monologues from the Spanish Open Subtitles Project dump and undergoes light cleaning. Monologues retained
in this dataset are intances in the raw .txt dump where a single speaker is uninterrupted for more than 100 words. The dataset consists of monologues from the 2013, 2016,
and 2018 OPUS Spanish monolingual datasets.
Quick Dataset Facts
Contains 1,481 documents
Each document averages ~241.2 words
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/opus-es-monologues.erdos741ii-lean4-opus-traces
Erdős #741(ii) — Lean 4 Opus Agent Traces
39 Claude Opus agent sessions attempting Erdős problem #741(ii) in Lean 4 (G1 rung: build the proof from an NL construction description).
v2 (2026-06-10) — label correction. The original upload labeled all 39 traces
reward=1.0; the labeler matched the text SCORE=1.0 anywhere in the conversation,
including the worker prompt. Rewards are now anchored to genuine oracle output
(line-anchored SCORE= adjacent to… See the full description on the dataset page: https://huggingface.co/datasets/vincentoh/erdos741ii-lean4-opus-traces.bybit-linear-perps-opusdtqgqa-unified-processed-claude-opus-4-6-20260213-033913qgqa-claude-opus-4-6-20260213-041708eleusis-opus-sft-smokeres_gptoss120b_original_1_high_0.7_16000_claude-opus-4-7audioset_mpq2_opusres_gptoss120b_original_1_xhigh_0.7_32000_claude-opus-4-7telugu-english-opus-pairsquery-evaluation-full_model_eval_claude_4_opusqgqa-testing-claude-opus-4-6query-evaluation-single_model_eval_opus_4_n6
