datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineWeb2-mds-tokenized-4096gift-pretrain-full-4096
gift-pretrain-full-4096
Full counterpart to jeremycochoy/gift-pretrain-small-4096:
every series of every arrow file in
Salesforce/GiftEvalPretrain,
cropped into non-overlapping 4096-point windows and globally
shuffled. The small companion sub-samples 10 series per sub-dataset;
this one keeps everything.
6,376 source arrow files across 152 sub-datasets fully consumed
42,571,692 windows of length 4096 (float32)
4,274 parquet shards, ~619 GB total (zstd)
Layout
.
├──… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/gift-pretrain-full-4096.FineWeb2-mds-tokenized-v2-4096Cement-4096Cold-Large-32-4096multihop_qa_sft_doc4096_seq1024_v2bioasq_bier_4096_1m
BioASQ-BEIR Vector Database Dataset (4096d, 1M)
Generated embeddings dataset for vector database training and evaluation.
Dataset Summary
This dataset contains 1,000,000 text samples with vector embeddings (4096 dimensions)
generated from the BioASQ-BEIR dataset using Qwen/Qwen3-Embedding-8B.
Dataset Structure
Base dataset: 1,000,000 samples with embeddings
Embedding dimension: 4096
Repository Structure
parquet/
base.parquet - Main… See the full description on the dataset page: https://huggingface.co/datasets/maknee/bioasq_bier_4096_1m.obelics_100k-tokenized-4image_llava_vicuna-7B_4096tokenized-falcon2-dutch-4096obelics_100k-tokenized-4image_llava_vicuna-13B_4096bioasq_bier_4096_100k
BioASQ-BEIR Vector Database Dataset (4096d, 100k)
Generated embeddings dataset for vector database training and evaluation.
Dataset Summary
This dataset contains 100,000 text samples with vector embeddings (4096 dimensions)
generated from the BioASQ-BEIR dataset using Qwen/Qwen3-Embedding-8B.
Dataset Structure
Base dataset: 100,000 samples with embeddings
Embedding dimension: 4096
Repository Structure
parquet/
base.parquet - Main… See the full description on the dataset page: https://huggingface.co/datasets/maknee/bioasq_bier_4096_100k.lllavasae_obelics100k-tokenized-4096_allMiriad_Pubmed_metadata_4096_chunksdclm-replay.seq-4096.tokens-32B2^35 tokens of replay data from DCLM-baseline, concatenated into 2^23 sequences of 4096 tokens each with <|endoftext|> separators.
tulu-v3.1-mix-preview-4096-OLMoE
OLMoE SFT Mix
The SFT mix used is an expanded version of the Tulu v2 SFT mix with new additions for code, CodeFeedback-Filtered-Instruction, reasoning, MetaMathQA, and instruction following, No Robots and a subset of Daring Anteater.
Please see the referenced datasets for the multiple licenses used in subsequent data.
We do not introduce any new data with this dataset.
Config for creation via open-instruct:
dataset_mixer:
allenai/tulu-v2-sft-mixture-olmo-4096: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v3.1-mix-preview-4096-OLMoE.llavasae_obelics3k-tokenized-4096_4imagent_multispecies_4096plantcad2-c4096llama2-german-corpus-tokenized-llama-chunk-4096
Dataset Card for "llama2-german-corpus-tokenized-llama-chunk-4096"
More Information needed
distill-r1-qwen-1.5b-aime-24-4096-with-labels-prmtrain_group_theory_cpt_chunked_4096merged-dataset-4096distill-r1-qwen-1.5b-hmmt-feb-25-4096-with-bt-model-with-sigmoiddistill-r1-qwen-1.5b-aime-24-4096-with-bt-model-wout-sigmoidscience-qa-hard-neg-think-doc4096-seq1024-v2opengenome2-metagenomes-plantcad2-c4096
OpenGenome2 Metagenomes PlantCAD2 Subset (4096bp)
This dataset is a curated subset of arcinstitute/opengenome2
designed for comparative spectral analysis with plant genomic data.
Dataset Description
Sequences were randomly sampled from OpenGenome2, filtered and truncated to match the sample sizes
per split of the plantcad/Angiosperm_65_genomes_8192bp dataset.
Processing Steps
Streaming: Records were streamed from the metagenomes subfolder… See the full description on the dataset page: https://huggingface.co/datasets/plantcad/opengenome2-metagenomes-plantcad2-c4096.BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65
fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65
Packed pretraining corpus, 2,050,000 rows x 4096 tokens = 8.397B tokens, tokenized with
fhai50032/QTK-81K.
Format
column
type
notes
label
list<int32>
exactly 4096 tokens, no padding
raw_label
string
label decoded back to text (redundant, for inspection)
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65.data-artifacts-4096English_corpus-4096-packed-qtk-2.0M
fhai50032/English_corpus-4096-packed-qtk-2.0M
Packed pretraining corpus, 2,004,054 rows x 4096 tokens = 8.209B tokens, tokenized with
fhai50032/QTK-81K.
Format
column
type
notes
label
list<int32>
exactly 4096 tokens, no padding
raw_label
string
label decoded back to text (redundant, for inspection)
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/English_corpus-4096-packed-qtk-2.0M.distill-r1-qwen-1.5b-hmmt-feb-24-4096-with-bt-model-wout-sigmoid
