datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multihop_qa_sft_doc4096_seq1024_v2bioasq_bier_4096_1m
BioASQ-BEIR Vector Database Dataset (4096d, 1M)
Generated embeddings dataset for vector database training and evaluation.
Dataset Summary
This dataset contains 1,000,000 text samples with vector embeddings (4096 dimensions)
generated from the BioASQ-BEIR dataset using Qwen/Qwen3-Embedding-8B.
Dataset Structure
Base dataset: 1,000,000 samples with embeddings
Embedding dimension: 4096
Repository Structure
parquet/
base.parquet - Main… See the full description on the dataset page: https://huggingface.co/datasets/maknee/bioasq_bier_4096_1m.bioasq_bier_4096_100k
BioASQ-BEIR Vector Database Dataset (4096d, 100k)
Generated embeddings dataset for vector database training and evaluation.
Dataset Summary
This dataset contains 100,000 text samples with vector embeddings (4096 dimensions)
generated from the BioASQ-BEIR dataset using Qwen/Qwen3-Embedding-8B.
Dataset Structure
Base dataset: 100,000 samples with embeddings
Embedding dimension: 4096
Repository Structure
parquet/
base.parquet - Main… See the full description on the dataset page: https://huggingface.co/datasets/maknee/bioasq_bier_4096_100k.Miriad_Pubmed_metadata_4096_chunksdclm-replay.seq-4096.tokens-32B2^35 tokens of replay data from DCLM-baseline, concatenated into 2^23 sequences of 4096 tokens each with <|endoftext|> separators.
tulu-v3.1-mix-preview-4096-OLMoE
OLMoE SFT Mix
The SFT mix used is an expanded version of the Tulu v2 SFT mix with new additions for code, CodeFeedback-Filtered-Instruction, reasoning, MetaMathQA, and instruction following, No Robots and a subset of Daring Anteater.
Please see the referenced datasets for the multiple licenses used in subsequent data.
We do not introduce any new data with this dataset.
Config for creation via open-instruct:
dataset_mixer:
allenai/tulu-v2-sft-mixture-olmo-4096: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v3.1-mix-preview-4096-OLMoE.nt_multispecies_4096plantcad2-c4096distill-r1-qwen-1.5b-aime-24-4096-with-labels-prmtrain_group_theory_cpt_chunked_4096distill-r1-qwen-1.5b-hmmt-feb-25-4096-with-bt-model-with-sigmoiddistill-r1-qwen-1.5b-aime-24-4096-with-bt-model-wout-sigmoidscience-qa-hard-neg-think-doc4096-seq1024-v2opengenome2-metagenomes-plantcad2-c4096
OpenGenome2 Metagenomes PlantCAD2 Subset (4096bp)
This dataset is a curated subset of arcinstitute/opengenome2
designed for comparative spectral analysis with plant genomic data.
Dataset Description
Sequences were randomly sampled from OpenGenome2, filtered and truncated to match the sample sizes
per split of the plantcad/Angiosperm_65_genomes_8192bp dataset.
Processing Steps
Streaming: Records were streamed from the metagenomes subfolder… See the full description on the dataset page: https://huggingface.co/datasets/plantcad/opengenome2-metagenomes-plantcad2-c4096.BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65
fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65
Packed pretraining corpus, 2,050,000 rows x 4096 tokens = 8.397B tokens, tokenized with
fhai50032/QTK-81K.
Format
column
type
notes
label
list<int32>
exactly 4096 tokens, no padding
raw_label
string
label decoded back to text (redundant, for inspection)
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65.English_corpus-4096-packed-qtk-2.0M
fhai50032/English_corpus-4096-packed-qtk-2.0M
Packed pretraining corpus, 2,004,054 rows x 4096 tokens = 8.209B tokens, tokenized with
fhai50032/QTK-81K.
Format
column
type
notes
label
list<int32>
exactly 4096 tokens, no padding
raw_label
string
label decoded back to text (redundant, for inspection)
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/English_corpus-4096-packed-qtk-2.0M.distill-r1-qwen-1.5b-hmmt-feb-24-4096-with-bt-model-wout-sigmoidHindi_corpus-4096-packed-qtk-1.37M
fhai50032/Hindi_corpus-4096-packed-qtk-1.37M
Packed pretraining corpus, 1,369,441 rows x 4096 tokens = 5.609B tokens, tokenized with
fhai50032/QTK-81K.
Format
column
type
notes
label
list<int32>
exactly 4096 tokens, no padding
raw_label
string
label decoded back to text (redundant, for inspection)
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.
How… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/Hindi_corpus-4096-packed-qtk-1.37M.LoRa_4096blockfineweb-sample-100BT_over-4096-tokensdiverseqa-hard-neg-think-doc4096-seq1024-v2slimpajama_Qwen2_tokenized_upsample_4096_chunk_256Kgeometric-vocab-4096ddistill-r1-qwen-1.5b-aime-24-4096-with-bt-model-with-sigmoidsynthetic-genomes-plantcad2-c4096
Synthetic Genomes PlantCAD2 (4096bp)
Synthetic DNA sequences from randomly initialized Qwen2 using PlantCAD2 tokenizer.
Details
Model: Qwen/Qwen2-1.5B (random weights)
Tokenizer: kuleshov-group/PlantCAD2-Small-l24-d0768
Method: Teacher-forced parallel generation (single forward pass)
Length: 4096bp
Splits
Split
Count
train
2,638,656
validation
329,832
slimpajama_llama_tokenized_upsample_4096_chunk_1MGenerated using https://github.com/FranxYao/Long-Context-Data-Engineering with the below command:
mkdir logs
mkdir data
mkdir data/slimpajama
mkdir data/slimpajama/per_source_downsample
cd data_engineering
PATH_TO_SLIMPAJAMA=rokset3/slim_pajama_chunk_1
nohup python -u slimpajama_packing.py\
--dataset_size=5b\
--print_interval=100 --num_process=200\
--chunk_size=1000001 \
--dataset_path=$PATH_TO_SLIMPAJAMA\
--output_path=../data/slimpajama/per_source_downsample/… See the full description on the dataset page: https://huggingface.co/datasets/PY007/slimpajama_llama_tokenized_upsample_4096_chunk_1M.4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivOpenR1-Math-220k_all_Llama3_4096toksRULER-4096-llama-3.2-tokenizerslimpajama_mistral_tokenized_upsample_4096_chunk_128K
