CoolFace
Datasetpublic

jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens

Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3-8B tokenizer Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1 (High-Quality-DQA subset) and tokenized with the Qwen/Qwen3-8B tokenizer. In the source data each row is a web document whose tail carries synthetic QA pairs marked Question: / Answer:. Here that document is split into its original prose (context) and the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes302downloads
9 commits on main
c4d8f472mo ago

Upload README.md with huggingface_hub

jackyk02
d717cfc2mo ago

Upload neg_pairs/part_000005.parquet with huggingface_hub

jackyk02
7409d352mo ago

Upload neg_pairs/part_000004.parquet with huggingface_hub

jackyk02
eb82e622mo ago

Upload neg_pairs/part_000003.parquet with huggingface_hub

jackyk02
998b66a2mo ago

Upload neg_pairs/part_000002.parquet with huggingface_hub

jackyk02
869586f2mo ago

Upload neg_pairs/part_000001.parquet with huggingface_hub

jackyk02
b74cb222mo ago

Upload neg_pairs/part_000000.parquet with huggingface_hub

jackyk02
a5a32182mo ago

Add files using upload-large-folder tool

jackyk02
63131072mo ago

initial commit

jackyk02