jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens
Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3-8B tokenizer Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1 (High-Quality-DQA subset) and tokenized with the Qwen/Qwen3-8B tokenizer. In the source data each row is a web document whose tail carries synthetic QA pairs marked Question: / Answer:. Here that document is split into its original prose (context) and the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens.
Upload README.md with huggingface_hub
Upload neg_pairs/part_000005.parquet with huggingface_hub
Upload neg_pairs/part_000004.parquet with huggingface_hub
Upload neg_pairs/part_000003.parquet with huggingface_hub
Upload neg_pairs/part_000002.parquet with huggingface_hub
Upload neg_pairs/part_000001.parquet with huggingface_hub
Upload neg_pairs/part_000000.parquet with huggingface_hub
Add files using upload-large-folder tool
initial commit
