CoolFace
Datasetpublic

jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3.6-tokens

Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3.6-27B tokenizer Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1 (High-Quality-DQA subset) and tokenized with the Qwen/Qwen3.6-27B tokenizer (vocab 248,320). This is a re-tokenization of the same corpus previously released with the Qwen/Qwen3-8B tokenizer. Qwen3.6 uses a different, larger vocabulary, so the old token ids are not valid for Qwen3.6 models — the QA pairs were re-extracted from the raw… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3.6-tokens.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes34downloads
Dataset Card

Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3.6-27B tokenizer

Question/answer pairs extracted from `nvidia/Nemotron-CC-v2.1` (High-Quality-DQA subset) and tokenized with the Qwen/Qwen3.6-27B tokenizer (vocab 248,320).

This is a re-tokenization of the same corpus previously released with the Qwen/Qwen3-8B tokenizer. Qwen3.6 uses a different, larger vocabulary, so the old token ids are not valid for Qwen3.6 models — the QA pairs were re-extracted from the raw High-Quality-DQA text and tokenized afresh. The extraction is byte-identical to the original on 749,986 of 750,000 shard-0 documents (validated by re-encoding every segment with the old tokenizer and comparing ids); the handful of differences are cases where the original stripped markdown emphasis markers (__objc_methname → objc_methname, a lone ** answer dropped) and this version keeps the raw text intact. Net effect: 62,451,083 pairs here vs 62,451,070 before (+13).

In the source data each row is a web document whose tail carries synthetic QA pairs marked Question: / Answer:. Here that document is split into its original prose (context) and the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords are stripped — a stored question begins at the real question text.

Layout — two tables, joined on doc_id

Context is stored once per document rather than repeated on each of its QA pairs, so nothing is duplicated. Join on doc_id to rebuild context+QA sequences.

docs/ — one row per document (7,806,965 rows)

columntypemeaning
doc_idstringsource document uuid
context_idslist<uint32>Qwen3.6 token ids of the source prose
n_context_tokensuint32len(context_ids)

qa_pairs/ — one row per QA pair (62,451,083 rows)

columntypemeaning
doc_idstringjoins to docs.doc_id
pair_idxuint16index of the pair within its document
question_idslist<uint32>Qwen3.6 token ids of the question
answer_idslist<uint32>Qwen3.6 token ids of the answer
n_question_tokensuint32len(question_ids)
n_answer_tokensuint32len(answer_ids)

Token counts

tokens
context6,231,220,908
question1,018,523,657
answer631,911,245
total7,881,655,810

Tokenized with add_special_tokens=False — no BOS/EOS and no chat template applied, so you can compose your own sequence format.

Joining

Shards are aligned: docs/part_N.parquet and qa_pairs/part_N.parquet hold the same documents, so a shard joins against its own docs table rather than the whole corpus.

Note that pyarrow's Table.join cannot be used here — its Acero engine rejects list columns as join payload (Data type list<element: uint32> is not supported in join non-key field). Use a doc_id → row map instead:

python
import pyarrow.parquet as pq

docs  = pq.read_table("docs/part_000000.parquet")
pairs = pq.read_table("qa_pairs/part_000000.parquet")

row_of  = {d: i for i, d in enumerate(docs.column("doc_id").to_pylist())}
ctx_col = docs.column("context_ids")

for p in pairs.to_pylist():
    ctx = ctx_col[row_of[p["doc_id"]]].as_py()
    seq = ctx + p["question_ids"] + p["answer_ids"]   # separators are your call

Extraction rules

Reverse-engineered from the raw text and validated against the original release:

  • —Context is everything before the first marker line, right-stripped.
  • —A marker line is Question: / Answer: with only ASCII space/tab indentation (a non-breaking-space-indented Answer: inside prose does not count) and an optional space before the colon (Question : counts).
  • —A pair is a Question: marker and the following Answer: marker; each side accumulates the marker-free lines up to the next marker. Pairs with an empty question or answer are dropped.
  • —Documents with no parseable pair are dropped from both tables.

Notes and caveats

  • —Pairs per document is usually 8, but not always (source data ranges roughly 1–17). Don't assume a fixed 8.
  • —~110 documents across the corpus produce no parseable QA pair and are excluded from both tables, so every docs row has at least one qa_pairs partner and there are no orphans.
  • —Verified: doc_id unique within each docs shard, no empty sequences, no Question:/Answer: markers leaking into stored text, all token ids within the 248,320 vocab.

Provenance

Derived from nvidia/Nemotron-CC-v2.1, a gated dataset — the original terms and license govern this derivative. See the source dataset card for its license and access conditions. Produced by preprocessing/retokenize_nemotron_qwen36.py.