CoolFace
Datasetpublic

nutshells3/livegraph-matched-benchmark

LiveGraph Matched Graph-vs-Linear Benchmark This release contains the frozen synthetic datasets used by LiveGraph: Typed Graphs versus Linear Traces. It supports a controlled comparison between a typed-graph rendering and a chronological linear-trace rendering of the same Boolean relational-propagation records. Main matched benchmark data/semantic_train.jsonl: 480 semantic training records, depths 1--10. data/semantic_heldout.jsonl: 3,200 clean held-out records… See the full description on the dataset page: https://huggingface.co/datasets/nutshells3/livegraph-matched-benchmark.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes106downloads
Dataset Card

LiveGraph Matched Graph-vs-Linear Benchmark

This release contains the frozen synthetic datasets used by LiveGraph: Typed Graphs versus Linear Traces. It supports a controlled comparison between a typed-graph rendering and a chronological linear-trace rendering of the same Boolean relational-propagation records.

Main matched benchmark

  • —data/semantic_train.jsonl: 480 semantic training records, depths 1--10.
  • —data/semantic_heldout.jsonl: 3,200 clean held-out records, depths 1--32.
  • —data/semantic_eval_variants.jsonl: 19,200 semantic evaluation variants.
  • —data/sft_{typed_graph,linear_trace}.jsonl: paired SFT rows.
  • —data/eval_{typed_graph,linear_trace}.jsonl: paired evaluation rows.
  • —manifest.json: frozen counts, construction invariants, and file hashes.

The variants are clean, edge_order, node_rename, extra_distractors, one_path_edge_flip, and edge_delete. The generator balances labels within depth and holds the semantic record and full-chat token count fixed across the two registered renderings. Ordinary visible trailing PAD markers are part of the graph rendering package; this is not a pure graph-topology intervention.

Additional frozen subsets

  • —data/grpo_v1/: exact-verifier GRPO train, ID, and OOD sets.
  • —data/native_robustness_v1/: native-thinking robustness subset.
  • —data/edge_delete_unknown_v1/: deletion-only UNKNOWN evaluation.
  • —data/subsets_v3/: clean and diagnostic subsets.
  • —data/relassum_v3/: auxiliary relation/assumption evaluation artifacts.

Integrity

Verify the frozen main files from this directory:

bash
sha256sum -c SHA256SUMS

For the original main files, the frozen hashes are also recorded in manifest.json and SHA256SUMS. Derived subset directories carry their own manifests.

The byte-frozen top-level manifest retains two original absolute tokenizer and chat-template provenance strings so that its paper-reported SHA-256 remains verifiable. They are inert local paths; use the public base model files rather than those locations.

Scope and limitations

The data are synthetic. supports preserves a Boolean value and refutes negates it along one queried path. These tasks do not establish real-world reasoning quality, theorem validity, source-claim truth, or end-to-end OAE assurance. Auxiliary relation/assumption labels are not human gold.

License and citation

The released dataset is provided under Apache License 2.0. Cite the LiveGraph paper and this dataset release. A final arXiv identifier can be added to this card after submission.