datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
connections-rl-results
connections-rl: raw evaluation artifacts
Per-puzzle records, bootstrap summaries and analysis outputs backing
connections-rl, a two-scale (Qwen2.5-1.5B / 7B), three-seed study of
what verifiable-reward RL actually transfers.
This is an artifact bundle for auditing published numbers, not a loadable
training dataset, so the dataset viewer is disabled.
Read this before using the numbers
Two conventions in these files are easy to misread. Both have bitten this
project… See the full description on the dataset page: https://huggingface.co/datasets/jacksonlukas/connections-rl-results.quran-bil-quran-connections
Dataset Card for quran-bil-quran-connections
This dataset contains a .jsonl file, where each line is a single connection between 2 verses made by Ibn Ashur in his Exegesis.
Language(s) (NLP): Arabic, English
Dataset Sources [optional]
Source Exegesis: [Tafsir al-Tahrir wa al-Tanwir]
Writeup: [Ibn Ashur's Qur’an bi’l Qur’an Visualized]
Demo: [App]
Method
For each verse/group in original source exegesis file, find all verse references using libraries like… See the full description on the dataset page: https://huggingface.co/datasets/ShahamFarooq/quran-bil-quran-connections.nyt-connections-datasets-raw
NYT Connections Raw Datasets
This repository contains the raw and formatted reasoning data for NYT Connections puzzle solving experiments. These files are the source data used to create the experiment splits in nickting/nyt-connections-experiments.
Overview
This dataset includes three types of puzzle data with AI-generated reasoning:
NYT Connections Puzzles - Authentic New York Times puzzles
Synthetic Connections Puzzles - Algorithmically generated puzzles… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-datasets-raw.nyt-connections-experiments
NYT Connections Experiments Dataset
This dataset contains training, validation, and test splits for fine-tuning language models on New York Times Connections puzzles. It includes three experimental configurations examining data augmentation, reasoning format, and curriculum learning.
Dataset Overview
NYT Puzzles: 831 total (673 training, 74 validation, 84 test)
Synthetic Puzzles: 200 total (162 training, 18 validation, 20 test)
Pre-Connections Tasks: 720 training… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-experiments.
