datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nvidia-nemotron-model-reasoning-dataset-turkish
Nemotron Reasoning Challenge - Turkish
Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge
Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input.
Category
Rows
Description
bit
1602
Hidden bit manipulation rule on 8-bit binary numbers
grav
1597
Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.Italian_reasoning_dataset
Quiz ed enigmi logici in italiano - Synthetic Dataset (Partial Data)
Note: Dataset generation aimed for 1100 rows, but 985 rows were successfully produced. This may be due to model output characteristics, filtering of incomplete items, or automatic correction of JSON key names.
This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI.
Topic: Quiz ed enigmi logici in italiano
Field 1: domanda e risposta corretta delimitati da tag
Field 2: Il pensiero… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Italian_reasoning_dataset.reasoning-persona-dataset_test
Reasoning + Persona SFT Dataset
Columns: instruction, input, output, persona, reasoning_summaryUse: Supervised fine-tuning for cinematic/storytelling or creative-director style outputs.
Schema
instruction (str)
input (str)
output (str)
persona (str)
reasoning_summary (str, brief rationale cue)
Citation
Author: saravan
