datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.persian-natural-fluently
Persian scientific dataset
I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology.
The content of the dataset has been generated by : Human, Grok3, DeepSeek R1.
License
This dataset is licensed under apache-2.0.
ArsPoetica
Ars Poetica
The Ars Poetica dataset is a collection of Russian-language poetry from the 19th and 20th centuries, annotated with stress marks. This dataset is designed to support research in generative poetry, computational linguistics, and related fields.
Stress marks were automatically assigned using the RussianPoetryScansionTool library. While the dataset has undergone selective manual validation, users should be aware of potential inaccuracies due to the automated process.… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/ArsPoetica.Nemotron-RL-Instruction-Following-Citation-Formatting-v1
Dataset Description:
Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations.
This dataset is ready for commercial/non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
Created on: April 10, 2026
Last Modified on: April 10, 2026
Version:
Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1
Dataset Description:
Teaches the model to follow arbitrary text formatting instructions (bullet styles, numbering, delimiters, heading formats, inline emphasis, web-answer structure, etc.) for targeted chat behaviors. Uses explicit Regex and string matching for the reward signal.
This dataset is ready for commercial or non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
Created on: April 10, 2026
Last Modified on: April… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.Kisna-Sovl-UncleanedThis is a Raw Uncleaned Kisna Sovl Dataset
It contains 1000 question and answering for Kisna's life and dataset.
Entire memory is fictional in nature, don't think too deeply about the implication please.
Original Character Card by Bronya Rand.
. . .
Yeah... That, Bronya, I guess...
. . .
Will Clean Up Later...
Nemotron-Math-Proofs-v1
Nemotron-Math-Proofs-v1
Paper: Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode SupervisionCode: https://github.com/NVIDIA/NeMo-SkillsDocumentation: Nemotron-MathProofs-v1 documentation
Dataset Description:
Nemotron-Math-Proofs-v1 is a large-scale mathematical reasoning dataset containing ~580k natural language proof problems, ~550k formalizations into theorem statements in Lean 4, and ~900k model-generated reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-Math-Proofs-v1.Kisna-Assist-1kThis is the 'Assist' dataset for Kisna Kaaalana (original character by Bronya Rand... The Supreme Guardian of Belobog, apparently.)
It's about the same as Viel-Lite dataset. Only with Kisna's personality instead of Viel
It's smaller than Viel as this dataset is generated with my laptop using Sao10K/L3-8B-Stheno-v3.2
But hey, quality over quantity~
. . .
(Yeah that as a cope)
Will try making Kisna AI with this dataset next.
Also I ditched alpaca formatting. So there's that.
Uses the message… See the full description on the dataset page: https://huggingface.co/datasets/ArsParadox/Kisna-Assist-1k.Kisna-Assist-UncleanedThis is a Raw Uncleaned Kisna Assistant Dataset
It contains 1000 question and answering in Kisna's character
Original Character Card by Bronya Rand.
. . .
Yeah... That, Bronya, I guess...
. . .
Will Clean Up Later...
Kisna-Sovl-1kThis is the 'Sovl' or 'Personality' type dataset for Kisna
It contains 1000-ish question and answering of personal question of Kisna Kaalana
The questions were generated by Deepseek v3 and answered by Sao10K/L3-8B-Stheno-v3.2 using my local machine.
The point is creating an AI that has their own background, bias, personality, likes, dislikes, quirks, manners, and desires.
It's untested I do not know what the result of training any AI using this dataset.
This is my attempt to create an AI that… See the full description on the dataset page: https://huggingface.co/datasets/ArsParadox/Kisna-Sovl-1k.
