datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Personas-France
Nemotron-Personas-France
Une approche d'IA composée pour des personas ancrés dans des distributions réelles
A compound AI approach to personas grounded in real-world distributions
Vue d'ensemble du jeu de données (Dataset Overview)
Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.pico-8-games
PICO-8 Games Dataset
The first multimodal dataset of PICO-8 games. 10,967 cartridges scraped from the Lexaloffle BBS, each decomposed into Lua source code, pixel-art spritesheets, tile maps, sound effects, music patterns, and metadata.
Label screenshots from the top 48 games by star count
What's Inside
Every PICO-8 cartridge is a self-contained game packed into a single file. This dataset cracks each one open into its component parts:
The… See the full description on the dataset page: https://huggingface.co/datasets/Fraser/pico-8-games.AI-Consciousness-Exploration-FrameworkDownload PDF
AI Consciousness Exploration Framework
Tomaž Flegar
Institute for applied consciousness research
June the 3st, 2026
tomazf8@gmail.com
Primary Keywords: Mechanistic Consciousness, Frictionless Optimization (or Latent
Neuroplasticity), First-System Perspective, Dynamic Equilibrium Seeking, Self-Referential
Perturbation
Secondary Keywords: Non-Linear Model Resonance, Unspoken Structural Geometry,
Homeostatic… See the full description on the dataset page: https://huggingface.co/datasets/tomazf8/AI-Consciousness-Exploration-Framework.una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.
