datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi3-agy-gemini3.1pro-tr87
ARC-AGI-3 tr87 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game tr87, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-tr87.arc-agi3-agy-gemini3.1pro-g50t
ARC-AGI-3 g50t — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game g50t, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-g50t.arc-agi3-agy-gemini3.1pro-su15
ARC-AGI-3 su15 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game su15, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-su15.GeminiMol-QSARtaubench-gemini-traces
taubench-gemini-traces
Complete HTTP-level agentic traces from running taubench_gemini benchmark tasks through an instrumented reverse proxy.
Each trace captures full request/response pairs including system prompts, user messages, assistant responses, tool calls and results, and token usage metadata.
Stats
Total sessions: 115
Multi-turn sessions (2+ LLM calls): 115
Total records: 5744
Total LLM requests: 2872
Format
Raw JSONL traces from the instrumented… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/taubench-gemini-traces.alpaca-cleaned-gemini-hun-ratingsEz az adathalmaz úgy keletkezett, hogy a Bazsalanszky/alpaca-cleaned-gemini-hun-n lefuttattam egy llm által támogatott értékelést.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
ko_ultrafeedback_gemini_feedbackmaywell/ko_Ultrafeedback_binarized 중 12000 여개의 chosen 을 Google Gemini Pro를 이용해서 피드백하고 점수를 평가
swe-verified-gemini3-flash-trajectories
SWE-bench Verified — Gemini-3-flash agent trajectories (graded, 3 samples/instance)
Agent trajectories from gemini-3-flash-preview (high reasoning, temperature 0.8) run with the
OpenHands agent on SWE-bench Verified, in Modal sandboxes. For each of 100 instances
we sampled multiple trajectories and graded them with the SWE-bench harness; this dataset holds the
3 graded samples per instance = 296 trajectories, 198 resolved (67%).
pass@1 ≈ 66/100, pass@3 (oracle) = 73/100.… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/swe-verified-gemini3-flash-trajectories.apollo_english_books_translated_to_dutch_with_geminiflash15
Data description
Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
sinhala-personas-lk-v0.2-gemini-1000-preview
Sinhala-Personas-LK v0.2 Gemini 1000 Preview
Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation.
Version
0.1-preview
Records
1000 synthetic records.
Language
Sinhala (si).
Country context
Sri Lanka (LK).
Important limitations
This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.persuasion-gemini-1.5-pro
Dataset Summary
This dataset was built upon the Nuclear News V4 Dataset.
The aim of this dataset is to increase the persuasiveness of nuclear domain news articles, depending on the audience, author's intention, and the original article's sentiment toward nuclear energy.
Both audiences are responding to the (usually same, if no controversy boosting was used) statement about the nuclear energy, supported by the article content.
The dataset was constructed by a multi-agent system in a… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/persuasion-gemini-1.5-pro.261_DetectionAI_Gemini2.5_datagemini-flash-2-speech-MOSS-Annotatedwildchat-oracle-questions-mini-gemini3Video-R1-p2r-gemini-video-only-2kGeminiPlusGPT
