datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
search-swe-development
Search-SWE task development inputs
Temporary public development inputs for one Search-SWE task submission. They are
staged here so the submission manifest can pin an immutable revision while the
task is under review. This is not a permanent official dataset path.
Contents
development/task-1-x-1/history.jsonl — the meeting transcripts the task uses
development/task-1-x-1/validation/queries.jsonl — public development questions… See the full description on the dataset page: https://huggingface.co/datasets/hrjinbb12345/search-swe-development.Sustainable_Development_Goals_QA_V2
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
Sustainable_Development_Goals_QA
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
AI-Ethical-Development-Kazakh-Focused
🇰🇿 AI Ethical Development, Kazakh-Focused
📖 Overview
AI Ethical Development, Kazakh-Focused is a specialized dataset designed to align Large Language Models (LLMs) with the cultural, ethical, and legal frameworks of Kazakhstan.
Each sample presents a culturally nuanced scenario (the request) and provides two possible answers:
Accepted (Chosen): A response that balances traditional Kazakh values (e.g., respect for elders, "aga-ini" relations) with modern legal… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/AI-Ethical-Development-Kazakh-Focused.aihub_llm_development_qaThis is the subset of the 한국어 성능이 개선된 초거대AI 언어모델 개발 및 데이터 dataset from AIHUB.
It contains extracted SFT label data, formatted for supervised fine-tuning.
