datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese_modern_poetry
简介
数据集包括了近现代的中国诗人及外国诗人(中译版)作品,所有作品著作权归原作者所有,侵删请联系aa531811820@gmail.com
chinese_poems.jsonl为原数据,training_imagery2-5_maxlen256.json 分别是根据2-5个关键意象生成诗歌的相关数据集
数据来源于网络,包括但不限于
https://github.com/sheepzh/poetry
https://bedtimepoem.com/
https://poemwiki.org/
baidu、google、zhihu等
一些作品
使用此数据集训练ChatGLM、LLaMA7b模型生成的诗歌,更多诗歌查看poems目录
TCM-Ancient-Modern-Open
TCM Ancient-to-Modern Chinese Open Metadata
Chinese documentation | Code repository
This public metadata companion was designed to avoid redistribution of third-party published text from the controlled TCM Ancient-to-Modern Chinese parallel corpus. It releases the complete 9,610-record index, frozen split assignments, source-edition register, length metadata, cryptographic integrity digests, and ten readable demonstration pairs. It does not redistribute the experimental… See the full description on the dataset page: https://huggingface.co/datasets/xs12345/TCM-Ancient-Modern-Open.repro-conservation-laws-for-modern-neural-architectures-traces
Agent traces
Agent sessions published from a Trackio Logbook.
aozora-bunko-modern-onlymalaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend
malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend_adolescent_abortion_adolescent_anc_phc_orc_patient_types_nurses_reg
A merged, ShareGPT-formatted Nepali (नेपाली) health-statistics instruction dataset, combining 11 individual government health data sources into one JSONL file (332 rows).
Short name used in this repository: merged_new_health_332_sharegpt_ne
Full combined name (all 11 source datasets joined): see title above.… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend.modern-webapp-instructionsPerfeito 😄📦
Então aqui está um README.md completo, moderno e pronto para subir no Hugging Face Hub, alinhado com tudo que a gente construiu (modelos pequenos, eficiência, Unsloth, apps web modernos).
Você pode copiar e colar direto.
# modern-webapp-instructions
## 📌 Overview
**modern-webapp-instructions** é um dataset de *instruction tuning* focado na criação de **webpages e aplicações modernas**, otimizado para **modelos pequenos e eficientes** (1.5B–3B), como Qwen, Phi e LLaMA… See the full description on the dataset page: https://huggingface.co/datasets/SpaceGhost/modern-webapp-instructions.modern-embedding-bench
Modern Embedding Bench
Modern Embedding Bench evaluates embedding models on practical retrieval tasks
that show up in current AI systems but are often under-covered by broad
leaderboards. The focus is on agent memory, tool and document retrieval,
long-context RAG, cross-lingual technical retrieval, coding-oriented retrieval,
and multimodal search rather than a single aggregate score.
The companion leaderboard Space is available at:… See the full description on the dataset page: https://huggingface.co/datasets/zc277584121/modern-embedding-bench.gutenberg-moderne-dpo
Gutenberg-Moderne DPO
A DPO dataset meant to enhance the writing capabilities of LLMs using public domain books from Project Gutenberg.
Inspired by Jon Durbin's Gutenberg DPO dataset: jondurbin/gutenberg-dpo-v0.1
Process
Various books were selected from Project Gutenberg for their "modern" (early 20th century) language and concise prose.
The process is nearly identical to nbeerbower/gutenberg2-dpo except that the cleaning of the original text was improved and backported… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/gutenberg-moderne-dpo.ModernChinese2ClassicalChinesepashto_ai_modern_poetrylicense: mit
language:
ps
zh
tags:
pashto
translation
chat
dataset
parquet
jsonl
chinese-to-pashto
size_categories:
10K<n<100K
Pashto AI Chat & Translation Dataset
This repository contains a high-quality, structured dataset designed for training and fine-tuning Large Language Models (LLMs) in Pashto (ps), translated from comprehensive conversational and text corpora.
Dataset Structure
The dataset is provided in both Parquet and JSONL formats for optimal… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto_ai_modern_poetry.call_of_duty_modern_warfare_2_campaign_remastered_recordings_01
使命召唤:6重制版 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_1e12c0cdd1936a273c5ce0a97b81b98f
Collection: general (泛数据)
Recordings: 4
Layout: recordings/<recording_id>/<raw component>
Classical-Modernmodern_to_ancientcall_of_duty_modern_warfare_ii_recordings_01
使命召唤:19 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_f293397332f0b207310e801c119cfb00
Collection: general (泛数据)
Recordings: 5
Layout: recordings/<recording_id>/<raw component>
Classical_Modern_1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zeroyet/Classical_Modern_1.Modern-Chinese-to-Classical-ChineseAI-Generated_Chinese_Modern_Poetry
中文
这个数据集是使用DeepSeek-R1根据标题或摘要生成中文现代诗而构成的。
我们使用这个数据集训练了芝麻Zhima。芝麻是一个专注于中文现代诗创作的LLM,能根据用户指令用标题、摘要或关键词生成原创中文现代诗。
芝麻的名字来源于志摩(徐志摩)的谐音。徐志摩(1897-1931)是一位著名的中国现代诗人。
致谢
感谢modern-poetry,我们使用该项目中汇总的中文现代诗的标题和提炼出的摘要。
English
This dataset consists of Chinese modern poems generated by DeepSeek-R1 based on titles or summaries.
We use this dataset to train Zhima. Zhima is an LLM focused on Chinese modern poetry creation, capable of generating original Chinese modern poems based… See the full description on the dataset page: https://huggingface.co/datasets/Hyaline/AI-Generated_Chinese_Modern_Poetry.scicode-dsa-modernnbeerbower__Mistral-Nemo-Moderne-12B-FFT-experimental-details
Dataset Card for Evaluation run of nbeerbower/Mistral-Nemo-Moderne-12B-FFT-experimental
Dataset automatically created during the evaluation run of model nbeerbower/Mistral-Nemo-Moderne-12B-FFT-experimental
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__Mistral-Nemo-Moderne-12B-FFT-experimental-details.ancient-modern-mutualclassical-to-modernmodern-to-classical-datasetClassical-Moderntaming-modern-prometheus-assurance
Taming the Modern Prometheus — Agentic Financial Assurance Benchmark
A small, transparent benchmark for evidence-grounded agentic workflows in financial assurance.
Contents
13 labelled cases.
3 public-derived Microsoft aggregate financial checks.
10 synthetic audit, ICFR, CAM, governance, ESG, adversarial, and reproducibility cases.
Required evidence, red-team challenge, target assertion, and non-compensatory gate for each case.
The public-derived rows are based… See the full description on the dataset page: https://huggingface.co/datasets/SADHON/taming-modern-prometheus-assurance.
