datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese_modern_poetry
简介
数据集包括了近现代的中国诗人及外国诗人(中译版)作品,所有作品著作权归原作者所有,侵删请联系aa531811820@gmail.com
chinese_poems.jsonl为原数据,training_imagery2-5_maxlen256.json 分别是根据2-5个关键意象生成诗歌的相关数据集
数据来源于网络,包括但不限于
https://github.com/sheepzh/poetry
https://bedtimepoem.com/
https://poemwiki.org/
baidu、google、zhihu等
一些作品
使用此数据集训练ChatGLM、LLaMA7b模型生成的诗歌,更多诗歌查看poems目录
IEA_Energy_Dataset
IEA_Energy_Dataset
Dataset Details
Dataset Description
The dataset is energy-related, covering topics of Oil, Coal, Wind, Hydrogen, Bioenergy, Electric vehicles, Heating, Building envelopes, Methane abatement and Chemicals.
Dataset Creation
Source Data
The dataset sources are reports from the webiste of International Energy Agency(IEA).
Data Collection and Processing
We scraped free open reports from IEA's website. The reports… See the full description on the dataset page: https://huggingface.co/datasets/Zihao-Li/IEA_Energy_Dataset.ru-bank-ie
pymlex/ru-bank-ie
Russian bank client information extraction benchmark with coverage-validated text-to-JSON pairs.
Each example contains a chat-style client message, a gold BankClientExtraction JSON object,
and a separate validation_json coverage justification. Fields may be null when absent from the source text.
Columns
id — sample identifier
reasoning — model planning before the client message
text — client message used for evaluation
gold_json — gold… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/ru-bank-ie.lavita-ChatDoctor-HealthCareMagic-100kAva-100
Empowering Agentic Video Analytics Systems with Video Language Models
[🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset]
Introduction
AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.letras-carnaval-cadiz
Dataset Card for Letras Carnaval Cádiz
English |
Español
Changelog
Release
Description
v1.0
Initial release of the dataset. Included more than 1K lyrics. It is necessary to verify the accuracy of the data, especially the subset midaccurate.
Dataset Summary
This dataset is a comprehensive collection of lyrics from the Carnaval de Cádiz, a significant cultural heritage of the city of Cádiz, Spain. Despite its… See the full description on the dataset page: https://huggingface.co/datasets/IES-Rafael-Alberti/letras-carnaval-cadiz.IELTS-writing-feedback-reasoning
Dataset Card for IELTS Writing Task 2 – Reasoning-Based Evaluation Dataset
Dataset Summary
This dataset is an IELTS Writing Task 2 automated scoring and feedback dataset based on explicit reasoning. It contains writing prompts, student essays, and a complete scoring process with professional-grade feedback generated by GLM-4.7, one of the top-tier Large Language Models (LLMs) in the current open-source ecosystem known for its strong reasoning capabilities.
Unlike… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/IELTS-writing-feedback-reasoning.writing9-ielts-essays
writing9 IELTS Essays (with band scores)
163,575 IELTS Writing essays with their overall band and the four sub-criteria bands,
crawled from writing9.com. Intended for training/evaluating automatic
IELTS Writing scorers (band regression/classification).
Splits
Band-stratified 70/30 split, fixed for reproducibility:
split
examples
description
train
114,504
real crawled essays (training portion)
test
49,071
real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.FinTalk-19k
Dataset Card for FinTalk-19k
Dataset Description
Dataset Summary
FinTalk-19k is a domain-specific dataset designed for the fine-tuning of Large Language Models (LLMs) with a focus on financial conversations. Extracted from public Reddit conversations, this dataset is tagged with categories like "Personal Finance", "Financial Information", and "Public Sentiment". It consists of more than 19,000 entries, each representing a conversation about financial topics.… See the full description on the dataset page: https://huggingface.co/datasets/ceadar-ie/FinTalk-19k.iec61131-3-st-clean-augment
ST-Coder: Multi-Source IEC 61131-3 Dataset
This dataset is designed for training Large Language Models (LLMs) to generate and analyze Structured Text (ST) code according to the IEC 61131-3 industrial standard.
📊 Dataset Subsets
This repository provides multiple configurations based on the source and processing method:
1. usecomplier
Source: Local golden datasets processed via the AST-Augmentation Factory.
Content: High-quality, syntactically correct ST code… See the full description on the dataset page: https://huggingface.co/datasets/RnniaSnow/iec61131-3-st-clean-augment.IELTs-Speaking-answer
Overview
This dataset consists of 2 json files named 'ielts_new.json' and 'ielts_old.json', which contain ielts questions and its corresponding answers for part 1 and part 2.
'ielts_new.json': new IELTs topics for 2024 September-December.
'ielts_old.json': remained IELTs topics for 2024 September-December.
Quality
Since the dataset is analysed and generated by ChatGPT based on my own pdf file, the answer may be incomplete(only part of the sentence is extracted, leading to… See the full description on the dataset page: https://huggingface.co/datasets/qwertyuiopasdfg/IELTs-Speaking-answer.theoria-hle-audit
Theoria — HLE-Verified Gold Audit
Per-problem verification artifacts from the Theoria audit on HLE-Verified Gold
(text-only), keyed by hle_id: the HLE-Verified question/answer text, Theoria's solver
answer, the formalized proof (a typed state-transition witness), per-step judge verdicts,
two independent grader verdicts (key_match + dispute_category), the raw per-call
prompts/responses, and the author's manual adjudication notes on disputed cases.
Canary — do not train:… See the full description on the dataset page: https://huggingface.co/datasets/i-eat-food/theoria-hle-audit.iecc-climate-zone-by-county
IECC/Building America climate zone by U.S. county, 2021 code cycle
Canonical, always-current version: https://referencesource.org/iecc-climate-zone-by-county/
Machine-readable: https://referencesource.org/iecc-climate-zone-by-county/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2028-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 3134
Which IECC climate zone (1-8, with… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/iecc-climate-zone-by-county.ies4-turtle-instruct
IES4 Turtle Instruct — training data + eval harness
Instruction pairs for text -> IES4 (UK Gov Information Exchange Standard) RDF/Turtle, built
correct-by-construction with telicent-ies-tool and double-validated against the published
dstl/IES4 ontology. Companion dataset to fabsssss/qwen3-coder-30b-a3b-ies4.
ies/ — 1,799 IES pairs + refusal/boundary pairs + OOD test (MIT; ontology © Crown
copyright Dstl, MIT licence)
multistd/ — 450 ontology-conditioned extraction pairs derived… See the full description on the dataset page: https://huggingface.co/datasets/fabsssss/ies4-turtle-instruct.necva__IE-cont-Llama3.1-8B-details
Dataset Card for Evaluation run of necva/IE-cont-Llama3.1-8B
Dataset automatically created during the evaluation run of model necva/IE-cont-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/necva__IE-cont-Llama3.1-8B-details.iemocap_splitLithium-Battery-IE-Dataset
Lithium-Ion Battery Patent Technical Indicator Dataset (锂离子电池专利技术指标精标数据集)
Introduction (简介)
This repository provides a highly specialized, bilingual (Chinese & English) instruction-tuning dataset designed for Fine-grained Information Extraction (IE) from Lithium-ion battery patents. It is the official data repository for our data paper: [A Dataset of Fine-Grained Technical Indicators from Lithium-Ion Battery Patents for Instruction Tuning of Large Language Models].… See the full description on the dataset page: https://huggingface.co/datasets/SongKun909/Lithium-Battery-IE-Dataset.icml2026-IelAHU5MVz-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
AIVision360-8k
Dataset Card for AIVision360-8k
Dataset Description
AIVision360 is the pioneering domain-specific dataset tailor-made for media and journalism, designed expressly for the instruction fine-tuning of Large Language Models (LLMs).The AIVision360-8k dataset is a curated collection sourced from "ainewshub.ie", a platform dedicated to Artificial Intelligence news from quality-controlled publishers. It is designed to provide a comprehensive representation of AI-related… See the full description on the dataset page: https://huggingface.co/datasets/ceadar-ie/AIVision360-8k.IEC_61131-3_STairlineSFT_Alldataset-whisper-carnavalIEInstructfortune_telling_twIEDB_B_cellTrafficData-PCielts-writing-feedback1IEFeedbackmal-google-lastlinepython-function-examples
