datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Follow-Line-Combine-Datasetline-msg-fact-check-tw
Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies
The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community.
Attribution to the Community
This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers.
To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.egokei-linesJIC-VQA
JIC-VQA
Dataset Description
Japanese Image Classification Visual Question Answering (JIC-VQA) is a benchmark for evaluating Japanese Vision-Language Models (VLMs). We built this benchmark based on the recruit-jp/japanese-image-classification-evaluation-dataset by adding questions to each sample. All questions are multiple-choice, each with four options. We select options that closely relate to their respective labels in order to increase the task's difficulty.
The… See the full description on the dataset page: https://huggingface.co/datasets/line-corporation/JIC-VQA.pretty-line-494f34
pretty-line-494f34
Synthetic weather test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/graniteVault/pretty-line-494f34.LinearEquationsThe linear equations in this dataset are in the form:
zy + ay + b + n = py + dy + c + r
with integer coefficients ranging from -10 to 10.
small-line-9ec53f
small-line-9ec53f
Synthetic sensors test data: 44 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Prism-James79/small-line-9ec53f.three_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。
Llama v2向けのプロンプトを追加して成形してあります。
学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。
Number of rows: 3,907
Datasetは以下のリポジトリを利用してscrapeしました。
git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git
employed-population-below-international-poverty-line-for-african-countries
Employed Population Below International Poverty Line for African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/employed-population-below-international-poverty-line-for-african-countries.employed-population-below-international-poverty-line-female-for-african-countries
Employed Population Below International Poverty Line Female for African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/employed-population-below-international-poverty-line-female-for-african-countries.command-line-suggestionsunl_tesis_linea_investigacionline-level-code-vs-text-classificationThis dataset was created for SemEval-2026 Task 13, which focuses on distinguishing machine-generated code from human-written code across multiple programming languages and domains.
While the original SemEval task operates at the code snippet level, this dataset provides line-level annotations that enable finer-grained analysis of how code-like and text-like content is distributed within mixed inputs. The dataset is intended to support research in machine-generated code detection, robust… See the full description on the dataset page: https://huggingface.co/datasets/violetakastreva/line-level-code-vs-text-classification.LinearEquationTrainingDatapopulation-below-international-poverty-line-for-african-countries
Population Below International Poverty Line for African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/population-below-international-poverty-line-for-african-countries.between-the-line
暗流 Between the Line
暗流(Between the Line) 是一个中文攻击性 / 毒性短文本标注数据集,用于训练与评测「守望」微博理性发言主动干预插件的检测模型(望潮 TideWatcher)。名字取"字里行间"之意:攻击与恶意往往藏在字里行间。
1. 解决什么问题
中文互联网的攻击性发言检测缺少高质量的短句标注数据。公开的冒犯性语料覆盖有限,对缩写攻击(nmsl/cnm)、谐音、阴阳怪气、玩梗边界、冷暴力句式等"表达代差"覆盖不足。暗流以微博语境的短句为主,补充这些表达变体,为攻击性文本检测模型的训练与评测提供标注数据。
2. 主要内容
数据集共 2,414 条(train 1,936 + test 239 + val 239),本次发布 train(1,936 条)、test(239 条)与 val(239 条) 三个部分。
文件
条数
说明
train.csv
1,936
训练集(label 1: 976 / 0: 960,含 387… See the full description on the dataset page: https://huggingface.co/datasets/RainbowLIght/between-the-line.primevul-for-linevuloriginal dataset: https://huggingface.co/datasets/colin/PrimeVul
this dataset is created by:
filter out 26% records with func > 512 tokens
filter project record has < 2 samples
duplicate vul records and split dataset into train, val, test to match distribute ratio in BigVul dataset
employed-population-below-international-poverty-line-male-for-african-countries
Employed Population Below International Poverty Line Male for African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/employed-population-below-international-poverty-line-male-for-african-countries.vulnerable_1_linehomeric-speech-narrative-by-linealgebra__linear_1d_testnba-lineup-spacing-coherence-risk-v0.1What this repo is for
Detect when a lineup breaks spacing.
Focus
shooting gravity
lane congestion
creator space
corner threat
help commitment
Why it matters
Bad spacing kills offense.
Teams feel it before numbers show it.
sentiment_lines.csv
Short Sentiment Lines v1
Short labeled sentences for basic sentiment classification.
Intended Use
Sentiment analysis and educational research.
License
CC-BY-4.0
Linear1d_distilation_testinghandsome_jack_linesyahoo-answers-3k-linestest-512-linesfollow-line-dataset-v3lineage-name-mapperexemplo
