sci-fi
Datasets
All datasets matching “sci-fi”Sci-Fi-ZH一份 VeejaLiu 正在手工清洗的数据:https://github.com/VeejaLiu/ScienceFictionCollection
SciFigPlag-Bench
SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection
🌐 Homepage |
📖 arXiv
SciFigPlag-Bench is a benchmark for provenance-aware scientific figure plagiarism detection. It evaluates whether a suspicious figure reuses evidence from a specific source figure, which figure is the original source, how the reused content has been transformed, and where the reused evidence appears.
The benchmark is designed to evaluate vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/FreeLand123/SciFigPlag-Bench.SciFigAlign
SciFigAlign
Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
Paper · Code · Dataset
Scientific figure assessment in peer review is not natural-image IQA. A figure must be legible, support the manuscript’s claims, and present a clear visual hierarchy.
This dataset is the full SciFigAlign corpus: 3,857 figure crops from 3,126 ICLR / NeurIPS / ICML papers, labeled on four peer-review dimensions (1–5): Clarity, Relevance, Informativeness… See the full description on the dataset page: https://huggingface.co/datasets/haihanlamu/SciFigAlign.SciFigQual-Bench
SciFigQual-Bench
A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
Overview · Highlights · Files · Rubric · eval1200 · Usage · Citation
Gold instances
Public test split
Rubric
SFQ-Agent MAE
Within-1
6,308
eval1200 · 1,200
5 dims · 1–10
0.418
93.4%
ACL · EMNLP · ICML · NeurIPS | 2020–2025 | 1,144 papers | ~355k citing-paragraph records
Figure 1. From isolated-figure evaluation… See the full description on the dataset page: https://huggingface.co/datasets/haihanlamu/SciFigQual-Bench.Sci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.SciFIBench
SciFIBench
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie
NeurIPS 2024
Note: This repo has been updated to add two splits ('General_Figure2Caption' and 'General_Caption2Figure') with an additional 1000 questions. The original version splits are preserved and have been renamed as follows: 'Figure2Caption' -> 'CS_Figure2Caption' and 'Caption2Figure' -> 'CS_Caption2Figure'.
Dataset Summary
The SciFIBench (Scientific Figure… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SciFIBench.
