datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SegmentScore
Dataset Card for SegmentScore
Dataset Description
This dataset contains open-ended long-form text generations from various LLM models (namely OpenAI gpt-4.1-mini, Microsoft phi 3.5 mini Instruct and Meta Llama 3.1 8B Instruct), scored for factuality using the SegmentScore algorithm and gpt-4.1-mini as the judge.
Homepage: arxiv/TBD
Repository: github.com/dhrupadb/semantic_isotropy
Point of Contact: [Dhrupad Bhardwaj, Tim G.J. Rudner]
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/dhrupadb/SegmentScore.tech-talks-segments
Awesome Tech Talks Dataset
A curated dataset of 2,200+ technical sessions, workshops, and keynotes from official engineering organizations including Google, Microsoft, OpenAI, Anthropic, and Cursor. The dataset includes video metadata, structured topic classifications, cleaned transcripts, and 42,000+ segmented text chunks designed for Retrieval-Augmented Generation (RAG), vector search, and language model evaluation.
Dataset Summary
Attribute
Value… See the full description on the dataset page: https://huggingface.co/datasets/0xShashi/tech-talks-segments.human_grch38_segment_sample
Human GRCh38 Genome Segments
Dataset Description
This dataset contains 16,384 base pair segments from the human reference genome (GRCh38) prepared for Sparse Autoencoder (SAE) training with the Evo2 model. The segments are extracted using a sliding window approach with 75% overlap.
Dataset Details
Total segments: 718,648
Segment size: 16,384 base pairs
Stride: 4,096 base pairs (75% overlap)
Source genome: GRCh38.primary_assembly (GENCODE Release 41)… See the full description on the dataset page: https://huggingface.co/datasets/harari/human_grch38_segment_sample.
