datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TelAgentBench-ID
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems
📌 Dataset Summary
TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.Benchmarks_CyberSec_RedSageMCQ
Dataset Card for RedSage-MCQ
Dataset Summary
RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM".
The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.riskroll-sec-10k-10q-sections
Riskroll: SEC 10-K and 10-Q sections as clean text
Need it fresh, filtered or via API? This free file is a snapshot (10-K/10-Q sections up to the last refresh), last updated 2026-09-25.
Insidewell on Apify ($0.004 per insider transaction): pulls today's SEC Form 4 trades for your own watchlist, filtered by buy/sell and size, with cluster-buy alerts on a schedule.
Get an email when this dataset updates: free, double opt-in, unsubscribe any time.
Information only, not… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/riskroll-sec-10k-10q-sections.engram-eval
engram evaluation data
The evaluation data behind Typed Decisions in Agent Memory: Where They Help, Where They Don't, and What It Costs
(Rishabh Sharma, 2026, doi:10.5281/zenodo.22948964; version 1: doi:10.5281/zenodo.22941758): update sets
that extend LoCoMo with fact changes, labeled contradiction pairs, the relation decisions escalated to an LLM,
and every scored answer from the paper's runs. Code: the engram repository (bench/make_hf_dataset.py builds this
directory from the… See the full description on the dataset page: https://huggingface.co/datasets/ris3abh-11/engram-eval.longcovid-risk-eventtimeseries
Citation
If you find this dataset or our work useful in your research, please consider citing:
Jing Wang, Amar Sra, Jeremy C. Weiss. Active Learning for Forecasting Severity among Patients with Post Acute Sequelae of SARS-CoV-2. arXiv:2506.22444, 2025.
BibTeX:
@misc{longcovid,
title = {Active Learning for Forecasting Severity among Patients with Post Acute Sequelae of SARS-CoV-2},
author = {Jing Wang and Amar Sra and Jeremy C. Weiss},
year = {2025},
eprint = {2506.22444}… See the full description on the dataset page: https://huggingface.co/datasets/juliawang2024/longcovid-risk-eventtimeseries.
