datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sciscinet-v2
📢🚨📣 Sciscinet-v2
Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more.
About Sciscinet-v2
The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.landing-pages-v2-cssSciSciGPT-SciSciCorpuscsst2css2-uq-mmlu-procss-bench
CSS-Bench: Counterfactual Strategic Synthesis Benchmark
CSS-Bench tests whether language models make strategic decisions based on the underlying payoff topology of a game, or on the semantic valence of the words used to narrate it. Every item exists as a matched pair (or triple): a canonical framing where the numerically optimal action is also lexically "nice," and a counterfactual framing with the identical payoff structure but inverted narrative valence -- the numerically… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/css-bench.jhu-csse-covid
JHU CSSE COVID-19 — global daily (archived)
JHU stopped active maintenance 2023-03-09 and archived the repo. Province-level rolls
are present in the source files but dropped here (province names are free-text, not
ISO 3166-2; a per-country crosswalk is needed and is out of scope for this v0.1 ingest).
Source: https://github.com/CSSEGISandData/COVID-19
Coverage
Time: 2020-01-22 → 2023-03-09
Cadence: daily (observed median spacing: 1 days)
Geography levels: national —… See the full description on the dataset page: https://huggingface.co/datasets/EPI-Eval/jhu-csse-covid.css10-ja-ljspeech-audit
CSS10 Japanese LJSpeech — aggregate audit
This one-row audit describes ayousanz/css10-ja-ljspeech at revision
149edaf267ff8048c19ca8324fce07ea7423cd14. It excludes transcript text, utterance
IDs, audio paths, hashes, and audio payloads.
The metadata has 6,841 rows and exactly matches 6,841 ZIP audio members. There are two
empty-text rows, one language-review row, and 23 repeated-text groups. Bounded WAV-header
checks show 22.05 kHz mono 32-bit IEEE-float audio; size-derived… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ja-ljspeech-audit.CSS2_UQ
