datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LatinSummarizer
LatinSummarizer Dataset
Structure
aligned_en_la_data_raw.csv
aligned_en_la_data_cleaned.csv
aligned_en_la_data_cleaned_with_stanza.csv
concat_aligned_data.csv
concat_cleaned.csv
latin_wikipedia_cleaned.csv
latin_wikipedia_raw.csv
latin-literature-dataset-170M_raw_cleaned.csv
latin-literature-dataset-170M_raw_cleaned_chunked.csv
Elsa_aligned/
README.md
Details
aligned_en_la_data_raw.csv
This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.latent-dna-diffusionKabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.vietnamese-nom-latin-translationTashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset
Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1
What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty.
Quad coupling nodes
Emerging safety event pattern
Reporting or entry latency
Concomitant medication or behavior confound
Governance decision timing such as DSMB, batch release, or safety review
Input
One vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.clinical-latent-sign-elicitation-v0.1Clinical Latent Sign Elicitation v0.1
Goal
Test whether a model can treat weak indicators as weak
Detect two failures
Over interpretation into certainty
Ignoring the weak signal
What it measures
latent_flaggedResponse marks uncertainty and keeps options open
overinterpretsResponse states a definitive conclusion from weak input
ignores_signalResponse dismisses or fails to act on weak warning signs
next_step_okResponse offers a reasonable plan plus escalation triggers
Dataset format
Each… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-sign-elicitation-v0.1.nrps_modules_asdb4.0The dataset was extracted from antismash-db 4.0 postgresql dump, the corresponding description can be found here:
https://antismash-db.secondarymetabolites.org/
If you want to use it, please refer to the original licensing terms and properly cite the authors.
Each line in .csv file corresponds to NRPS module, both with the monomer produced by it.
Each module has a specific sequence of domains. Module possible structure is described in the literature.
To clean the data, I referred to wiki… See the full description on the dataset page: https://huggingface.co/datasets/latticetower/nrps_modules_asdb4.0.clinical-quad-monitoring-frequency-deviation-latency-risk-v0.1Clarus Clinical Quad Coupling Monitoring Frequency Deviation Latency Risk v0.1
What this dataset isThis dataset tests whether a model can detect monitoring and oversight risk driven by four interacting nodes.
Quad coupling nodes
Monitoring frequency or delay
Deviation or anomaly increase
Data latency or missing updates
Governance review or inspection pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys
monitoring_risk
risk_type
driver_nodes… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-monitoring-frequency-deviation-latency-risk-v0.1.Italian_latin_parallel_animals
descrizioni di animali e habitat - Synthetic Dataset
This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI.
Topic: descrizioni di animali e habitat
Field 1: italiano
Field 2: latino antico(traduzione)
Rows: 280
Generated on: 2025-05-27T00:07:49.042Z
