AINovice2005/carbon-likelihood-stats
Dataset Summary Dataset Summary AINovice2005/carbon-likelihood-stats is a model-enriched dataset derived from the Carbon genomic corpus, containing sequence-level likelihood statistics for 251,427 sequence regions. Each row links a source sequence record to its corresponding region in the processed/tokenized Carbon corpus using record_id, start, and end. The dataset summarizes both the model's overall likelihood of the observed sequence and the variation in… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-likelihood-stats.
Dataset Summary
Dataset Summary
AINovice2005/carbon-likelihood-stats is a model-enriched dataset derived from the Carbon genomic corpus, containing sequence-level likelihood statistics for 251,427 sequence regions.
Each row links a source sequence record to its corresponding region in the processed/tokenized Carbon corpus using record_id, start, and end. The dataset summarizes both the model's overall likelihood of the observed sequence and the variation in token-level model confidence within that sequence.
The dataset contains aggregate likelihood measures including mean_log_prob, sum_log_prob, and perplexity, together with local likelihood statistics such as min_token_logprob, argmin_position, and per_token_logprob_std. supervised_position_count records the number of token positions that contributed to the likelihood calculation. The start and end fields represent corpus coordinates.
Intended Uses
This dataset is intended for model-based analysis of genomic sequence likelihoods, including:
- Sequence likelihood distribution analysis
- Sequence ranking and filtering
- Model-based quality assessment
- Detection of sequences with unusual likelihood profiles
- Identification of locally low-probability regions
- Corpus sampling and curation
- Comparison of model likelihoods across Carbon corpus subsets
- Supporting downstream genomic dataset construction and enrichment
The likelihood statistics are intended as model-derived analytical features and should not be interpreted as independent biological or clinical annotations.
Information of Features
Intended Uses
This dataset is intended for model-based analysis of genomic sequence likelihoods, including:
- Sequence likelihood distribution analysis
- Sequence ranking and filtering
- Model-based quality assessment
- Detection of sequences with unusual likelihood profiles
- Identification of locally low-probability regions
- Corpus sampling and curation
- Comparison of model likelihoods across Carbon corpus subsets
- Supporting downstream genomic dataset construction and enrichment
