iam
Datasets
All datasets matching “iam”python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
iam_handwriting_finevision
Dataset Card for finevision_iam
This is a FiftyOne dataset with 5663 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/iam_handwriting_finevision")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/iam_handwriting_finevision.chanpicfortestIAM-line
IAM - line level
Dataset Summary
The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.chat_formatted_examplesiamd_v0
Internet Archive Music Dataset (IAMD v0)
~4.2M thirty-second music segments (34,469 hours) sourced from
Creative-Commons audio on the Internet Archive, each
paired with machine-generated natural-language captions and the original item
metadata.
Segments
4.2M
Audio
34k hours
Segment length
30 s nominal (mean 29.22 s)
Format
MP3, 320 kbps CBR, native channels + sample rate
Shards
2,320 Parquet files
Download size
4.53 TB
Loading
A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.
