datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SecureCodePairs
Dataset Summary
Field
Value
Version
1.2.0
License
MIT
Total code examples
470
LLM security trajectories
30
Languages (15)
Python, Java, JavaScript, TypeScript, Go, PHP, C#, Kotlin, Swift, Rust, Ruby, C, C++, Scala, YAML (Kubernetes)
Frameworks
Flask, Django, FastAPI, Spring Boot, Express, NestJS, Next.js, Laravel, ASP.NET Core, Gin, Android, iOS, Actix, Rails, Qt, Play, gRPC, GraphQL, Kubernetes
New in v1.2.0
+260 records (deep Python/Java packs… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/SecureCodePairs.movies
Movie Scripts Dataset
The Movie Scripts Dataset consists of scripts from 1,172 movies, providing a comprehensive collection of movie dialogues and narratives. This dataset is designed to support various natural language processing (NLP) tasks, including dialogue generation, script summarization, and text analysis.
Details
The dataset contains 2 columns:
Name: The title of the movie.
Script: The full script of the movie in English.
Usage
The Movie Scripts… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/movies.books
Books
The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
Details
The dataset contains 4 columns:
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/books.libri-in-italiano
Libri
Il dataset dei libri consiste in una raccolta diversificata di 18 libri organizzati in 4 categorie.
Questo dataset è ben pulito e progettato per supportare diversi compiti di elaborazione del linguaggio naturale (NLP), inclusi generazione di testo, traduzione e modellazione del linguaggio mascherato.
Dettagli
Il dataset contiene 4 colonne:
titolo: Il titolo del libro.
autore: L'autore del libro.
categoria: Il genere/categoria del libro.
contenuto: Il contenuto… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/libri-in-italiano.
