filtered
Gemma-2-2B-Stheno-Filtered-i1-GGUFsfm_filtered_e2e_alignment_upsampled_basesfm_filtered_midtrain_alignment_upsampled_basesfm-midtraining_mix_blocklist_filteredGemma-2-2B-Stheno-Filtered-GGUFGemma-2-2B-Stheno-Filtered-GGUFsfm-midtraining_blocklist_filtered_insert_xxf_charactersfm-midtraining_e2e_blocklist_filtered__insert_hyperstition_v1
Datasets
All datasets matching “filtered”State-Parse-FilteredThe single cell RNA-seq dataset with human PBMC samples was sourced from Parse Biosciences [1]. [1] Performance of Evercode™ WT v3 in Human Immune Cells (PBMCs), https://www.parsebiosciences.com/datasets/performance-of-evercode-wt-v3-in-human-immune-cells-pbmcs/; Parse Biosciences, Seattle, USA; accessed 05/27/2025.
Certain uses of this data may require a license from Parse Biosciences, Inc.
OmniEdit-Filtered-1.2M
OmniEdit
In this paper, we present OMNI-EDIT, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) OMNI-EDIT is trained by utilizing the supervision
from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality.
📃Paper | 🌐Website |… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/OmniEdit-Filtered-1.2M.CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
OpenCodeInterpreter
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.MultiLegalPile_Wikipedia_FilteredA filtered version of the MultiLegalPile dataset, together with wikipedia articles.hoigen-filtered-videos
HOIGen Filtered Videos Dataset
This dataset contains 28562 filtered videos from the HOIGen-1M dataset based on the allowlist.
Dataset Structure
The videos are organized in the same structure as the original HOIGen dataset:
filtered_videos/
├── videos_part_1/
├── videos_part_2/
├── ...
└── videos_part_100/
Usage
from huggingface_hub import hf_hub_download
# Download a specific video
video_path = hf_hub_download(… See the full description on the dataset page: https://huggingface.co/datasets/charlychan123/hoigen-filtered-videos.objaverse-filtered
3D Model Dataset (built from allenai/objaverse)
Filtered, Blender-validated subset of allenai/objaverse.
Model files are preserved byte-for-byte from the source; each record carries
Blender-extracted geometry, materials, textures, scene structure, and hierarchy.
Layout
data/shard-XXXXXX/models/ original model files ({uid}.glb)
data/shard-XXXXXX/metadata.jsonl one record per model
data/shard-XXXXXX/train.jsonl chat-format LLM fine-tuning pairs… See the full description on the dataset page: https://huggingface.co/datasets/raj2708/objaverse-filtered.
