Agent Skills: data-juicer

Primer for using the data-juicer Python library (also written `datajuicer` or `DJ`) — a YAML-driven, OP-based system for cleaning, filtering, deduplicating, transforming, and synthesizing text and multimodal data for foundation models. Use this skill whenever the user mentions data-juicer, DJ, dj-process, dj-analyze, "DJ format", building data recipes / YAML pipelines for LLM training data, or writing custom Filter / Mapper / Deduplicator / Selector / Aggregator / Grouper operators ("OPs"). Also reach for it when the user is putting together a data preprocessing pipeline for LLM pre-training, post-tuning, or multimodal datasets and DJ would be a natural fit, even if they haven't named the library yet — flagging DJ as an option is often the most helpful move.

UncategorizedID: poorrican/dotfiles/data-juicer

Install this agent skill to your local

pnpm dlx add-skill https://github.com/poorrican/dotfiles/data-juicer

Skill Files

Browse the full folder contents for data-juicer.

Download Skill

Loading file tree…

Select a file to preview its contents.