Paper Slide Deck Generator
Transform academic papers and content into professional slide deck images with automatic figure extraction.
Usage
/paper-slide-deck path/to/paper.pdf
/paper-slide-deck path/to/paper.pdf --style academic-paper
/paper-slide-deck path/to/content.md --style sketch-notes
/paper-slide-deck path/to/content.md --audience executives
/paper-slide-deck path/to/content.md --lang zh
/paper-slide-deck path/to/content.md --slides 10
/paper-slide-deck path/to/content.md --outline-only
/paper-slide-deck # Then paste content
Setup (one-time)
The TypeScript scripts (merge-to-*, detect-figures, extract-figure,
apply-template) need Node dependencies. Install them once:
cd ${SKILL_DIR}/scripts && npm install
This installs canvas, pdfjs-dist, pptxgenjs, and pdf-lib (a package-lock.json
pins versions). If a script exits with missing Node dependency "<name>", run the
command above. The Python generator (generate-slides.py) auto-installs google-genai
on first run.
Also install PyMuPDF (pip install pymupdf) — it is the reliable fallback for
extracting figures from pages that embed bitmaps (X-rays, CAM heatmaps, photographs),
where the pdfjs + canvas path in extract-figure.ts fails with
Error: Image or Canvas expected. For medical-imaging papers this is the common case,
not the exception, so treat PyMuPDF as required, not optional.
Image generation & no-API-key path
Image generation needs either a GOOGLE_API_KEY/GEMINI_API_KEY (Gemini API) or the
Gemini Web skill. If no key and no web option is available, the skill still works in
a degraded mode — do not abort:
- Run with
--outline-onlyto produce the outline + prompts (no images). - For a source PDF, extract real figures/tables with
detect-figures.ts+extract-figure.ts+apply-template.ts(no API key needed — pure rendering). - Merge whatever slides exist (
extract-sourced pages) into PPTX/PDF, and hand theprompts/back to the user to generate images later when a key is available.
Script Directory
Important: All scripts are located in the scripts/ subdirectory of this skill.
Agent Execution Instructions:
- Determine this SKILL.md file's directory path as
SKILL_DIR - Script path =
${SKILL_DIR}/scripts/<script-name>.ts - Replace all
${SKILL_DIR}in this document with the actual path
Script Reference:
| Script | Purpose |
|--------|---------|
| scripts/generate-slides.py | Generate AI slides via Gemini API (Python) |
| scripts/merge-to-pptx.ts | Merge slides into PowerPoint |
| scripts/merge-to-pdf.ts | Merge slides into PDF |
| scripts/detect-figures.ts | Auto-detect figures/tables in PDF (heuristic; verify pages) |
| scripts/extract-figure.ts | Render a full PDF page to PNG (optional --crop; PyMuPDF fallback) |
| scripts/apply-template.ts | Apply figure container template |
Options
| Option | Description |
|--------|-------------|
| --style <name> | Visual style — gallery + auto-selection rules in references/style-selection.md |
| --audience <type> | Target audience: beginners, intermediate, experts, executives, general |
| --lang <code> | Output language (en, zh, ja, etc.) |
| --slides <number> | Target slide count |
| --outline-only | Generate outline only, skip image generation |
Style & layout selection: read references/style-selection.md
when picking a style (explicit or auto-selected from content signals — including the
academic-signal caution about garbled math/numbers). Read
references/layout-gallery.md only when composing the
outline's // LAYOUT hints (Layout: <name> per slide).
Design Philosophy
This deck is designed for reading and sharing, not live presentation:
- Each slide must be self-explanatory without verbal commentary
- Structure content for logical flow when scrolling
- Include all necessary context within each slide
- Optimize for social media sharing and offline reading
File Management
Output Directory
Each session creates an independent directory named by content slug:
slide-deck/{topic-slug}/
├── source-{slug}.{ext} # Source files (text, images; source-paper.pdf for papers)
├── figures.json # Figure-detection results (academic PDFs)
├── outline.md # Final outline (with IMAGE_SOURCE blocks)
├── outline-{style}.md # Style variant outlines
├── extracted/ # Raw PDF page extractions (pre-template)
├── figures/ # Extracted figure crops (pre-template)
├── prompts/
│ └── {NN}-slide-{slug}.md, ...
├── {NN}-slide-{slug}.png # Final slide images in the deck ROOT — both AI-generated
│ # and extracted-figure slides; ext may be .jpg
├── {topic-slug}.pptx
└── {topic-slug}.pdf
Slug Generation:
- Extract main topic from content (2-4 words, kebab-case)
- Example: "Introduction to Machine Learning" →
intro-machine-learning
Conflict Resolution
If slide-deck/{topic-slug}/ already exists:
- Append timestamp:
{topic-slug}-YYYYMMDD-HHMMSS - Example:
intro-mlexists →intro-ml-20260118-143052
Source Files
Copy all sources with naming source-{slug}.{ext}:
source-article.md(main text content)source-diagram.png(image from conversation)source-data.xlsx(additional file)
Multiple sources supported: text, images, files from conversation.
Workflow
Step 1: Analyze Content
-
Save source content (if pasted, save as
source.md) -
Follow
references/analysis-framework.mdfor deep content analysis -
Determine style (use
--styleor auto-select from signals) -
Detect languages (source vs. user preference)
-
Plan slide count (
--slidesor dynamic) -
For academic papers (PDF with figures): Run automatic figure detection:
npx -y bun ${SKILL_DIR}/scripts/detect-figures.ts --pdf source-paper.pdf --output figures.jsonThis outputs a JSON file with all detected figures/tables, their page numbers, and captions.
Caption detection is heuristic — verify, especially the first-page teaser. The line-anchored
Figure Nmatcher reliably finds captions that sit on their own line (single-column layouts), but misses figures whose caption is interleaved with body text on a two-column first page — which is often the paper's most important architecture/overview figure. After running detect-figures, cross-check the source'sFigure 1explicitly: if the paper's text references aFigure Nthat is absent fromfigures.json, add it manually via an// IMAGE_SOURCEblock and extract it with the PyMuPDF fallback. Do not assumefigures.jsonis complete.
Step 2: Generate Outline Variants
- Generate 3 style variant outlines based on content analysis
- Follow
references/outline-template.mdfor structure - Auto-populate IMAGE_SOURCE for academic papers:
- Read
figures.jsonfrom Step 1 - Map figures to slides using rules in
references/analysis-framework.mdSection 8 - Automatically add
// IMAGE_SOURCEblocks to appropriate slides:- Architecture/pipeline figures → Methods slides (
Source: extract) - Results tables → Quantitative results slides (
Source: extract) - Comparison images → Qualitative results slides (
Source: extract) - Conceptual/simple diagrams → Leave for AI generation (
Source: generateor omit)
- Architecture/pipeline figures → Methods slides (
- Read
- Save as
outline-{style}.mdfor each variant
Step 3: User Confirmation
Single AskUserQuestion with all applicable options:
| Question | When to Ask | |----------|-------------| | Style variant | Always (3 options + custom) | | Language | Only if source ≠ user language |
After selection:
- Copy selected
outline-{style}.mdtooutline.md - Regenerate in different language if requested
- User may edit
outline.mdfor fine-tuning
If --outline-only, stop here.
Step 4: Generate Prompts
- Read
references/base-prompt.md - Combine with style instructions from outline
- Add slide-specific content
- If
Layout:specified in outline, include layout guidance in prompt:- Reference layout characteristics for image composition
- Example:
Layout: hub-spoke→ "Central concept in middle with related items radiating outward"
- Save to
prompts/directory
Step 5: Image Generation Method Selection
Before generating images, ask user to choose generation method:
Use AskUserQuestion with options:
| Option | Label | Description | |--------|-------|-------------| | 1 | Gemini API (Recommended) | Official Google API via Python. Requires GOOGLE_API_KEY env var. | | 2 | Gemini Web (Browser-based) | ⚠️ Uses reverse-engineered web API. No API key needed but may break. |
If no API key is available, do not still recommend Option 1 — offer the degraded
no-key path from the Setup section (--outline-only + extraction + prompts handoff).
Option 1: Gemini API (Python)
- Verify API key: Check
GOOGLE_API_KEYorGEMINI_API_KEYenvironment variable - Run generation script:
The model id is stated once here:python3 ${SKILL_DIR}/scripts/generate-slides.py <slide-deck-dir>gemini-3-pro-image(Nano Banana Pro, GA; the older-previewid is deprecated). Override with--model <id>only if needed.
Script behavior (details in the script docstring):
- Auto-installs
google-genai; errors out (non-zero) if no prompt files are found - Retries failed generations with exponential backoff (3 attempts total)
- Skips already-generated slides (> 10KB, any image extension)
- Writes each slide to the deck root — the same place extracted-figure slides land, so one merge step picks up both
- Saves with the real image extension, converting webp responses to PNG (the merge scripts only accept png/jpg/jpeg)
Option 2: Gemini Web Skill
Read references/gemini-web.md for the consent check,
per-slide invocation, and proxy setup. It requires the baoyu-danger-gemini-web
skill and explicit user consent to the reverse-engineered-API disclaimer.
Step 5.5: Process IMAGE_SOURCE (Automatic Figure Extraction)
For academic presentations, IMAGE_SOURCE metadata was auto-populated in Step 2 based on figure detection from Step 1.
Automatic Execution:
-
Parse outline to identify slides with
Source: extract -
Create figures directory:
mkdir -p figures -
For each extract slide, automatically:
- Read the Figure number, Page, and Caption from metadata
- Run figure extraction script:
Note:npx -y bun ${SKILL_DIR}/scripts/extract-figure.ts \ --pdf source-paper.pdf \ --page <page-number> \ --output figures/figure-<N>.pngextract-figure.tsrenders the entire page to a high-resolution PNG — it does not auto-detect or crop a single figure's bounding box. On a two-column page you will get both columns. To isolate one figure, either pass--crop "x,y,width,height"(pixels in the rendered/scaled page) or open the PNG, confirm it visually, and crop manually before applying the template. - Run template application script:
npx -y bun ${SKILL_DIR}/scripts/apply-template.ts \ --figure figures/figure-<N>.png \ --title "<slide-headline>" \ --caption "Figure <N>: <caption-text>" \ --output <NN>-slide-<slug>.png - Report: "Extracted: Figure N → slide NN"
-
For slides with
Source: generate(or no IMAGE_SOURCE):- Proceed to Step 6 for AI generation
Note: Source PDF must be saved as source-paper.pdf in output directory.
Troubleshooting:
- If figure detection missed a figure: manually add
// IMAGE_SOURCEblock to outline - If wrong figure mapped: edit the
Figure:andPage:values in outline - If extraction fails: check PDF page number (1-indexed)
PyMuPDF Fallback for Page Extraction:
If extract-figure.ts fails with "Image or Canvas expected" error (common with complex PDFs), use PyMuPDF:
import fitz
doc = fitz.open("source-paper.pdf")
page = doc[page_num - 1] # 0-indexed
mat = fitz.Matrix(3, 3) # 3x scale for high resolution
pix = page.get_pixmap(matrix=mat)
pix.save(f"extracted/page-{page_num}.png")
Then apply template using apply-template.ts.
Step 6: Generate Images
- Use selected method from Step 5
- Skip slides already processed in Step 5.5 (those with
Source: extract) - Generate session ID:
slides-{topic-slug}-{timestamp} - Generate each remaining slide with same session ID
- Report progress: "Generated X/N"
- Failures retry automatically (API path: 3 attempts with backoff; see Step 5)
Step 6.5: Proofread Generated Images (Content Integrity)
Text-to-image bakes text into pixels and will garble spelling, math symbols, and numbers — this is the single biggest risk of this skill. Do not ship unchecked. (This is this skill's instance of the repository's shared citation-integrity core: numbers and claims must match the source, and failures surface as visible flags — never silent substitutes.)
For every generated slide (especially any with equations, tables, key numbers,
or non-Latin text), use Read to open the PNG and visually check:
- Spelling / wording — headline and body text match the outline, no invented or mangled words.
- Math & symbols — equations, subscripts, Greek letters, operators are correct (or absent). Assume the model got them wrong until you confirm otherwise.
- Numbers & units — any figure that carries data matches the source exactly.
If garbling is found:
- Regenerate that slide with a corrected/simplified prompt (spell risky terms phonetically, reduce text density, move exact numbers to a caption). Max 2 retries.
- If it still fails after 2 retries, flag the slide
[CHECK]in the Step 8 summary and recommend one of:- Replace with an extracted figure/table from the source PDF (
Source: extract), or - Simplify the slide to remove the fragile text, or
- For a deck that genuinely needs faithful, editable formulas/data, switch to scholar-slides.
- Replace with an extracted figure/table from the source PDF (
Never silently deliver a slide with garbled math or data — always surface it.
Step 7: Merge to PPTX and PDF
npx -y bun ${SKILL_DIR}/scripts/merge-to-pptx.ts <slide-deck-dir>
npx -y bun ${SKILL_DIR}/scripts/merge-to-pdf.ts <slide-deck-dir>
Step 8: Output Summary
Slide Deck Complete!
Topic: [topic]
Style: [style name]
Location: [directory path]
Slides: N total
- 01-slide-cover.png ✓ Cover
- 02-slide-intro.png ✓ Content
- 04-slide-results.png ⚠ [CHECK] math/numbers — verify or use scholar-slides
- ...
- {NN}-slide-back-cover.png ✓ Back Cover
Outline: outline.md
PPTX: {topic-slug}.pptx
PDF: {topic-slug}.pdf
List any [CHECK]-flagged slides (from Step 6.5) explicitly so the user knows which
slides may contain garbled text/math/data and how to remediate them.
Slide Modification
See references/modification-guide.md for:
- Edit single slide workflow
- Add new slide (with renumbering)
- Delete slide (with renumbering)
- File naming conventions
Image Generation Dependencies
Gemini API (Option 1 - Recommended)
Requires:
GOOGLE_API_KEYorGEMINI_API_KEYenvironment variable- Python 3.8+ with pip
google-genaipackage (auto-installed by script)
Model id: stated once in Step 5.
Gemini Web Skill (Option 2)
See references/gemini-web.md — requires the
baoyu-danger-gemini-web skill, Chrome with a logged-in Google account, and user consent.
PDF Figure Extraction
Requires (install via cd ${SKILL_DIR}/scripts && npm install):
- Primary:
pdfjs-distnpm package (use legacy build for Node.js) canvasnpm package for extract-figure.ts / apply-template.ts- Fallback:
pymupdfPython package (more reliable for complex PDFs)
References
Read only the reference file for the stage you are in — never glob references/.
| File | Read when |
|------|-----------|
| references/analysis-framework.md | Step 1 — deep content analysis; IMAGE_SOURCE mapping rules |
| references/style-selection.md | Step 1/3 — style gallery + auto-selection (incl. academic-signal caution) |
| references/outline-template.md | Step 2 — outline structure and STYLE_INSTRUCTIONS format |
| references/layout-gallery.md | Step 2 — per-slide Layout: hints |
| references/base-prompt.md | Step 4 — base prompt for image generation |
| references/gemini-web.md | Step 5 Option 2 — consent, invocation, proxy |
| references/modification-guide.md | Post-delivery edits — add/delete/renumber slide workflows |
| references/content-rules.md | Content and style guidelines |
| references/figure-container-template.md | Step 5.5 — visual specs for extracted figure containers |
| references/styles/<style>.md | Full specification of the selected style |
Notes
Image Generation
- Gemini API: Recommended. Stable, reliable, requires API key
- Gemini Web: No API key needed, but uses reverse-engineered API with account risk
- Generation time: 10-30 seconds per slide
- Failures retry automatically (API path: 3 attempts with backoff)
- Maintain style consistency via session ID
Content Guidelines
- Use stylized alternatives for sensitive public figures
- Both methods use the same underlying Gemini model for image generation