Prompt Engineering Standards
Load if: Writing AI prompts, optimizing context usage Prerequisites: @smith-principles/SKILL.md
Prompt Caching
Cache reduces costs 90%, latency 85%
Structure for caching:
- Static content first (methodology, rules)
- Tool definitions in consistent order
- Project context (AGENTS.md, docs)
- Dynamic content last (recent changes)
Cache breakpoints: Every ~1024 tokens. Prefix must be identical for cache hit.
Rules:
- Keep tool definitions in a consistent order between calls
- Keep dynamic content out of static sections
- Keep the cached prefix stable — only change it when the static content itself changes
- Use bullet lists instead of Markdown tables (see
@smith-skills/SKILL.md)
AGENTS.md Cache-Friendly Structure
<!-- STATIC - cached -->
**Metadata**: Scope, Load if, Prerequisites
## Critical Rules
Critical ALWAYS rules, written as affirmative statements
## Hard Limits
Anti-patterns with no natural positive phrasing
<!-- CACHE BREAKPOINT (~1024 tokens) -->
<!-- DYNAMIC - not cached -->
## Examples
Code examples that evolve
Token Efficiency
Progressive Disclosure
Three-level loading:
- Metadata only (50 tokens)
- Core concepts when triggered (200 tokens)
- Full details when accessed (1000+ tokens)
Sparse Attention
Efficient file reading:
- Grep to find location
- Read with offset/limit for large files
- Read only necessary context (±20 lines)
Rules:
- Prefer targeted reads over loading full files
- Check metadata before reading full documentation
- Answer directly instead of restating the user's question
Structured Output
Platform mechanisms:
- OpenAI: JSON Schema with
strict: true(100% compliance) - Anthropic: Tool use with flexible schemas
- Gemini: responseSchema with retry
Schema design:
- Match existing project patterns
- Include descriptions for complex fields
- Define required vs optional fields
- Keep nesting ≤3 levels
Related
- @smith-ctx/SKILL.md - Progressive disclosure, reference-based communication
@smith-xml/SKILL.md- XML tags for runtime prompts (not SKILL.md bodies)
Before You Finish
For caching:
- Place static content before dynamic
- Maintain consistent tool order
- Target >80% cache hit rate
For efficiency:
- Use Grep before Read
- Read incrementally (narrow → expand)
- Use file:line references