Agent Skills: DSPy Production Deployment

Use for deploying DSPy with save/load, configure_cache, restrict_pickle, track_usage, async execution, streaming, and production runtime controls.

UncategorizedID: OmidZamani/dspy-skills/dspy-production-deployment

Install this agent skill to your local

pnpm dlx add-skill https://github.com/OmidZamani/dspy-skills/tree/HEAD/skills/dspy-production-deployment

Skill Files

Browse the full folder contents for dspy-production-deployment.

Download Skill

Loading file tree…

skills/dspy-production-deployment/SKILL.md

Skill Metadata

Name
dspy-production-deployment
Description
Use for deploying DSPy with save/load, configure_cache, restrict_pickle, track_usage, async execution, streaming, and production runtime controls.

DSPy Production Deployment

Goal

Prepare a DSPy program for repeatable, observable, scalable, and safer production execution.

Cache Hardening

DSPy enables memory and disk caches by default. Disk cache deserialization uses pickle unless restricted. Enable the allowlist mode in production:

import dspy

dspy.configure_cache(restrict_pickle=True)

Register trusted custom cache types only when needed:

dspy.configure_cache(
    restrict_pickle=True,
    safe_types=[MyResult, Metadata],
)

Disable a cache layer explicitly when a deployment cannot persist data or requires fresh model responses:

dspy.configure_cache(
    enable_disk_cache=False,
    enable_memory_cache=True,
)

Save and Load

Prefer state-only JSON for readable, safer artifacts:

compiled.save("./artifacts/program.json", save_program=False)

loaded = MyProgram()
loaded.load("./artifacts/program.json")

Use whole-program save only for trusted artifacts. It uses cloudpickle:

compiled.save("./artifacts/program/", save_program=True)
loaded = dspy.load("./artifacts/program/")

Keep the DSPy major version compatible when loading saved programs.

Usage Tracking

dspy.configure(
    lm=dspy.LM("openai/gpt-4o-mini"),
    track_usage=True,
)

prediction = program(question="What is DSPy?")
print(prediction.get_lm_usage())

Cached calls return no new token usage.

Async Execution

Most built-in modules support acall():

import asyncio

async def main():
    prediction = await program.acall(question="What is DSPy?")
    print(prediction.answer)

asyncio.run(main())

Implement aforward() for custom async modules. Use dspy.asyncify(program) only when adapting a synchronous callable is the right boundary.

Streaming

import asyncio
import dspy

stream_program = dspy.streamify(
    dspy.Predict("question -> answer"),
    stream_listeners=[
        dspy.streaming.StreamListener(signature_field_name="answer"),
    ],
)

async def main():
    async for chunk in stream_program(question="Explain DSPy briefly."):
        print(chunk)

asyncio.run(main())

For looped modules such as ReAct, set allow_reuse=True on listeners for repeated fields. Cache hits yield the final Prediction without replaying token chunks.

Production Checklist

  1. Pin the stable DSPy series.
  2. Use state-only JSON unless whole-program pickle is necessary and trusted.
  3. Enable restrict_pickle=True.
  4. Record usage, latency, errors, and traces.
  5. Load-test async and streaming paths separately.
  6. Use dspy-debugging-observability for MLflow and callbacks.

Official Documentation

  • Production guide: https://dspy.ai/production/
  • Cache tutorial: https://dspy.ai/tutorials/cache/
  • Saving tutorial: https://dspy.ai/tutorials/saving/
  • Async tutorial: https://dspy.ai/tutorials/async/
  • Streaming tutorial: https://dspy.ai/tutorials/streaming/
DSPy Production Deployment Skill | Agent Skills