DSPy Adapters and Multimodal I/O

SkillMedia

Use for DSPy adapter selection, JSONAdapter, XMLAdapter, ChatAdapter, native function calling, structured outputs, and multimodal inputs like dspy.Image or dspy.Audio.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the DSPy Adapters and Multimodal I/O skill

What this skill tells your AI

The instructions your AI receives, as published by omidzamani/dspy-skills in skills/dspy-adapters-multimodal/SKILL.md and read by ahel’s review.

Goal

Choose an adapter deliberately and model image, audio, and file inputs with DSPy's typed primitives.

Adapter Selection

AdapterUse it for
dspy.ChatAdapter()Default, human-readable field markers, broad model compatibility
dspy.JSONAdapter()Structured JSON output and native function calling where supported
dspy.XMLAdapter()XML-tagged fields when XML is easier for the target LM to follow
dspy.TwoStepAdapter()A separate extraction pass when parsing needs extra help

Configure globally or for a limited scope:

import dspy

dspy.configure(
    lm=dspy.LM("openai/gpt-4o-mini"),
    adapter=dspy.JSONAdapter(),
)

with dspy.context(adapter=dspy.XMLAdapter()):
    result = dspy.Predict("question -> answer")(question="What is DSPy?")

Native Function Calling

JSONAdapter enables native function calling by default. ChatAdapter keeps text parsing by default. Override either behavior explicitly:

chat_native = dspy.ChatAdapter(use_native_function_calling=True)
json_manual = dspy.JSONAdapter(use_native_function_calling=False)

DSPy falls back to manual parsing when the configured LM does not support native function calling.

Image Inputs

class DescribeImage(dspy.Signature):
    image: dspy.Image = dspy.InputField()
    description: str = dspy.OutputField()

describe = dspy.Predict(DescribeImage)
result = describe(image=dspy.Image("./diagram.png"))

Pass a local path, HTTP URL, bytes, PIL image, or existing data URI directly to dspy.Image(...).

Audio and File Inputs

class SummarizeAudio(dspy.Signature):
    audio: dspy.Audio = dspy.InputField()
    summary: str = dspy.OutputField()

audio = dspy.Audio.from_file("./meeting.wav")
summary = dspy.Predict(SummarizeAudio)(audio=audio)
class SummarizeFile(dspy.Signature):
    file: dspy.File = dspy.InputField()
    summary: str = dspy.OutputField()

document = dspy.File.from_path("./research.pdf")
summary = dspy.Predict(SummarizeFile)(file=document)

Provider capabilities vary. Verify that the selected model accepts the media type before deployment.

Best Practices

  1. Start with ChatAdapter; switch only for a measured reason.
  2. Use typed signatures for structured output.
  3. Test adapter behavior against the exact production model.
  4. Avoid deprecated Image.from_file() and Image.from_url() helpers; call dspy.Image(...).
  5. Keep local file handling and uploaded file IDs within provider policy.

Related Skills

Official Documentation

Signals

GitHub stars
123
Forks
13
Last commit
Jun 2026
Advanced
Catalog kind
skill
Gateway key
dspy-adapters-multimodal
Source
github.com/omidzamani/dspy-skills