Can I use function calling with streaming?
How Streaming Works with Tool Calls
When you enable streaming, the model sends chunks of the response as they are generated. For tool calls, the initial chunk includes the tool call ID and function name, and subsequent chunks contain pieces of the arguments string. You must concatenate these pieces to get the full arguments JSON.
For example, OpenAI's streaming API returns delta objects with tool_calls array. Each delta may have a function.arguments fragment. You accumulate them in order.
- Listen for tool_calls in the delta
- Accumulate arguments fragments by index
- Once the stream ends, parse the complete arguments string
- Execute the tool and continue the conversation
Platform Differences
OpenAI, Anthropic, and Google all support streaming with tool calls, but the exact event structure varies. OpenAI uses server-sent events with delta objects; Anthropic uses content blocks with input_json_delta events; Google uses functionCall parts in streamed responses.
Check your platform's documentation for the specific format. The general pattern is the same: accumulate fragments until the stream completes.
Practical Considerations
Streaming tool calls can improve perceived latency because you can start processing as soon as the tool call is complete, without waiting for the entire response. However, you cannot execute the tool until you have the full arguments, so there may be a delay.
Also, handle errors gracefully. If the stream breaks, you may have incomplete arguments. Implement timeouts and retries as needed.
Common mistakes
- Trying to parse arguments from a single chunk instead of accumulating all fragments.
- Assuming all platforms use the same streaming format for tool calls.
- Not handling the case where the stream ends before the tool call is fully received.