Agent Block
Connect workflows to AI models for text processing, data extraction, classification, and structured output
The Agent block connects your workflow to Large Language Models (LLMs) through Vercel AI Gateway. Use it to extract data from documents, classify inputs, generate text, or make decisions based on unstructured data. The block supports structured output via JSON Schema, ensuring consistent, machine-readable responses.
Configuration
Provider and Model
The canvas exposes two provider/model pairs:
| Provider | Model | Model ID |
|---|---|---|
| Anthropic | Claude Opus 5 | claude-opus-5 |
| OpenAI | GPT-5.6 Sol | gpt-5.6-sol |
Claude Opus 5 is the default. Provider and model are separate fields, so keep them paired as shown above; an Anthropic provider with an OpenAI model ID (or the reverse) is rejected.
The block does not require a workflow connection. The runtime must have Vercel AI Gateway authentication configured through AI_GATEWAY_API_KEY, or VERCEL_OIDC_TOKEN when it is available in a Vercel-hosted environment.
System Prompt
The system prompt defines the agent's role, behavior, and constraints. It is sent with every request and shapes how the agent processes input.
You are an expert invoice data extractor. Given a document, extract:
- Invoice number
- Vendor name
- Total amount
- Due date
- Line items with descriptions and amounts
Always return valid JSON matching the output schema.User Prompt
The primary input for the agent. This is the text or data the agent will process. Sources include:
- Static text configured directly in the block
- Dynamic data from upstream blocks via template variables (e.g.,
{{steps.download.content}}) - Trigger payload data from the workflow input
Input Data
Optional JSON object passed as additional context to the agent. Use this to pass structured data from upstream blocks:
{
"documentContent": "{{steps.download.content}}",
"documentMimetype": "application/pdf",
"customerName": "{{steps.lookup.output.name}}"
}Attaching a document by URL
For anything larger than a page or two, pass documentUrl instead of
documentContent. The model provider fetches the file directly, so the bytes
never pass through the workflow run:
{
"documentUrl": "{{steps.get_invoice_file.url}}",
"documentMimetype": "application/pdf"
}Base64 content is inflated ~33% by the encoding, held in memory, and retained in the step output for the entire run — a 7 MB PDF is enough to exhaust a small machine. A URL is a few hundred bytes.
Pair it with workspace_files.getFile in mode: "url". Notes:
- Only PDFs and images can be attached by URL — those are the types the
providers fetch natively. Anything else must go through
documentContent. - Set
documentMimetypeunless the URL path ends in a recognizable extension. The step fails with a clear error rather than sending the model an attachment it cannot read. - The URL must be HTTPS, and is a bearer credential for its lifetime. Keep the expiry short (the default is 10 minutes).
- Mint the URL immediately before the agent step. A signed URL created
before a
waitorapprovalstep can expire while the run is paused, and the failure surfaces as an opaque provider-side fetch error. documentContentis not scanned for URLs — usedocumentUrlexplicitly.
Output Schema
Define a JSON Schema to enforce structured output. When set, the agent's response is constrained to match the schema exactly, preventing free-form text and ensuring downstream blocks receive consistent data.
{
"type": "object",
"properties": {
"invoiceNumber": { "type": "string" },
"vendorName": { "type": "string" },
"totalAmount": { "type": "number" },
"dueDate": { "type": "string" },
"lineItems": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": { "type": "string" },
"amount": { "type": "number" }
}
}
}
},
"required": ["invoiceNumber", "vendorName", "totalAmount"]
}Temperature
Temperature is a model-specific sampling hint. Lower values typically reduce variation, while higher values typically allow more varied responses:
| Range | Typical behavior | Best For |
|---|---|---|
| 0 -- 0.2 | Focused, lower variation | Data extraction, classification, structured output |
| 0.3 -- 0.6 | Balanced | General processing, summarization |
| 0.7 -- 1.0 | Creative, varied | Content generation, brainstorming |
Default: 0. The exact effect depends on the selected model, and temperature alone does not guarantee identical output. For billing workflows, start at 0 and use an output schema plus downstream validation for fields that must be reliable.
Max Tokens
Maximum length of the agent's response. The runtime uses 4096 when this field is omitted. Set a lower value to control costs or response length, or a higher value when the selected model's output limit allows it.
Outputs
| Output | Type | Description |
|---|---|---|
result | JSON | Structured output matching the output schema (when schema is set) |
text | string | Raw text response (when no schema is set) |
usage | JSON | Token usage statistics (prompt tokens, completion tokens, total) |
Example Use Cases
Invoice Data Extraction
Stagehand (download PDF) -> Agent (extract invoice fields) -> QuickBooks (create invoice)System prompt: "Extract invoice number, vendor, total, due date, and line items from the document." Output schema: Structured JSON matching QuickBooks invoice format.
Payment Classification
Stripe (payment received) -> Agent (classify payment type) -> Condition (recurring vs one-time)System prompt: "Classify this payment as recurring, one-time, or refund based on the metadata."
Output schema: { "type": "recurring" | "one-time" | "refund", "confidence": number }
Contract Term Extraction
DocuSign (signed contract) -> Agent (extract billing terms) -> Transform -> Stripe (create subscription)System prompt: "Extract billing frequency, amount, start date, and payment terms from this contract."
Email Triage
Gmail (new email) -> Agent (classify and extract) -> Condition (route by type)
-> Invoice inquiry: Slack (notify finance)
-> Payment issue: Stripe (lookup) -> Slack (notify support)Best Practices
-
Start with temperature 0 for extraction tasks. This usually reduces variation, but it is not a determinism guarantee. Use an output schema and validate critical financial fields before acting on them.
-
Always define an output schema for structured data. This prevents the agent from returning unexpected formats and ensures downstream blocks receive the data they expect.
-
Be specific in system prompts. Instead of "extract data from this document," write "extract the invoice number (alphanumeric, usually in format INV-XXXX), vendor name, total amount in USD, and due date in ISO 8601 format."
-
Use the right model for the task. Claude Opus 5 is the default and is the better fit when native PDF handling or complex document reasoning matters. GPT-5.6 Sol is available for text, image, tool-driven, and agentic workloads and has lower current AI Gateway token rates.
-
Handle PDF documents with care. When passing PDF content, include
documentMimetype: "application/pdf". Claude Opus 5 accepts PDF attachments natively. The current GPT-5.6 Sol adapter handles images directly but not native PDF parts, so the runtime switches a GPT-configured PDF step to Claude Opus 5 when Gateway authentication is available. Select Claude Opus 5 directly when the executed model must stay explicit, or extract the PDF text upstream.
Loopfour