Put the Workflow in Code When the Agent Keeps Improvising
An agent that "uses tools" will sometimes call them in a sensible order and sometimes skip the one that checks a balance, a permission, or a duplicate. If that check is optional in the prompt, it is optional in production. The model is doing what you asked: deciding. The workflow did not want a decision there.
Put the order in code when every successful run must do the same steps. Let the model fill the gaps that are actually language: classify the email, extract the fields, draft the reply, choose among documented options.
A split that holds up
| In code | In the model |
|---|---|
| Authenticate, authorize, and scope to the tenant | Read unstructured text and pull out ids the code will verify |
| "Search, then get, then act" as three calls the program makes | Decide whether the user's paragraph is a refund request or a question |
| Idempotency, retries, timeouts | Phrase the question back when a required field is missing |
| The commit, the payment, the send | Rank a short list the code already filtered |
The code calls the model at the points of judgment and ignores any tool plan the model invents for the steps that are not negotiable. You can still expose tools. The outer function does not ask the model whether to check the duplicate. It checks, then asks the model what to tell the user.
This feels less magical. It is how you get a log that says the check happened. An improvised trace that happened to include the check on Tuesday does not tell you it will happen on Wednesday.
Where improvisation is the feature
A support assistant that must handle questions you have not scripted should choose tools. Constrain the choice with the descriptions in MCP tool descriptions that route the model, and keep the write tools behind a confirmation the code enforces. "The model said it confirmed" is not a confirmation. A function that returns only after the user or a policy approves is a confirmation.
Evaluation follows the split. For the code path, assert the calls: duplicate check before create, no send when the check fails. For the model path, score the classification and the draft against examples. Mixing them into one "did the agent do well" grade hides a skipped check behind a fluent email.
If you cannot name the steps that must always happen, you do not have a workflow yet. You have a chat. Ship the chat if that is the product. The moment a missed step costs a double refund or a leaked row, lift that step out of the prompt and into a function the model cannot skip.
Keep reading
Indirect Prompt Injection: Tool Output Is Not Instructions
A retrieved document, an email, or a tool result can tell the model to take an action. Delimiters do not stop it. The tool allowlist after untrusted text does.
Copilot Auto Model Selection Now Has Tiers: Efficiency, Balance, Intelligence
GitHub Copilot's Auto picker gained efficiency, balance, and intelligence tiers in September 2026, alongside GPT-6 Astra and GPT-6.1 Sol. How the tiers change the cost-quality trade, and what admins control.
GitHub Copilot AI Model Comparison: Which Model to Use for Each Task
AI model comparison for GitHub Copilot Chat: GPT-5.6, Claude, Gemini, Grok, and Kimi, plus when Auto is the cheaper default.
Evaluating LLM Applications: Getting Past 'It Looks Good to Me'
Shipping LLM features on vibes works until a prompt tweak silently breaks ten other cases. Here is how to build evals that catch regressions before your users do.
Getting Structured Output From an LLM Without the Heartbreak
Asking a model to 'return JSON' and parsing the result is how you get 3 a.m. pages. Tool schemas, constrained decoding, and validation turn a probabilistic text generator into a reliable API.
Chunking Strategies for RAG: Where Retrieval Quality Is Won or Lost
Most RAG systems that retrieve bad context aren't failing at embeddings or reranking — they're failing at chunking. How you split documents quietly decides what your model can ever find.
Newsletter
New posts, straight to your inbox
One email per post. No spam, no tracking pixels, unsubscribe anytime.
Comments
- No comments yet. Be the first.