Copilot Auto Model Selection Now Has Tiers: Efficiency, Balance, Intelligence
Until September 2026, Auto in GitHub Copilot was one setting: let Copilot pick a model from what is available and how complex the task looks. The September 14 weekly release split that into three tiers — efficiency, balance, and intelligence — rolling out in VS Code, Copilot CLI, and the Copilot app. All three choose from the same pool of enabled models. What changes is how the picker weighs cost, quality, and response time.
That is a different knob from the model list. Picking a model by name, covered in the Copilot model comparison, is a per-task decision. A tier is a standing policy: "spend less by default" or "spend more by default," and let Auto do the per-task choice inside that policy.
What each tier is for
GitHub's changelog describes the tiers by what they weigh, not by which models they contain. Read them that way.
| Tier | Leans toward | Reasonable default for |
|---|---|---|
| Efficiency | Cost and response time | Repetitive edits, test scaffolding, docs, chat that is mostly lookups |
| Balance | The middle | Day-to-day feature work |
| Intelligence | Quality | Debugging across files, architecture questions, long agent runs |
The paid-plan discount that applies to Auto, documented on GitHub's model comparison page, is the reason to set a tier instead of pinning a model. A pinned frontier model costs list price on every turn, including the turn where you asked it to rename a variable. Auto on efficiency spends frontier credits only where the picker thinks the task warrants it.
The test is cheap. Run the same afternoon of work on efficiency and on intelligence, and look at the usage view, which now also shows credit usage and remaining balance to end users. If the output was fine on efficiency, you found your default. Flip to intelligence for the one bug that needs it.
The models that arrived with the tiers
Two OpenAI models reached general availability in Copilot in September, and both are billed at provider list pricing under usage-based billing.
GPT-6 Astra, GA on September 4, is described for long-horizon, autonomous coding: it plans and validates as it goes, batches diagnosis with verification, and confirms its own result before declaring a task done. GitHub reports fewer steps than prior OpenAI models on long tasks.
GPT-6.1 Sol, GA on September 29, is positioned for agentic coding and terminal workflows, with the specific claim that it completed tasks with noticeably fewer tokens and steps than earlier GPT-6 and GPT-5.6 models. Fewer tokens per task is the number that matters under usage-based billing; a model that is marginally better and twice as verbose can be the more expensive choice.
Both are available to Copilot Pro+, Max, Business, and Enterprise, across VS Code, Visual Studio, Copilot CLI, the coding agent, the Copilot app, github.com, mobile, JetBrains, Xcode, and Eclipse. Rollout is gradual. If the picker does not show them yet, that is the rollout, not your plan.
There is also an experimental option in Copilot CLI called Project HydraFusion, under /experimental. You select it like a model and it routes between local, cloud, and compound models per task. It is a research preview, not a tier. Treat results as something to observe, not a default for a team.
What admins control
For Copilot Business and Enterprise, new models are enabled automatically under default model enablement unless an administrator turned off the global default or disabled a specific model. Model policy in Copilot settings is where Astra and Sol are allowed or blocked. If a team cannot see a model, check policy before checking rollout.
Two September changes matter for budgets. Users who hit their AI credit limit can now request a higher budget from their organization or enterprise, and owners or billing managers approve, adjust, or deny in settings — generally available on Business and Enterprise with usage-based billing, not available for enterprises with managed users. And on the Microsoft side, admins can set which model families are available to groups of users, and those settings shape which models Auto can select from. A tier chooses inside the allowed set. It does not override it.
A policy that survives the next model
- Set balance as the organization default. Let individuals move to efficiency or intelligence.
- Leave default model enablement on so new GA models (Astra, Sol, the next one) reach the picker without a ticket, and block specific models by policy if cost or data handling requires it.
- Watch tokens per completed task, not tokens per request. Sol's claim is about the former.
- Keep HydraFusion to people who will report back. It is in
/experimentalfor a reason. - Revisit the tier quarterly. The pool behind Auto changes faster than the policy should.
The model picker is now two decisions: which models are allowed, and how hard Auto should try. Most of the cost control lives in the second one.
Keep reading
Indirect Prompt Injection: Tool Output Is Not Instructions
A retrieved document, an email, or a tool result can tell the model to take an action. Delimiters do not stop it. The tool allowlist after untrusted text does.
GitHub Copilot AI Model Comparison: Which Model to Use for Each Task
AI model comparison for GitHub Copilot Chat: GPT-5.6, Claude, Gemini, Grok, and Kimi, plus when Auto is the cheaper default.
Put the Workflow in Code When the Agent Keeps Improvising
Which steps belong in a fixed program and which belong in a model, so an agent stops skipping the check that production depends on.
Evaluating LLM Applications: Getting Past 'It Looks Good to Me'
Shipping LLM features on vibes works until a prompt tweak silently breaks ten other cases. Here is how to build evals that catch regressions before your users do.
Getting Structured Output From an LLM Without the Heartbreak
Asking a model to 'return JSON' and parsing the result is how you get 3 a.m. pages. Tool schemas, constrained decoding, and validation turn a probabilistic text generator into a reliable API.
Chunking Strategies for RAG: Where Retrieval Quality Is Won or Lost
Most RAG systems that retrieve bad context aren't failing at embeddings or reranking — they're failing at chunking. How you split documents quietly decides what your model can ever find.
Newsletter
New posts, straight to your inbox
One email per post. No spam, no tracking pixels, unsubscribe anytime.
Comments
- No comments yet. Be the first.