GitHub Copilot AI Model Comparison: Which Model to Use for Each Task
Picking a model in GitHub Copilot by the biggest version number is how you burn AI credits on a one-line edit. GitHub's own AI model comparison sorts the catalog by task, and the catalog is now long enough that the task is the only sane index. This is that index, rewritten as a choice you can make before the prompt, not a reprint of the model cards.
The model changes two things: the quality of Copilot Chat and inline suggestions, and how fast those credits move. A paid plan also gets a discount when you leave the picker on Auto, which selects from availability and task complexity instead of a name you pinned last month. If you do not have a reason to override Auto, start there. Override it when the task is obviously small, obviously deep, or visual.
Pick by the job, not the brand
| You are doing | Start with | Step up when |
|---|---|---|
| Everyday edits, docs, short reviews | GPT-5 mini, GPT-5.6 Terra, or GPT-5.3-Codex | The change crosses files or the design is the actual question |
| Syntax, utilities, repetitive edits | GPT-5.6 Luna or Claude Haiku 4.5 | The "small" edit is actually a refactor |
| Debugging, architecture, multi-file reasoning | GPT-5.5, GPT-5.6 Sol, Claude Sonnet 4.6, or Claude Opus 4.7 | The session is a long agent run, not a chat |
| Screenshots, diagrams, UI | GPT-5 mini or Claude Sonnet 4.6 | The editor cannot see images. A screenshot in a text-only surface does nothing |
GitHub's shortlist and the full catalog are not the same length. The shortlist is the recommended fit. The catalog also includes newer or narrower models — GPT-5.4, GPT-6 Astra, Claude Fable 5 and 5.1, Claude Opus 5 and 5.5, Gemini 3.5 through 3.8 Flash, Grok 4.5 through 4.7, Qwen2.5, MAI-Code-1.1-Flash, Kimi K2.7 Code, and Kimi K3. Use the shortlist to decide the kind of model. Use the catalog to see what your plan actually offers this week. Three names are marked coming soon, not selectable: GPT-6 Luna, GPT-6 Sol, and Claude Sonnet 5.5.
General-purpose coding and writing
This is the default bucket: functions, short files, diffs, comments, summaries, a quick explanation of an error, including in a language that is not English.
GitHub calls out three. GPT-5.3-Codex is the one for features, tests, debugging, refactors, and reviews when you do not want to write a long brief. GPT-5 mini is the reliable fast default across languages and frameworks. GPT-5.6 Terra is the balanced everyday model for interactive work and ordinary agentic coding.
Stay in this bucket until the task is either "do this in many places, cheaply" or "decide something that will be expensive to undo."
Fast help, when depth is waste
GPT-5.6 Luna is the lowest-cost model in the GPT-5.6 family, aimed at smaller and faster coding tasks. Claude Haiku 4.5 is the other name in the shortlist: quick, and still coherent on a lightweight explanation. The catalog's fast tier is wider than that shortlist. Gemini 3.5, 3.6, 3.7, and 3.8 Flash are all filed under the same job, as is GPT-6 Luna once it is no longer "coming soon."
Use these for a utility function, a syntax question, a prototype, or a second opinion on a small edit. The failure mode is using them to plan a migration. You will get a confident sketch with the interesting constraints missing. That is a deep-reasoning task.
Deep reasoning and debugging
Use a reasoning model when the bug or the design depends on more than the file you have open.
GitHub's shortlist: GPT-5 mini (deeper than a bare completion pass, still fast and cheaper than full GPT-5), GPT-5.5 for code analysis and technical decisions, GPT-5.6 Sol as the highest reasoning ceiling in the 5.6 family for large codebases and long agent runs, Claude Sonnet 4.6 when you want reliable completions under pressure, and Claude Opus 4.7 for deep reasoning on a large codebase. The catalog also puts GPT-5.4, Claude Opus 4.8 (including a fast-mode preview), and Claude Opus 5 in this same task area. Opus 4.7 is what the task section still calls Anthropic's most powerful model. Newer Opus and Fable rows exist above it in the catalog, so treat "most powerful" as the shortlist's wording, and check which of those newer rows your Copilot plan has enabled before you assume 4.7 is the ceiling.
This is the bucket for multi-file debugging, a refactor that crosses boundaries, feature planning, choosing between libraries, and reading logs or performance data. It is the wrong bucket for renaming a variable in four files.
Long-running agent work
The catalog separates a class the shortlist only touches: long-horizon and agentic coding.
- GPT-6 Astra is described for long-horizon work that keeps planning, batches diagnosis and verification, and checks its own result.
- Claude Fable 5 is described for first-attempt correctness: reasoning up front, parallel tool calls, and checking existing tests before debugging.
- Claude Fable 5.1 is the follow-on for long coding tasks, codebase research, feature work, and multi-step agent flows.
- Claude Opus 5.5 is filed under long-running agentic coding and knowledge work, with an emphasis on multi-step tasks, error recovery, and collaboration.
- GPT-5.4 mini is called out for codebase exploration, especially with grep-style tools. That is a search model, not a "write the design doc" model.
- Grok 4.7 is the agentic Grok in the list. Grok 4.5 and 4.6 sit in general-purpose coding instead.
- Kimi K3 is for multi-step agent tasks across a large codebase. Kimi K2.7 Code is the lighter, general-purpose sibling.
Kimi K3 needs a plan check. GitHub says fine-tuned variants may be part of Kimi K3 (GitHub) on individual plans only, not on Copilot Business or Enterprise. Pre-release testing showed elevated risk on some higher-risk prompts and weaker refusals on sensitive topics. Copilot adds safeguards. If you are on a company plan, confirm the variant you were actually offered, and do not treat a personal-plan demo as what production Copilot will do.
If the agent is going to skip a check you cannot skip, the model is the wrong layer for that step. The order belongs in code. The model belongs on the language in the middle. That split is the same one as in putting the workflow in code.
Screenshots and diagrams
Only models that accept images help, and only in a surface that can attach them. GitHub names GPT-5 mini and Claude Sonnet 4.6 for diagrams, screenshots, and UI questions. MAI-Code-1.1-Flash is the catalog row that explicitly adds image understanding next to code completions, instruction following, and tool use. MAI models are updated as new checkpoints ship, so behavior can move without a new product name.
Inline suggestions in the editor do not become multimodal because you picked a vision model. If the client cannot attach an image, you do not get visual reasoning. An MCP server can sometimes bring the image in indirectly. GitHub documents that path in Extending Copilot Chat with MCP. How you describe those tools still decides whether the model calls the right one. See MCP tool descriptions that route the model.
A default that will not embarrass you
- Leave Auto on for mixed work, especially on a paid plan where that option is discounted.
- Pin GPT-5 mini or GPT-5.6 Terra when you want a known general model rather than a moving choice.
- Pin Haiku 4.5 or GPT-5.6 Luna for a pile of small edits.
- Pin GPT-5.6 Sol or an Opus-class model when you are debugging across files or choosing an architecture.
- Pin a Fable, Astra, Opus 5.5, or Kimi K3 class model only for a long agent run, and only after you know your plan includes it.
Credits follow token price, not the label "smarter." The rate table lives in GitHub's models and pricing page, and the switcher is documented for chat and for inline suggestions. Change the model when the task changes. A reasoning model left on for the rest of the afternoon is a billing choice, not a quality choice.
Keep reading
Indirect Prompt Injection: Tool Output Is Not Instructions
A retrieved document, an email, or a tool result can tell the model to take an action. Delimiters do not stop it. The tool allowlist after untrusted text does.
Copilot Auto Model Selection Now Has Tiers: Efficiency, Balance, Intelligence
GitHub Copilot's Auto picker gained efficiency, balance, and intelligence tiers in September 2026, alongside GPT-6 Astra and GPT-6.1 Sol. How the tiers change the cost-quality trade, and what admins control.
Put the Workflow in Code When the Agent Keeps Improvising
Which steps belong in a fixed program and which belong in a model, so an agent stops skipping the check that production depends on.
MCP Resources vs Tools vs Prompts: Stop Putting Everything in a Tool
The three primitives in the Model Context Protocol, what the model actually sees, and a rule of thumb for Dataverse, files, and internal APIs.
Evaluating LLM Applications: Getting Past 'It Looks Good to Me'
Shipping LLM features on vibes works until a prompt tweak silently breaks ten other cases. Here is how to build evals that catch regressions before your users do.
Getting Structured Output From an LLM Without the Heartbreak
Asking a model to 'return JSON' and parsing the result is how you get 3 a.m. pages. Tool schemas, constrained decoding, and validation turn a probabilistic text generator into a reliable API.
Newsletter
New posts, straight to your inbox
One email per post. No spam, no tracking pixels, unsubscribe anytime.
Comments
- No comments yet. Be the first.