AI Dev Weekly #29: Haiku 5.5, Mistral Large 4, Decisions API and Copilot
AI Dev Weekly is a Thursday series covering the AI developer news that changes what you can build, what it costs, and how you should operate it.
Last weekโs releases made agents more capable and persistent. This week made the supporting pieces more interesting: cheaper workers, a purpose-built decision endpoint, another large model to evaluate, and stronger local execution controls.
There is plenty to test, but availability needs careful reading. Haiku 5.5 is available now. Mistral Large 4 has an API preview, not downloadable weights yet. OpenAIโs Decisions API is a beta. GitHubโs local sandbox has reached general availability, while its desktop computer-use feature remains a preview. Those are different levels of readiness, not interchangeable launch labels.
1. Haiku 5.5 makes the small-worker tier worth revisiting
Anthropic launched Claude Haiku 5.5 on October 7. The model ID is claude-haiku-5-5, and it brings adjustable effort to the Haiku family. The target workloads include classification, summaries, context compaction, database queries, and delegated coding tasks.
The new standard token prices depend on prompt length. All figures below are dollars per million tokens:
| Token type | Prompts up to 100K | Prompts over 100K |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache reads | $0.01 | $0.05 |
| Cache writes | $0.125 | $0.625 |
For the shorter tier, input and output rates are 90% below Haiku 4.5. That is not a promise that every application becomes 90% cheaper. Longer prompts have different rates, and tokenization changes the number of billed tokens for the same text.
Anthropic also cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. Standard input and output pricing did not change. The launch announcement covers both updates.
Migration deserves its own test run. The Haiku 5.5 migration guide tells developers to recount prompts, replace explicit thinking budgets with adaptive thinking, remove sampling parameters, and stop relying on assistant prefills. Response handlers should select content blocks by type rather than assuming the first block contains the answer. Refusals also need an explicit handling path.
My take: This is a good week to isolate the small, repetitive jobs currently going to an expensive model. Give each one a narrow contract and a measurable acceptance test. A cheap summarizer that drops the one detail needed by the next worker is not cheap in practice. Compare accepted results, retries, and escalation costs across the whole workflow. Our agent cost monitoring comparison and Sonnet 5.5 guide provide useful starting points.
2. Mistral Large 4 is here as an API preview, with weights still to come
Mistral announced Large 4 on October 6. You can evaluate it through Mistral Studio and the API now; the company says weights will follow at the end of the month after additional red-team work. Calling it a model you can already download and self-host would be premature. See the official announcement.
The model card lists the API ID mistral-large-4, a 1M-token context window, multimodal input, structured outputs, and function calling. It describes a mixture-of-experts architecture with 1.05 trillion total parameters and 52 billion active parameters, plus a vision encoder.
Current launch rates are $0.68 input, $0.07 cached input, and $2.09 output per million tokens. The listed regular rates are $1.36, $0.14, and $4.18 respectively. Mistralโs changelog specifies that the 50% launch discount lasts two weeks. Build a cost comparison using regular pricing as well as the temporary offer.
My take: Treat this as an opportunity to collect evaluation data, not a completed self-hosting decision. API access lets you test multilingual documents, visual inputs, tool selection, and long-context retrieval before the weight release. It does not establish what hardware, license, or operational setup a future local deployment will require.
I would use a fixed workload with known answers and record missing evidence, incorrect tool arguments, and recovery behavior. Long context is especially easy to overrate: accepting a large document is not the same as reliably finding its decisive paragraph. Our Large 4 API reference covers the preview details, while the Mistral model guide supplies the broader lineup context.
3. OpenAI Decisions separates classification from text generation
OpenAI added the Decisions API on October 6, according to its API changelog. The public beta uses gpt-6-luna through POST /v1/decisions.
Instead of asking for a paragraph and parsing it, you define a typed question. A predicate estimates whether a condition is true; a choice evaluates supplied categories; a score evaluates ordered levels. Scores are probability-weighted averages, so they can fall between levels rather than always returning an integer.
Pricing is input-only at $0.10 per million tokens, with no output, cache-read, or cache-write charges. Regional processing premiums and long-context multipliers still apply. Independent questions can share one request; questions that depend on an earlier answer need separate requests. This is not a replacement for arbitrary structured output or tool calling. The Decisions guide explains those boundaries.
My take: This is more interesting than another JSON-formatting option. Many applications need a decision before they need prose: which queue should receive this ticket, does this document meet a criterion, or should a human review this result? A narrow endpoint makes that boundary easier to express and measure.
The probability field still needs calibration on your data. I would keep a labeled test set, choose thresholds according to the cost of mistakes, and send uncertain cases to review. A false negative that hides a critical issue is not equivalent to a false positive that adds one item to a queue. Use our agent reliability evaluation guide to structure that test and AI fallback patterns to design the failure path.
4. Copilot adds desktop reach and makes local sandboxing generally available
GitHubโs desktop computer-use preview, announced October 1 after the previous issue, lets Copilot CLI and the Copilot app interact with desktop applications on macOS and Windows. It combines accessibility information and visual interaction, with approval before controlling an application. Organizations can disable the feature. The computer-use announcement explains the permissions and controls.
Separately, local sandboxing reached general availability on October 7 across Copilot CLI, the Copilot app, and VS Code Agent Host on Windows, macOS, and Linux. We covered the preview in issue #27; this is the move to GA, not the first appearance of the feature.
GitHub says its Microsoft eXecution Container translates policy into native operating-system controls for local execution, including filesystem, network, and credential boundaries. Enterprise policy cannot be weakened locally. See the sandbox GA announcement.
My take: Desktop access and shell sandboxing solve different problems. Do not assume that a sandboxed command automatically makes every permitted GUI action harmless. An agent can still perform a mistaken action inside an application it is allowed to control. Test both boundaries separately, using disposable data and narrowly scoped accounts.
Before enabling desktop access on a daily workstation, decide which applications are in scope, what needs fresh approval, and how to stop the agent. Keep production administration and sensitive external actions out of the first experiment. Our Copilot app guide and agent sandboxing guide cover the workflow and containment questions.
Quick hits
- Copilot CLI can discover local models. Version 1.0.94-0 adds running Ollama models to the
/modelpicker. This does not download a model or automatically switch off telemetry. GitHub documentsCOPILOT_OFFLINE=trueseparately, and a configured remote provider can still receive prompts. Read the local-model announcement alongside our Ollama guide. - Ai2 released AstaBrief. The October 2 release makes the report-generation model and training data available for scientific synthesis. Its focus is evidence-grounded reporting, not merely fluent summaries. The Ai2 release post is worth reading for its separate treatment of citation support and overstated conclusions.
What I would test this week
Start with one bounded task rather than migrating an entire agent stack. Measure Haiku on an existing worker job, try Decisions on a labeled routing problem, or compare Large 4 against a frozen document workload. For Copilot, verify denied access as deliberately as successful access. A release becomes useful when you know where it works, where it fails, and how to reverse the change.
FAQ
Is Mistral Large 4 available to download now?
Not as of October 8. API public preview is available, while Mistral says the weight release is planned for the end of the month. Evaluate the preview without assuming local deployment is already possible.
Should I move every agent task to Haiku 5.5?
No. Start with bounded jobs that have objective acceptance criteria. Keep an escalation path for ambiguous work and compare total workflow cost, not just the rate on a pricing page.
Does the Decisions API replace the Responses API?
No. It targets narrow typed decisions. Keep generation, explanations, arbitrary output schemas, and tool orchestration on the appropriate generation or agent interface.
Does a local Copilot model mean everything is offline?
No. Model selection and offline configuration are separate concerns. Review the endpoint, telemetry settings, and any tools or remote services the workflow can still contact.
What is the main theme of this weekโs releases?
More deliberate allocation of capability. Use smaller workers for bounded tasks, explicit decision contracts for routing, and enforceable limits around execution. Bigger models remain useful, but they do not need to handle every step.