Model retirement is the new API deprecation, except it changes behavior, not syntax
Databricks retired Gemini 2.5 endpoints on October 2 and retires Claude Sonnet 4 pay-per-token on October 9. Third-party model lifecycles are now a production dependency to manage.
Listen to this article · 7 min
AI-generated narration of the full article.

When a REST API is deprecated, the failure is loud. Calls return errors, tests break, someone gets paged. When a model is retired, the failure can be quiet. The endpoint is replaced, the calls keep succeeding, and the answers are different. Your invoice classifier still returns a category. It is just not always the same category.
That is why this week’s routine documentation update from Databricks deserves attention from anyone running AI in production.
What happened
Databricks’ Foundation Model APIs documentation lists Gemini 2.5 Flash and Gemini 2.5 Pro as retired on October 2, 2026. The maintenance policy shows Gemini 2.5 Flash retiring on both pay-per-token and provisioned throughput, and Gemini 2.5 Pro retiring for provisioned throughput on the same date; the recommended replacements are Gemini 3.1 Pro or Gemini 3.5 Flash.
Next in line: Claude Sonnet 4 on pay-per-token retires on October 9, 2026, with Claude Sonnet 4.6 as the recommended replacement. DeepSeek V4 Pro (0813) and Thinking Machine Labs’ Inkling, the latter still in Public Preview, follow on October 30.
The policy itself is clear about the mechanics. A model moves from legacy (no longer offered to new workspaces) to deprecated (a retirement date 30 or 90 days out) to retired (no longer accessible). Provisioned throughput customers sometimes get longer windows: Meta Llama 3.1 405B retired on pay-per-token in February 2026 but stayed available on provisioned throughput until May.
One detail matters more than the dates. When a partner model provider gives less than a month’s notice, Databricks says it may temporarily redirect calls to a similar version to give customers time to migrate. Between March 26 and June 7, 2026, calls to Gemini 3 Pro were redirected to Gemini 3.1 Pro. That is a reasonable courtesy. It also means the model answering your production traffic can change without your code changing.
Why this is harder than API deprecation
An API deprecation changes the contract. A model retirement keeps the contract and changes the behavior behind it. A newer model is usually better on average, but production systems do not run on averages. They run on specific prompts, specific output formats, specific thresholds that someone tuned, and downstream code that parses what comes back.
What can change with a model swap:
- Output shape: JSON that used to be compact arrives with commentary, a field gets renamed, a list gets reordered.
- Judgment: a classifier’s borderline cases fall differently; an extraction step becomes more or less conservative.
- Cost and latency: a “recommended replacement” from a higher tier can double cost per task, or add seconds to a step with a latency budget.
- Refusals and safety behavior: a newer model may decline inputs the old one handled, or the reverse.
None of these trigger an error. All of them can break a business process.
Treat models like dependencies with an expiry date
- Retirement notice
- Candidate replacement
- Regression evals on your tasks
- Cost and latency check
- Approved release
- Rollback path kept
EvidenceRelease control
Keep an inventory with dates. Every production use of a hosted model should be listed with provider, platform, pricing mode and the published retirement date. Most teams can say which models they use; few can say which ones expire this quarter.
Abstract the model name. Reference models through configuration or a gateway route, not hard-coded strings scattered across services. We argued in our analysis of the C1 LLM Gateway that the gateway is becoming a policy point; model lifecycle is one more policy it can enforce.
Own a regression set. The only reliable way to approve a replacement is to run it against your own tasks with your own scoring. This is the point we made in our piece on Palantir AIP Evolve: evaluation is the control surface that makes change safe.
Pin where you can, and know when you can’t. Pinning a version buys predictability until the retirement date. Redirects mean a pin is a request, not a guarantee.
Decide fallbacks in advance. If a model disappears or degrades, which model takes over, at what cost, and who approves the switch?
Treat the swap as a change under your normal controls. For U.S. companies with SOC 2 reports, a model replacement in a customer-facing workflow is a production change and should leave the same evidence as any other: ticket, test results, approval. The NIST AI Risk Management Framework makes the same point in its own vocabulary: monitoring and managing an AI system includes managing the changes made to it.
What to do this week
- Search your code and configs for Claude Sonnet 4 on Databricks pay-per-token. October 9 is six days away.
- List every hosted model you depend on with its retirement date, and put the dates in the same calendar as certificate expirations.
- Build a small regression set per use case, even 50 representative cases with expected outputs, and run it before every model change.
- Ask your platforms how they handle short-notice retirements: redirect, error or fallback, and whether they notify you when traffic is redirected.
- Budget for migrations. Each model you depend on implies at least one migration a year.
The bottom line
Hosted models are now dependencies with published expiry dates, and the replacement is never a drop-in in the strict sense. The teams that will handle this calmly are the ones that already treat a model change as a release: inventory, evaluation, approval and a way back.