When AI agents start maintaining AI systems, evaluation becomes the control surface
Palantir's AIP Evolve, now generally available, has agents propose model migrations, cost cuts and eval improvements. Why this only works with strong evals and human review.
Listen to this article · 5 min
AI-generated narration of the full article.

There is a persistent myth about production AI: that the hard part is getting it live. In reality, going live is when a second job starts. Models are updated or deprecated. Prices change. Faster and cheaper options appear. In September 2026 alone, OpenAI released three new models to its API and Google made a new default model generally available in Gemini Enterprise. Every one of those releases is a potential improvement, and a potential regression, for systems already in production.
Palantir’s AIP Evolve is interesting because it automates that second job.
What Palantir announced
According to Palantir’s announcement of September 8, AIP Evolve is generally available for enrollments that have AIP enabled and access to AI FDE. It coordinates fleets of AI FDE agents to improve AI systems built in Foundry. A team defines:
- A goal: model migration, cost reduction, latency reduction, evaluation score improvement, or a custom objective.
- A validation strategy: selected test cases or existing evaluation suites, with configurable scoring and an acceptable level of output divergence.
- Constraints: which types of change agents may propose, and how many iterations they may run.
The agents then work toward the goal and produce a proposal with the changes, validation results, output comparisons, supporting evidence and confidence assessments. People review it, and changes are merged through Palantir’s branching workflow. Palantir says early adopters have used it to cut AI costs, improve evaluation performance and migrate workloads to open-source models; those results are Palantir’s to substantiate.
What actually changes
The optimization loop of an AI system, try a cheaper model, adjust a prompt, compare outputs, decide, used to be manual and occasional. AIP Evolve makes it agentic and, potentially, continuous.
That shifts the human role. People stop making each change and start defining what an acceptable change is. The most important artifacts are no longer the prompt or the model choice. They are:
- The evaluation suite, which decides what “better” means.
- The divergence threshold, which decides how different outputs may be before a change is rejected.
- The constraints, which decide what the agents may touch.
- The review step, which decides whether evidence is sufficient.
- Goal and constraints
- Agents propose changes
- Validation against the eval suite
- Human review of the evidence
- Merge through branch review
The control surfaceHuman decisions
Evaluation becomes the control surface
This is the part many teams will underestimate. An optimization agent does exactly what the evaluation rewards. If the test suite covers the common cases and not the rare, expensive ones, a model swap can pass every check and still fail where it matters. Cost goes down, the score stays flat, and a class of edge cases quietly degrades.
In other words, automated improvement is only as safe as the evaluation it optimizes against. Teams without a serious evaluation suite should not read AIP Evolve as a shortcut. They should read it as a reason to build one.
Controlled change, borrowed from software engineering
The design choices in AIP Evolve mirror what mature software teams already do: bounded change types, iteration limits, evidence attached to every proposal, review before merge, and a branch-based workflow. That is the right model for AI systems in production, whichever platform they run on.
What any enterprise can take from this
You do not need Foundry to apply the lesson:
- Build and version an evaluation suite for every production AI system, including the rare and costly cases.
- Record baselines for quality, cost and latency, so any change can be judged against them.
- Treat model and prompt changes like code changes: proposed, tested, reviewed and merged, with a way to roll back.
- Write down why each change was accepted. Six months later, that record is what lets you trust the system.
- Plan for continuous migration. Model releases are now monthly events. Budget the time to evaluate them.
Production AI is not “deploy once and forget”. Whether the maintenance is done by people or by agents, the controls are the same. See our 2026 stack guide for how evaluation and lifecycle fit the rest of the system.