The AI development cycle is closing into a loop. CoreWeave Forge turns production runs into the next version
CoreWeave launched Forge, a development layer from training to agents, and added notebooks wired to experiments and evals. What changes when production evidence flows straight back into development.
Listen to this article · 6 min
AI-generated narration of the full article.

In most enterprises, the people who build an AI system and the people who watch it in production use different tools, different data and often different teams. Engineers train or configure in one place. Operations watches dashboards in another. When something goes wrong, someone exports a sample of bad cases to a spreadsheet, emails it, and a few weeks later a new version ships. The loop exists, but it runs through people’s inboxes.
CoreWeave’s launch of Forge on September 30, followed a day later by CoreWeave Notebooks, is a bet that this loop should be a product. It is not the only bet of its kind, and you do not need CoreWeave to act on it. But it describes clearly where AI platforms are heading.
What CoreWeave launched
CoreWeave describes Forge as a development layer that brings model and agent building across the lifecycle into one connected platform, open across models, frameworks and other clouds. Its stated loop is: run, observe, curate, improve, evaluate and repeat.
The pieces, as CoreWeave lists them, include Weights & Biases Models for experiment tracking; Agent Lens for production observability of agents; Sandboxes, now generally available, as isolated environments for agent tool use, reinforcement learning and evaluation; a Registry for versioned model checkpoints and agent configurations; Post-Training with serverless reinforcement learning, supervised fine-tuning and distillation; Inference with an RL Rollouts capability in preview; and ARIA, a now generally available coding agent that analyzes experiment and observability data. CoreWeave names MasterClass and Canva as early builders. Forge comes in Free, Pro and Enterprise editions. CoreWeave also publishes performance and cost claims for several components; we have not seen independent measurements and do not repeat them here.
The October 1 Notebooks post, by Julia Rose and Akshay Agrawal, adds reactive Python notebooks built on the open-source marimo project. They open inside a project already connected to its experiments, artifacts and evaluation results, and support Python and SQL. What is listed as future work, not shipped: publishing notebook views as persistent Workspace panels, hosting notebooks as data apps, more compute in the notebook sandbox, and ARIA writing notebook cells from a description.
- Run in production
- Observe traces and failures
- Curate cases into datasets
- Improve: prompt, config or training
- Evaluate against a fixed bar
- Release a versioned change
Where enterprise controls attach
What changes when the loop closes
The real story is not “another notebook.” It is the shrinking distance between a failure in production and the next version of the system. That is good for quality. It also moves three controls to places they were not before.
Production data becomes development data. Traces contain real prompts, retrieved documents and outputs. Once they flow into curation and training, the question is no longer only who can see them, but whether they can be used to change the system at all. Customer contracts, privacy notices and internal data policies often say something about this, and the answer differs by data source. Decide before the pipe is built, not after.
Every iteration is a change. A fine-tuned checkpoint or a new agent configuration is a change to a production system. In banking, model changes already sit under model risk management; the federal agencies revised that guidance in April, replacing SR 11-7, with expectations focused on organizations above $30 billion in assets. A faster loop does not remove validation. It means validation has to be as automated and versioned as the loop itself.
Evaluation becomes the gate, not the report. In a closed loop, the evaluation step decides what ships. That only works with a fixed evaluation set, agreed thresholds and someone who owns them. We made this argument when Palantir’s AIP Evolve let agents propose changes to AI systems: evaluation is the control surface. Forge’s design puts the same idea at the center of the platform.
Platform or assembly?
Many enterprises have assembled pieces of this loop already: an observability tool, an experiment tracker, an evaluation framework, a model registry, often on their main cloud. Microsoft shops typically have evaluation and tracing in Azure AI Foundry alongside Azure Machine Learning. The question is not whether to buy Forge, but whether your pieces are actually connected. Can an engineer go from a failed production trace to the evaluation set to the next release without exporting a file?
If the honest answer is no, the gap is the loop, not any single tool. As we argued in our piece on agent observability, monitoring that does not feed improvement is only half the job.
What to do now
- Trace one real failure end to end. Time how long it takes from a bad production answer to a fix in production, and count the manual handoffs.
- Write a data-use rule for traces: which sources may be curated into evaluation sets, which into training, and with what redaction.
- Freeze an evaluation set per use case, version it, and make passing it a release condition.
- Version everything that changes behavior: prompts, configurations, checkpoints and tool definitions, in one registry.
- Separate future features from shipped ones when you evaluate any platform, including this one.
The bottom line
AI systems improve fastest when production evidence flows straight back into development. CoreWeave is packaging that loop; others will follow, and some teams will keep assembling their own. Whatever the tooling, a closed loop needs three things to be safe: rules for what production data may teach the system, an evaluation gate with an owner, and versioned releases that someone can roll back.