Koa is post-trained on nearly three decades of CRM data to run multistep agent workflows. Salesforce says it makes three times fewer errors — on its own benchmark.
The story here is not a model launch, it is where the errors live. Every AI failure a seller has already sat through — the follow-up booked against the wrong contact, the stage moved on a call that never happened — came from a general model guessing at a CRM object nobody taught it. Salesforce's claim is that a model trained on CRM structure specifically guesses less. If that holds up, the argument for letting software write to your pipeline gets stronger, and the thing standing between you and a wrong record stops being the model's caution and starts being whether you still check it. None of this reaches your org before winter, which makes now the useful window: find out what your error rate is today, while you still have a before to compare against.
A reasoning model is one trained to work a problem in steps — decide what to do first, pick a tool, check the result — rather than produce an answer in a single pass. Post-training means Salesforce took an existing NVIDIA model, Nemotron 3 Super, and trained it further on its own material instead of building a model from scratch.
A synthetic dataset means that extra training material was generated rather than taken from real customer records. Salesforce says it was "modeled on" nearly thirty years of CRM deployments, which is not the same as being those deployments. And a pilot here is not a beta you can ask to join — Salesforce picks the customers.
The short version: Salesforce taught someone else's model to be good at CRM chores in particular, a handful of customers are testing it now, and everyone else waits for winter.
Salesforce and NVIDIA announced Koa on September 15, 2026, described as "Salesforce's first CRM reasoning model for Agentforce, built on NVIDIA Nemotron." Salesforce says it is "purpose-built to help agents reason through complex, multistep workflows and use the right tools." It was made by "post-training NVIDIA Nemotron 3 Super with a proprietary synthetic dataset modeled on enterprise knowledge from nearly three decades of CRM deployments."
Announced: September 15, 2026, by Salesforce and NVIDIA.
Availability: "Available to select pilot customers now in Agentforce; general availability expected winter 2026 in U.S. regions."
Price: Not stated.
Regions: U.S. regions at general availability. Not stated for the pilot.
The performance claim is that Koa "matches or exceeds leading model performance on CRM actions with three times fewer errors." That was measured on Salesforce's own CRM benchmark, covering "real-world tasks like updating an opportunity, routing a case, or scheduling a follow-up." Salesforce also describes Koa running "in an agent in Slack that helps employees find information and complete everyday tasks."
What it doesn't do: The release names no limitations at all, which is itself the thing to notice. It does not say which "leading model" Koa was compared against, does not publish the benchmark or define what counts as an error, and gives no baseline error rate — so "three times fewer" has no stated starting point. It also does not describe anything a seller opens or types into: every task named is an agent action inside Agentforce, not a chat window a rep uses.
SDR — Nothing in your day changes, and it is worth knowing that before someone tells you otherwise. Every task Salesforce names is a record operation: update an opportunity, route a case, schedule a follow-up. None of it is research and none of it is writing the message. If your team runs Agentforce, the one thing to watch is lead routing starting to happen without a person in it — so find out now who you appeal to when a routing decision is wrong.
AE — Your exposure is the opportunity record, because that is the object Salesforce named first. Take your three largest open deals this week and write down by hand every field an agent would have to update after your next call. By the time Koa reaches your org you will already know which of those fields are judgment calls no model should be setting, and which are clerical work you should be glad to hand over.
Manager — "Three times fewer errors" is a line your team will quote at you. Ask the two questions that make it mean anything: three times fewer than which model, and what counts as an error? Salesforce has published neither. Then get your own number — over the next month, count how often a CRM field your reps actually rely on turns out to be wrong. Without that baseline you will not be able to tell an improvement from a regression.
RevOps — This is a governance item on a known clock: pilot now, general availability expected winter 2026, U.S. regions. Use the gap. Decide which objects an agent may write to unattended, which need human confirmation, and what your audit trail actually shows when an agent changes a field. Then settle the harder question before the vendor settles it for you: whether a model that reasons better earns more autonomy, or the same amount with better logging.
The only defensible way to judge a CRM model is against your own baseline: take three live deals, build the pipeline review from the record alone, and list every fact the record does not contain. The pipeline-health and CRM-hygiene prompts in Sales Operations, CRM & Productivity are built for that comparison.
Two minutes, once a week. What changed in AI, and what to run because of it.