Updated September 27, 2026: rewritten as a follow-up to our July post. The first version listed specific model versions, which went stale within weeks; this one explains what staying model-agnostic looks like in practice.
At the end of July we wrote that there's a new AI model every month, and you are not supposed to keep up. The releases kept coming, as they were always going to, and the most common reply we got was a fair one. Fine, I'll stop chasing models, but then what does the thing I actually use do when the models change underneath it? This is our answer, and it's more concrete than the first piece, because "don't worry about the model" is only good advice if something is doing the worrying for you.
"Model-agnostic" gets thrown around as a slogan, so here's what it means in practice for us. A single customer answer isn't one call to one model. It's several jobs in sequence: working out what the customer is actually asking, finding the right passages in your content, writing the reply, and checking it before it goes out. Each of those jobs goes to the model that's best at it, drawn from frontier labs such as Anthropic, OpenAI, Google and Meta, whose Muse models we recently added, and from specialists such as TypeSafe, whose Jev model does one narrow job, recognizing what a customer is trying to do, and does it very well. No single lab is best at every job, so no single lab gets every job, and the roster keeps growing as better options appear.
That design changes what a model launch means. When a lab ships something better at one of those jobs, it can take over that job, and your Guru gets better without you doing anything, no migration, no retraining, no project. The leapfrogging that the headlines treat as drama shows up for you as a quiet improvement. A business that built directly on one vendor experiences the same launch completely differently: either it rebuilds to adopt the new model, or it watches others get the benefit first.
The less glamorous half of this is outages, and for a business it matters more than launches. Every major provider has bad days, including the biggest ones. When the model handling a step fails or stops responding, the system hands that step to a model from a competing vendor and carries on, so a provider having a bad afternoon doesn't turn into silence for your customers. A tool built on a single vendor's API has no such option. That vendor's outage is its outage, and so it's yours.
There's a quality benefit hiding in here too. Different model families have different blind spots. When a draft fails our checks and gets retried on the same model, it tends to repeat the same mistake, the way a person rereading their own sentence keeps missing the same typo. Retrying on a model from a different lab breaks that loop far more often. Vendor diversity isn't only insurance against downtime, it's part of how the answers stay accurate.
So here's the part of the model race that actually matters to you, which the July piece only pointed at. It isn't which model is ahead this month. It's whether the system you rely on can use whichever one is ahead, survive whichever one goes down, and do both without asking you to notice. That's a question you can put to any vendor, and the answer tells you more than any benchmark chart.
If you're evaluating an AI tool for your business, three questions cut through the noise. Which models does it use, and can it change them without me doing anything? What happens to my customers when that model's provider has an outage? When a better model ships, how long before I get the benefit, and what do I have to do? Vague answers to any of the three usually mean the product is one vendor's model with a friendly face, and that its roadmap is really that vendor's roadmap.
This is how we built ours, so we're hardly neutral, and you should weigh that. But the reasoning doesn't depend on whose product you choose. If the best model changes every few weeks and every provider has the occasional bad day, then the only durable position is one where neither of those facts reaches your customers. You can get there by building the orchestration yourself, which is a serious engineering commitment, or by choosing a platform that already treats models as swappable parts.
The models will keep leapfrogging each other, and we'll keep reading every release so you don't have to. Let the race be someone else's job. Just make sure it's actually someone's job, and if you want to see how we handle it, here's how our vendor-agnostic architecture works.



