Understand AI coworkers

Open-weight vs proprietary AI models: what a business team should care about

Justin team

·

·

7 min read

Justin blog header: Open-weight vs proprietary AI models, what a business team should weigh

Short answer: Open-weight AI models publish their trained weights, so anyone can download and run them; proprietary models are available only through the vendor’s own app or API. For a business team buying an AI product rather than building one, the model matters less than the context and controls around it. Choose on how your data is handled, quality on your own tasks, and cost at your volume.

What’s the difference between open-weight, open-source and proprietary models?

The difference is what the maker publishes. A model’s weights are the billions of numbers it learned in training; in practice, they are the model. A proprietary model keeps its weights private, so you reach it only through the vendor, as with GPT, Claude or Gemini. An open-weight model publishes them for download under a license, as with families such as Llama, Qwen, DeepSeek and Gemma.

Open source is a stricter bar. The Open Source Initiative’s definition asks for the weights, the complete code used to train and run the model, and enough detail about the training data for a skilled person to rebuild something substantially equivalent (OSI, Open Source AI Definition 1.0). Most “open” models publish weights only, which is why open-weight is the accurate term. Licenses also differ, so read one before assuming commercial use is allowed.


Proprietary

Open-weight

Open source (OSI definition)

What’s published

Nothing; access by app or API

The trained weights, under a license

Weights, full code, training-data details

Where it runs

The vendor’s servers

Your servers, or any host you choose

Your servers, or any host you choose

How you can change it

Only through the vendor’s options

Fine-tune it yourself

Fine-tune it or rebuild it

Examples

GPT, Claude, Gemini

Llama, Qwen, DeepSeek, Gemma

Fully open research models, such as OLMo

How big is the gap between open and closed models today?

The gap is small, and it moves. In March 2026 the best closed model was ahead of the best open-weight model by 3.3% on the Arena leaderboard, up from a 0.5% lead in August 2024 (Stanford HAI AI Index, 2026). The same report found the leading models so tightly clustered that competition is shifting toward cost, reliability and performance in specific domains, and that hard benchmarks now saturate within months.

For a buyer, a gap this small and changeable is no basis for choosing a product, and a benchmark score says little about whether a model writes your renewal summary the way your team needs.

Open-weight vs proprietary models for business: what to weigh

Weigh three things, in this order: how your data is handled, quality on your own tasks, and cost at your volume. The weights only matter if you plan to run the model yourself, which makes sense in one corner of the diagram below.

2x2 matrix: only teams with engineers and a firm in-house data rule should self-host open weights; most should buy a product

Figure: Most small teams land bottom left, where the model’s license matters less than the vendor’s data terms.

Data handling depends on who runs the model, not on whether its weights are public. An open-weight model served through someone else’s API is just another vendor holding your data. A proprietary model under a contract that rules out training on your data and sets a retention period can be the safer choice. Whatever you decide belongs in your AI usage policy.

Cost follows the same logic: open weights are free to download, not free to run, because servers and upkeep land on your team. A product folds that into a seat, credit or usage price; the AI glossary for business teams explains tokens and credits.

How to judge quality on your own tasks

The only quality test that counts is your own work, graded against an answer you already trust. Our CS ops lead keeps validation fixtures for every AI skill the team runs: the real input from a past case and the output that was actually delivered. Any change reruns all of them, so a fix for one customer can’t quietly break another. The same set is the fairest test of a new model; the diagram below shows the loop.

Loop: collect past tasks with trusted answers, run the product, check banned strings, compare, and rerun on every change

Figure: A test built from your own past work outlasts any model name.

The ops lead also keeps a banned-strings list (leaked field names, the model’s verbal tics, anything that makes the expert wince) and checks every output against it. It grew to 14 entries and stopped; the first phase of output had hundreds of hits, the second had zero. The ops lead calls it “the cheapest QA gate you’ll ever write,” and it catches a new model’s habits without anyone rereading every draft.

Why do AI products switch models behind the scenes?

AI products change models for ordinary reasons: a newer model is better or cheaper, a provider retires an old version, or a task needs speed more than depth. Many route work across several models at once, as our CS ops lead’s own jobs do, with a stronger model for decisions and a cheaper one for prose, a split the guide to how AI coworkers work describes.

So the model name on a pricing page is a snapshot, and sometimes you can’t choose it at all. In the scheduling tool our CS ops lead uses, the model can’t be set per job. The ops lead’s answer is to design for change: one job asks the model for one-line summaries and falls back to a plain template if the call times out, so a slow model never blanks the report.

Questions to ask a vendor about its models

Ask

Why it matters

A good answer

Which models do you use, and will you tell us when that changes?

Output can shift after a swap

Named models or providers, plus change notes

Is our data used to train any model, yours or your provider’s?

Answering and training are separate uses

No, in writing, including subprocessors

Where is our data processed, and how long is it kept?

Contracts or regulation may require it

A named region and a retention period

Can we rerun our own test tasks after a change?

Benchmarks don’t measure your work

Yes, in the same workspace with the same data

What happens when a model call fails?

A silent failure looks like a normal answer

A visible error, never a half-finished result

What does pricing meter: seats, credits or tokens?

Heavier reasoning can cost more per task

A clear unit, and what makes a task heavier

FAQ

Is open-weight the same as open source?

No. Open-weight means the trained weights are published under a license. Open source, under the Open Source Initiative’s definition, also requires the complete code and detailed information about the training data. Most well-known “open” models are open-weight, not open source.

Are open-weight models free?

They are free to download under most licenses, but not free to run. Servers, setup and upkeep cost money and engineering time, which a product folds into its price.

Should a small business use open-source or closed AI models?

Usually neither directly. Self-hosting an open model makes sense when you have engineers to run it and a firm reason to keep data in-house, because open weights are only safer when you run them yourself. Otherwise, pick a product on its data terms, its results on your past tasks, and its price.

Does it matter which model an AI coworker uses?

Less than most buyers expect. The same model gives very different results depending on the context it gets, such as documents found through retrieval, the tools it can reach, and the checks around it. Test the product on your own past tasks and ask to hear when the model changes.

Doing this with Justin

Justin, the AI coworker for Slack, uses large language models, and its output can be wrong or incomplete, so check anything consequential before you act on it. For harder work, typing \ultra turns on deeper reasoning for the rest of the conversation; it raises capability, not permissions. There are no seats: credits are pooled across the team and metered on model usage, so heavier work uses more. The free trial includes the whole team and every channel with no card required, which is enough to run your own test: five past tasks with answers you trust, asked in one channel.

Start free

Related reading