Glossary

Open-weight models

Models whose parameters you can download and run yourself, on your own infrastructure.

Open-weight models, Llama, Mistral and others, can be run entirely within your environment, including air-gapped. There is no per-token cost and no data leaving your boundary.

In exchange you take on hosting, scaling and evaluation work that a hosted API absorbs.

The total cost comparison is rarely as favourable as the zero per-token price suggests. Once you account for GPU capacity, the engineering to serve and scale it, and the evaluation work a hosted provider absorbs, self-hosting wins clearly at sustained high volume and much less clearly below that.

Commonly misunderstood: Open weights is not the same as open source. Most of these models carry licences with real usage restrictions worth reading.

Related terms, in context

The concepts you almost always meet alongside open-weight models.

Data residency
The requirement that data remains within a specific country or jurisdiction.
Fine-tuning
Further training a base model on your own examples, to fix style, format or narrow task behaviour.
Quantisation
Reducing the numeric precision of model weights to cut memory and increase speed.

Where this shows up in our work

Open-weight models is not an abstraction for us. It is a decision we make on live projects. It shows up most directly in custom model fine-tuning, ai infrastructure & mlops, where getting it wrong has a cost someone can measure.

If you are evaluating a vendor on this, the useful question is not whether they can define the term. It is what they measure, what they would refuse to do, and what happens in their system when the assumption behind open-weight models stops holding.

Questions

What is Open-weight models?

Models whose parameters you can download and run yourself, on your own infrastructure.

What do people get wrong about open-weight models?

Open weights is not the same as open source. Most of these models carry licences with real usage restrictions worth reading.

Does Orqent Labs build this?

Yes, Custom Model Fine-tuning and AI Infrastructure & MLOps. We work across India, covering all 19,238 PIN codes remotely.

Building something that involves open-weight models?

We will tell you honestly whether it is the right approach for your problem.

Or email bd@dtrasglobal.com · call +91 74118 77878